本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新1028篇论文,其中:

  • 自然语言处理118篇
  • 信息检索19篇
  • 计算机视觉186篇

自然语言处理

1. 【2610.10536】Decoupling Exploration from Optimization in RLVR

链接:https://arxiv.org/abs/2610.10536

作者:Saif Punjwani,Micah Goldblum

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Modern language models, undergo reinforcement learning, Modern language, language models undergo, models undergo reinforcement

备注: 20 pages, 16 figures, 9 tables. Code: [this https URL](https://github.com/SaifPunjwani/Exploration-Distillation) . Checkpoints: [this https URL](https://huggingface.co/SaifPunjwani/expdis-checkpoints)

点击查看摘要

Abstract:Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.

2. 【2610.10533】EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory

链接:https://arxiv.org/abs/2610.10533

作者:Hongru Cai,Ran Wei,Wenjie Wang,Chengfa Wu,Ning Song,Yongqi Li,Wenjie Li

类目:Computation and Language (cs.CL)

关键词:limited additional computation, DeepSeek Engram, Engram use input, large language models, expanding the capacity

备注:

点击查看摘要

Abstract:Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.

3. 【2610.10526】Rephrase Before You Act: Characterizing and Mitigating Language Sensitivity in Vision-Language-Action Models

链接:https://arxiv.org/abs/2610.10526

作者:Mikey Watts(Independent Researcher),Yuchen Cui(University of California, Los Angeles)

类目:Robotics (cs.RO); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:strikingly sensitive, vision-language models, language robustness, LIBERO stove, tasks

备注: 9 pages, 8 figures, 3 tables. Project page: [this https URL](https://sttawm.github.io/rephrase-before-you-act)

点击查看摘要

Abstract:Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: $\pi_{0.5}$ turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a $\pi_0$ checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we score many phrasings of a few training tasks, have a large language model distill the evidence into ten to twenty rephrasing rules, and at deployment rewrite each incoming instruction once under these rules. The rules improve the frozen $\pi_0$ by 16 to 27% relative on twelve held-out tasks across adversarial, VLM-generated, and human-generated phrasings, with gains concentrated on out-of-distribution tasks. The pipeline replicates on $\pi_{0.5}$ and LIBERO, lifting in-finetune success from 93.6% to 97.8%. The method requires no retraining and no per-step verification, and applies zero-shot to unseen tasks and instructions. Project website: this https URL

4. 【2610.10508】Your Prompt Should Do More: Effects of Retrieval Instructions in Embedding Models

链接:https://arxiv.org/abs/2610.10508

作者:Amanda Myntti,Jenna Kanerva,Veronika Laippala,Filip Ginter

类目:Computation and Language (cs.CL)

关键词:received increasing attention, Prompted embedding models, recently received increasing, Prompted embedding, increasing attention

备注:

点击查看摘要

Abstract:Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.

5. 【2610.10506】Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models

链接:https://arxiv.org/abs/2610.10506

作者:Daniel Robert Kling Alexander,Catherine Louise Kling

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Economics (econ.GN)

关键词:policy is worth, user should choose, questions now put, put to large, option a user

备注:

点击查看摘要

Abstract:Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of \$75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model's answers are coherent, not that they are correct.

6. 【2610.10455】PHRBench: A Behavioral Evaluation of Post-Hallucination Reasoning in LLMs

链接:https://arxiv.org/abs/2610.10455

作者:Linghao Meng,Feng He,Xuan Yang,Junyuan Mao,Pinze Ren,Deqing Mu,Hesen Yang,Qiankun Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:multi-stage LLM systems, information can propagate, propagate through multi-stage, Hallucinated information, reasoning

备注:

点击查看摘要

Abstract:Hallucinated information can propagate through multi-stage LLM systems and become part of the context for subsequent reasoning. Existing studies of post-hallucination reasoning (PHR) mainly characterize changes in final outcomes and aggregate reasoning dynamics, leaving how models resolve hallucinated premises at the response level insufficiently understood. In this work, we introduce PHRBench, a controlled benchmark for behaviorally structured PHR across four domains and 18 large language models. PHRBench characterizes each reasoning trajectory independently of final-answer correctness through Hallucination Compliance, Hallucination Avoidance, and Heuristic Correction, and defines an insightful trajectory as successful correction that ultimately reaches the correct answer. Across 4820 controlled instances, we find that successful recovery remains relatively rare and is associated with more frequent belief updates along the reasoning trajectory. We further find that properties of the hallucinated prompt contain substantial predictive signal for successful recovery, with a lightweight predictor achieving an AUROC of 0.847. These findings provide a behavioral view of post-hallucination reasoning, characterizing how LLMs resolve erroneous context and when successful recovery is likely to occur.

7. 【2610.10444】RunningTab: Direct Workspace Interaction with Environment-Side Tabs

链接:https://arxiv.org/abs/2610.10444

作者:Jinheon Baek,Soyeong Jeong,Yumin Choi,Dongsu Han,Sung Ju Hwang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:knowledge work produces, direct workspace interaction, knowledge work, work produces, direct workspace

备注:

点击查看摘要

Abstract:Much knowledge work produces new deliverables from files a workspace already holds, and LLM agents are beginning to take such work over. Through direct corpus interaction, an agent can search and read any of those files from a terminal with no indexing, and producing a deliverable from many of them in this way is what we call direct workspace interaction (DWI). Reaching the files, however, is only half the task: nothing keeps track of what the task asks for, what has been read, and what was listed but never opened, all of which slip through the context window without leaving a trace, so an agent may extract a figure and still deliver a report without it. To address this, we present RunningTab, a framework that equips direct workspace interaction with an environment-side tab: a per-task record of what the task still owes, kept by the environment alongside the agent. Specifically, the agent adds its requirements, while the environment records every file read as an excerpt with its provenance and every listed but unopened file as a candidate; the agent can then see each requirement beside its best-matching excerpts and top unopened candidates, resolve it against matching content or set it aside with a reason, and, should it try to finish with requirements still open, receive them in a finish check. We validate RunningTab on three benchmarks with three LLMs, where it consistently outperforms plain DWI and baselines that keep the record in the model, while its tab usually holds the values a deliverable needs once seen.

8. 【2610.10426】CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

链接:https://arxiv.org/abs/2610.10426

作者:Jixuan Chen,Jiaxin Zhang,Qinyuan Ye,Yada Pruksachatkun,Haoxiang Zhang,Jingming Zhuo,Yifan Zhang,Yutong Dai,Juntao Tan,Xiangyu Peng,Silvio Savarese,Zeyuan Chen,Lianhui Qin,Chien-Sheng Wu

类目:Computation and Language (cs.CL)

关键词:Terminal-agent capability depends, handles error recovery, Terminal-agent capability, capability depends jointly, binds tools

备注: Preprint. 32 pages, 7 figures, 17 tables

点击查看摘要

Abstract:Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.

9. 【2610.10422】Which Rollout Taught It That? BehaviorTrace and the Limits of Training-Data Attribution in Online RL

链接:https://arxiv.org/abs/2610.10422

作者:Amit Nautiyal

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:reinforcement learning teaches, reinforcement learning, learning teaches, teaches a language, rollouts that taught

备注: 11 pages, 2 figures, 4 tables. Code and data: [this https URL](https://github.com/AmitoVrito/BehaviorTrace)

点击查看摘要

Abstract:When reinforcement learning teaches a language model a new behavior, can we find the training rollouts that taught it? And when an attribution method says it can, how do we know the answer is real? We study both questions on online RL fine-tuning with GRPO, using a planted behavior with a known cause. We release BehaviorTrace, an open evaluation harness that combines full-gradient sketching, the planted-behavior setup, and controls for gradient magnitude, fluency, headroom, and variation across seeds and generation draws. Across three seeds on Qwen2.5-1.5B, much of the apparent attribution signal comes from confounds. A control that ranks training steps by gradient size alone, with no behavior target, reaches 4.2 to 4.5 times chance and matches or beats the best targeted estimator on two of three seeds. At saturated checkpoints, model fluency predicts the behavior label at least as well as every gradient method we compared it with. Once fluency is controlled, the per-rollout results change from seed to seed and from one generation draw to the next, so a single run cannot settle the question. One signal does hold on all three seeds. The gradient of the trigger tokens aligns with a target built where the behavior actually occurs. We turn these findings into a checklist for evaluating attribution in RL. We test existing estimators, including GAS (renormalized TracInCP) and a TRAK-style estimator, and do not propose a new one.

10. 【2610.10411】raining Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

链接:https://arxiv.org/abs/2610.10411

作者:Yunxiao Zhao,Changxiao Cai

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:large language model, language model inference, accelerates large language, low-cost draft model, decoding accelerates large

备注:

点击查看摘要

Abstract:Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.

11. 【2610.10405】Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

链接:https://arxiv.org/abs/2610.10405

作者:Maverick Morales,Tomáš Dominik,Vermut Gao,Katrina Shirey,Paulius Rimkevičius,Aaron Schurger,Uri Maoz

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:reasoning artificial intelligence, artificial intelligence, key approach, approach to detecting, forms of misbehavior

备注: 20 pages, 9 figures, 3 tables. Code: [this https URL](https://github.com/Wakaranaino/token-spike-project) ; Data: [this https URL](https://doi.org/10.5281/zenodo.21895296)

点击查看摘要

Abstract:Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.

12. 【2610.10378】Document-Level Text Simplification in Estonian Using Large Language Models

链接:https://arxiv.org/abs/2610.10378

作者:Meeri-Ly Muru,Eduard Barbu

类目:Computation and Language (cs.CL)

关键词:text simplification involves, simplification involves transformations, Document-level text simplification, anaphora resolution, sentence-internal edits

备注: 12 pages, 2 figures, 2 tables. Published at LREC 2026

点击查看摘要

Abstract:Document-level text simplification involves transformations that go beyond sentence-internal edits, addressing discourse coherence, anaphora resolution, and cross-paragraph consistency. Despite advances in sentence-level simplification for high-resource languages, document-level simplification in morphologically rich, low-resource languages such as Estonian remains largely unexplored. This study presents a comprehensive evaluation of five state-of-the-art multilingual large language models (LLMs) for document-level simplification in Estonian. Three prompting strategies are examined: single-pass generation, pipeline-based modular agents, and guideline-augmented pipelines. The evaluation framework integrates automatic metrics assessing readability, semantic preservation, and discourse coherence, alongside a structured manual annotation protocol. The findings indicate that Gemini-2.0 and LLaMA-3.3 produce outputs with near-native fluency and strong meaning preservation, whereas other models display notable grammatical and semantic limitations. This work contributes novel document-level coherence metrics, evidence-based prompting strategies, and publicly available resources for reproducibility.

13. 【2610.10368】Input-Blind Controls Produce Substantial Oracle Headroom for Layer Programs in Multiple-Choice Evaluation

链接:https://arxiv.org/abs/2610.10368

作者:Yibei Guo,Rui Liu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:improve language-model inference, Adaptive computation aims, Adaptive computation, aims to improve, improve language-model

备注:

点击查看摘要

Abstract:Adaptive computation aims to improve language-model inference by tailoring execution to each input. For layer programs, oracle evaluations use known answers to estimate the potential gain from this flexibility, before a practical selector is available. However, a gain from selection does not by itself explain why the chosen programs help. This study examines this distinction using 32 layer-skipping and repetition programs on two models and 4,413 multiple-choice items. The analysis compares their gains over a fixed action selected without the evaluation prompt with those of input-blind perturbations at the same sites, re-evaluating selections on another prompt. With shared option order, the controls give 10.2-11.8 and 15.6-19.4 percentage points of headroom on Qwen3-4B-Base and Llama-3.1-8B, exceeding the real programs' 9.0 and 10.1 in all three random-direction draws per model. They match answer-change rate only, and the ordering depends on the menu: in post hoc comparisons, real programs lead on Llama's repeat-only menu in every draw. A smaller KL-calibrated comparison, including an input-dependent control, favours real programs in point estimate, with inconclusive corrected tests. Fixed letter offsets produce headroom of similar scale. Rotating options sharply reduces both families' headroom, while leaving positive real-minus-control differences of 1.4-2.3 and 3.7-4.5 points; their magnitudes and statistical support depend on further adjustments and the reference. A supplementary generated-answer test finds that search-selected programs keep a 26.0-point advantage over programs selected for other problems after rewording, without a placebo comparison. These results show that substantial headroom can persist across prompts with shared option order without establishing a benefit specific to the selected layer computation; neither ordering against these controls identifies that benefit.

14. 【2610.10332】Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision

链接:https://arxiv.org/abs/2610.10332

作者:Wenxi Gan

类目:Computation and Language (cs.CL)

关键词:Learning from large-model, complete recurring tasks, large-model demonstrations offers, complete recurring, calling a large

备注:

点击查看摘要

Abstract:Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4\% mean unseen task success with either TPD or action-only supervision, compared with 48.3\% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0\% to 67.7\% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9\% at 808 demonstrations. Shared-history analyses link part of TPD's local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.

15. 【2610.10318】Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

链接:https://arxiv.org/abs/2610.10318

作者:Himarsha R. Jayanetti,Sivakanesan Dhanushkanda,Shuai Hao,Michael L. Nelson,Michele C. Weigle

类目:Computation and Language (cs.CL)

关键词:real-time public sentiment, sentiment analysis, understanding their limitations, rich source, source of real-time

备注: 11 pages, 1 figure, 2 tables, accepted for publication at TPDL 2026

点击查看摘要

Abstract:Social media is a rich source of real-time public sentiment, but widely used sentiment analysis tools are often applied without understanding their limitations. In this study, we evaluate the inter-rater reliability of three bespoke sentiment analysis tools (TextBlob, VADER, and Twitter-roBERTa-base) and three large language models (LLMs: Qwen3-32B, GPT-OSS-120B, Llama-4-Maverick-17B) against six human raters across 100 tweets. We measured agreement using two statistical measures: Cohen's kappa for pairwise comparisons and Fleiss' kappa for multiple raters. Even among the human raters, our results showed only fair agreement, highlighting the subjectivity of sentiment analysis. Higher agreement was observed under the binary sentiment classification (negative vs. non-negative and positive vs. non-positive) than under the three-class classification across both humans and automated tools. The Twitter-roBERTa-base model showed the strongest alignment with human ratings, outperforming both bespoke sentiment tools and LLMs, particularly in distinguishing negative versus non-negative sentiment. LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, performing better in positive vs. non-positive classifications. Our findings underscore that domain-specific fine-tuning remains crucial for reliable social media sentiment analysis, and human-centered evaluation remains essential for establishing gold-standard labels.

16. 【2610.10304】SemanticFold: Latent Sequence Compression SeparatesLanguage Modeling, Decodability, and Reasoning

链接:https://arxiv.org/abs/2610.10304

作者:Mingyan Liu,Min Huang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:prompt prefixes preserves, large language models, language models rely, prompt prefixes, prefixes preserves

备注:

点击查看摘要

Abstract:We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B. We use a fixed-target protocol: a frozen prefix is executed natively or compressed, and both arms teacher-force identical continuation tokens. This design rules out target-selection explanations for likelihood changes. We examine five endpoint families: fixed-target negative log-likelihood, finite-label reasoning accuracy, linear probe accessibility, open-ended generation, and systems-level memory and latency. We find that compression moves these endpoints non-monotonically and that they do not share a single compression threshold. On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws. On SmolLM2 at R=1.2, the mean change is 0.013 higher than native. On both Pythia checkpoints, NLL is effectively unchanged. An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone. MLP-only, which applies the transform without shortening, achieves 0.082 lower NLL than Full SemanticFold. Linear probe accuracy and macro AUC change by less than 0.03 in absolute value across conditions, with confidence intervals crossing zero. We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.

17. 【2610.10276】PatchBench: Measuring Collateral Damage in Activation Patching

链接:https://arxiv.org/abs/2610.10276

作者:Alexi Canesse,Mathis Le Bail,Maël Jenny,Clément Elliker,Mahammed El Sharkawy,Sonia Vanier

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Cryptography and Security (cs.CR)

关键词:LLM safety patch, LLM safety, LLM, harmful, prompts

备注: Accepted to NeurIPS 2026 (Datasets and Benchmarks Track)

点击查看摘要

Abstract:An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing protocols primarily test whether models can be broken, while aggregate metrics (attack success, refusal rates, global capability) cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures inducing actionable harmful answers. Starting from 27,870 prompts from 37 public datasets, we curate 15,314 English prompts and query 8 open-source instruction-tuned models. Combining WildGuard filtering, pairwise Elo ranking, and manual verification, we retain a curated bank of 400 high-confidence jailbreak failures. We further introduce PatchBench-Local, an evaluation protocol testing whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants preserving malicious intent, benign prompts with matched structure, and benign prompts reusing key harmful terms. It evaluates harmful-neighbour correction and benign-neighbour preservation, distinguishing selective repair from broader local suppression. Evaluating four activation steering methods with PatchBench-Local and MMLU shows that global capability can remain nearly unchanged while local benign regressions are severe, confirming aggregate metrics miss important collateral damage. PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.

18. 【2610.10232】LLM Persuasion Is in the Eye of the Evaluation

链接:https://arxiv.org/abs/2610.10232

作者:Kamile Dementaviciute,Julija Vaitonyte,Tijl De Bie

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:Large language models, Large language, exceed human experts, shown to match, match or exceed

备注:

点击查看摘要

Abstract:Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman $\rho = 0.25$). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.

19. 【2610.10227】From Prompts to Trees: Effective LLM-Guided Tree Generation for Few-Shot Tabular Classification

链接:https://arxiv.org/abs/2610.10227

作者:Yue Qiu,Zekang Du,Yiqun Diao,Bingsheng He,Qinbin Li

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large Language Models, Language Models, possess rich world, impressive generalization capabilities, Large Language

备注: Accepted to EMNLP 2026 Main as an oral presentation. Code available: [this https URL](https://github.com/yueqiu0/LLMTree)

点击查看摘要

Abstract:While Large Language Models (LLMs) possess rich world knowledge and impressive generalization capabilities, their direct application to tabular data classification is hindered by high inference costs and limited interpretability. In contrast, decision trees are fast and transparent but often underperform in low-data regimes. In this work, we propose a novel framework that bridges these paradigms by distilling LLM knowledge into interpretable decision trees under a few-shot learning setting. Instead of directly prompting the LLM to generate full trees, which is often unstable and inefficient, we develop a three-stage paradigm that prompts the LLM to generate rules and organize the rules into a tree. Experiments on multiple real-world tabular datasets demonstrate that our method achieves superior accuracy and interpretability with significantly lower prompting overhead compared to existing baselines.

20. 【2610.10201】GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

链接:https://arxiv.org/abs/2610.10201

作者:Jingyao Zhang,Yun Li,Lu Han

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:reasoning requires translating, function reasoning requires, perceived spatial configuration, satisfies geometric constraints, curve satisfies geometric

备注: 15 pages, 1 figure, 7 tables

点击查看摘要

Abstract:Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.

21. 【2610.10179】Beyond Outcome Rewards: Constructing and Assigning Retrieval Credit for Search Agents

链接:https://arxiv.org/abs/2610.10179

作者:Wenyu Huang,Xinyu Hou,Pavlos Vougiouklis,Ruofei Lai,Jeff Z. Pan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large Language Models, enable Large Language, Language Models, Large Language, agents enable Large

备注:

点击查看摘要

Abstract:Search agents enable Large Language Models (LLMs) to iteratively retrieve and use information for complex multi-hop questions. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising approach for post-training such agents, but its reliance on sparse, outcome-based supervision can make credit assignment difficult and limit learning efficiency. In this paper, we systematically investigate how intermediate supervision can improve reinforcement learning for search agents. We study a range of reward-shaping and credit-assignment strategies that provide learning signals from intermediate retrieval steps. Building on these insights, we develop a training framework that combines intermediate signals with final outcome rewards to improve learning from multi-step search trajectories. Experiments across multiple benchmarks under matched training conditions demonstrate improvements in aggregate search-agent performance and show that both the choice of intermediate signal and where its credit is assigned affect training behaviour. These findings show that reward design and credit assignment are important design dimensions for training effective search agents.

22. 【2610.10154】HySPE: Positional Encoding via Symplectic Dual Shears

链接:https://arxiv.org/abs/2610.10154

作者:Zhongping Ji

类目:Computation and Language (cs.CL)

关键词:Symplectic Positional Encoding, grounding positional attention, Hyperbolic Symplectic Positional, Positional Encoding, non-compact symplectic transformations

备注: 11 pages

点击查看摘要

Abstract:We introduce Hyperbolic Symplectic Positional Encoding (HySPE), grounding positional attention in non-compact symplectic transformations. While canonical Rotary Position Embedding (RoPE) parameterizes the compact, elliptic branch of $\Sp(2,\R)$ via rotations, HySPE operationalizes its hyperbolic branch via a damped symmetric composition of dual shears, yielding a conformally symplectic contraction with two spectral decay rates per channel pair. To eliminate the exponential representation drift inherent to naive absolute factorizations, we diagonalize the operator in its invariant eigenbasis and introduce blockwise coordinate rebasing with adaptive centered execution. This guarantees length-independent numerical bounds while matching cached RoPE forward latency (7.21\,ms on an RTX 4090). On TinyShakespeare, HySPE-UltraLong maintains an invariant perplexity of 4.810 up to $16\times$ zero-shot extrapolation ($L=4096$), whereas RoPE degrades to 131.198. Scaled to a 51M-parameter subword Transformer on WikiText-103 ($L_{\text{train}}=512$), HySPE closely matches RoPE in-domain while robustly extrapolating to length 8192, reducing tail perplexity by 83.9\% over RoPE. While these controlled experiments establish HySPE's extrapolation robustness and numerical stability, evaluating its scaling behavior on large-scale foundation models remains an important direction for future investigation.

23. 【2610.10145】InterView-C: A Synchronized Multimodal Corpus of VR Avatar-Mediated Survey Interviews

链接:https://arxiv.org/abs/2610.10145

作者:Patrick Schrottenbacher,Leon Hammerla,Lydia Kleine,Doris Stingl,Alexander Mehler

类目:Computation and Language (cs.CL)

关键词:German multimodal corpus, survey interviews conducted, virtual reality, represented by avatars, German multimodal

备注:

点击查看摘要

Abstract:We present InterView-C, a German multimodal corpus of 27 survey interviews conducted entirely in virtual reality, with both interlocutors represented by avatars. The corpus aligns spoken interaction with synchronized behavioral data, including gaze, head and body movement, facial behavior, hand and finger tracking. Its reference transcripts and linguistic annotations provide a reliable interface between this multimodal spoken interaction and predominantly text-based NLP methods. This interface is important because automatically transcribing speech can distort linguistically relevant information, while downstream models trained on existing resources may additionally face transfer challenges when applied to transcribed spoken data. InterView-C therefore provides word-timed and manually post-edited verbatim transcripts for all 54 recordings, interview-item timings, questionnaire responses and negation cue and scope annotations for 1,422 sentences, 1,398 of them doubly annotated ({\alpha}=0.87 for cues; {\alpha}=0.81 for scopes). We demonstrate both challenges empirically: nine open-weight ASR systems disproportionately misrecognize short closed answers and number words, while negation models trained on existing corpora show lower and highly variable performance on our transcribed interviews than a model trained on the InterView-C annotations. InterView-C thus enables linguistic analyses of spoken interaction while retaining their alignment with rich multimodal behavior.

24. 【2610.10138】LLM4Impact: Integrating Heterogeneous Information for Scientific Impact Prediction

链接:https://arxiv.org/abs/2610.10138

作者:Yong Cao,Markus Flicke,Haoyu He,Katrin Renz,Andreas Geiger

类目:Computation and Language (cs.CL)

关键词:Predicting the future, newly published paper, newly published, Predicting, future impact

备注: 27 pages, 12 figures, 11 tables

点击查看摘要

Abstract:Predicting the future impact of a newly published paper is challenging because it must be inferred from heterogeneous evidence available at publication time. Existing approaches often rely on a single source of information or combine multiple sources without accounting for their different predictive roles. In this paper, we present LLM4Impact, an evidence-aware method for scientific impact prediction that learns to represent, integrate, and calibrate heterogeneous information. LLM4Impact combines semantic, graph, LLM, and temporal representations, and injects graph information into a frozen LLM through continuous prefix tokens. A context aware gating mechanism adaptively weights different evidence, while a separate calibration module accounts for domain and temporal variation in citation scales. We further construct a large-scale benchmark dataset with 2 million papers, leakage-safe point-in-time heterogeneous ego graphs, temporal splits, and both year-level and month-level citation targets. Experiments show that LLM4Impact consistently outperforms strong semantic, graph, and LLM based baselines, with a 10.13% reduction in year RMSE on the in distribution test set and a 6.87% reduction under out-of-domain distribution. Our results reveal that the value of such evidence is context dependent: different papers benefit from different sources, while domain and publication time affect how evidence translates into citations. This finding motivates adaptive evidence selection and context-conditioned calibration rather than simply richer representations. We will release our code, benchmark, and an interactive web demonstration upon publication.

25. 【2610.10118】YANchor-4B: Effective Long-Horizon Reasoning in O(N) Time with O(1) Memory

链接:https://arxiv.org/abs/2610.10118

作者:Huishan Ji,Hua Xu,Weiming Zhang,Qirui Ye

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:manageable generation cost, reasoning demands access, demands access, access to earlier, earlier information

备注: 24 pages, 8 figures. Code: [this https URL](https://github.com/RocoreMatrix/YANchor) ; Model: [this https URL](https://huggingface.co/HuishanJi/YANchor-4B)

点击查看摘要

Abstract:Long-horizon reasoning demands access to earlier information at a manageable generation cost. Full-history attention incurs growing storage and computation, while recurrent compression can lose precise details. Therefore, we present YANchor-4B, a general-purpose recurrent model that preserves crucial memory as ANchors for retrieval during subsequent reasoning. Beyond $O(N)$-time generation and $O(1)$ memory, YANchor enables effective long-horizon reasoning through its multidimensional memory mechanism. For example, on challenging math problems, it achieves 82.93% mean pass@1 on AIME 2024--2026 and 63.64% on HMMT, substantially outperforming linear-time, constant-state counterparts, including larger models. It also delivers several-fold higher batched long-generation throughput than Transformer and hybrid baselines on H100. Furthermore, evaluations across dozens of benchmarks demonstrate YANchor's superiority in general-purpose capabilities.

26. 【2610.10114】Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position

链接:https://arxiv.org/abs/2610.10114

作者:Xiaoran Liu,Ziwei He,Xipeng Qiu

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, Language Models, hybrid models, improve long-context efficiency

备注: 60 pages, 36 figures, 25 tables, under review

点击查看摘要

Abstract:The architectural design of Large Language Models (LLMs) is shifting from traditional full-attention-only models to hybrid models, which combine different attention modules to improve long-context efficiency and performance in length extrapolation and context extension. To explain why hybrid models work and how to design them better, we propose Mechanics of Long-Context Hybrid Models. As Part 1.1 of this series, we begin with hybrids of full attention and either sliding-window attention (SWA) or gated variants of linear attention (LA), represented by GLA and GDN. We first observe a Seesaw Effect in Context Extension: LA hybrids benefit more from long-context continual pretraining, whereas SWA hybrids perform better under length extrapolation. We attribute this behavior to differences in the positional inductive biases induced by these attention mechanisms. We find that SWA hybrids suffer from a Short-Context Learning Trap, Short-Window Weariness, and Long-Window Laziness, and require extended windows to enhance performance in continual long-context pretraining. For LA hybrids, we summarize the Matthew Effect of Hybrid Position Extrapolation and propose Sliding-Window Linear Attention, achieving 16$\times$ training-free length extrapolation while maintaining 100\% accuracy on NIAH-SK1 in 64k context length.

27. 【2610.10092】I would rather quit NLP than read another paper like this: The rise of antithesis in NLP papers

链接:https://arxiv.org/abs/2610.10092

作者:Olga Zamaraeva,Adrián Gude,Roi Santos-Ríos,Carlos Gómez-Rodríguez

类目:Computation and Language (cs.CL)

关键词:recent models fill, models fill papers, scientific writing., unnecessary antithesis, routinely for scientific

备注: 30 pages, 2 figures, 48 tables

点击查看摘要

Abstract:For better or worse, LLMs are by now used routinely for scientific writing.\footnote{This paper is no exception; we did use AI to assist with writing some of the sections (see Acknowledgments).} Many have noticed that recent models fill papers with unnecessary antithesis, stating over and over what the work does not do, in ways that do not contribute to its precision or quality of expression and annoy reviewers \emph{rather than impressing them}. We study the construction \emph{rather than} in ACL papers from 2019, ACL-style arXiv papers from 2026, and papers written by GPT models from the same titles and abstracts. Its rate in 2026 is seven times the 2019 rate, and higher still in the GPT papers. Two annotators, blind to the source, find almost no 2019 use \emph{annoying} and about one in ten 2026 uses; they seldom agree on which, yet about half of 2026 papers contain a use that annoys each of them. \emph{Annoying} uses present the rejected alternative less favorably than legitimate uses. Raters of preference data and open reward models favor the construction, and an instruction to be honest promotes it. We conjecture that it is a side effect of post-training on pairwise preferences, which credit a disavowal in a single response and cannot register its cost across a text.

28. 【2610.10091】ExperienceIndex: Artifact-Grounded Memory

链接:https://arxiv.org/abs/2610.10091

作者:Peter Baile Chen,Geoffrey X. Yu,Xinming Liu,Samuel Madden,Dan Roth,Jacob Andreas,Doug Downey,Michael Cafarella

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Knowledge-intensive tasks require, tasks require answering, court cases, Knowledge-intensive tasks, scientific literature

备注:

点击查看摘要

Abstract:Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solving traces, but they primarily focus on user preferences, factual attributes, or abstract reasoning patterns rather than persistent artifact-specific knowledge. We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces. ExperienceIndex stores two complementary forms of experience: (i) single-artifact experiences that summarize an artifact's contribution to prior tasks and (ii) artifact-pair experiences that encode structural relationships discovered during past reasoning. Integrated as lightweight middleware, ExperienceIndex uses an experience retrieval mechanism to guide agents toward the complete set of relevant artifacts for new tasks, improving both answer quality and efficiency. Across diverse corpora and agentic solutions with different search frameworks, ExperienceIndex delivers consistent gains, raising answer quality by up to 11.0 points and reducing online dollar cost by up to 50.5%. We further demonstrate two benefits: (i) cross-task generalization, where experiences accumulated from text-to-SQL tasks transfer to factoid QA tasks over the same artifact corpus, and (ii) teacher-student learning, where experiences from a stronger model enable a weaker model to reach comparable performance.

29. 【2610.10088】SkillSandbox: Skill Verification via Dynamic Scenario Synthesis

链接:https://arxiv.org/abs/2610.10088

作者:Serin Kim,Kwangwook Seo,Dokyung Song,Jinyoung Yeo,Dongha Lee

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Self-evolving agents distill, agents distill task-solving, encode incorrect procedures, distill task-solving experience, Self-evolving agents

备注:

点击查看摘要

Abstract:Self-evolving agents distill task-solving experience into skills for future reuse, but these skills can encode incorrect procedures or non-transferable knowledge. It is therefore critical to verify each skill's reusability: whether its guidance remains useful beyond the experience from which it was distilled. Such verification requires observing how a skill affects execution in new tasks, yet existing tasks may not expose the situations where the target skill can actually be exercised. To construct such situations, we propose SkillSandbox, a framework that dynamically synthesizes a task and its environment for each skill that are skill-relevant yet novel. A Proposer specifies the conditions to preserve and the source-specific details to vary, a Builder constructs an executable scenario, and a Verifier compares executions with and without the skill. The Verifier assesses executability, utility, and efficiency to assign a Keep or Reject verdict, determining whether the skill enters the library. Across ALFWorld and WebShop with three models, SkillSandbox consistently yields the strongest downstream performance and improved execution efficiency. Further analyses examine whether these gains reflect accurate assessment of skill reusability and identify which components of SkillSandbox contribute to them.

30. 【2610.10058】Cache the Encoder Within:Compact, Reusable Memory across LLM Queries

链接:https://arxiv.org/abs/2610.10058

作者:Hanzuo Liu,Chunyu Liu,Chaofan Lin,Alex Lamb,Mingyu Gao

类目:Computation and Language (cs.CL)

关键词:incur redundant encoding, caching model states, model states introduces, documents incur redundant, Repeated queries

备注: 17 pages, 3 figures, 7 tables

点击查看摘要

Abstract:Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

31. 【2610.10049】he Long Road to the Same Answer: Cognitive Bias Under Escalating Reasoning Budgets in Large Language Models

链接:https://arxiv.org/abs/2610.10049

作者:Obada Kraishan

类目:Computation and Language (cs.CL)

关键词:allocate extra computation, models allocate extra, Reasoning models allocate, deliberate thought, allocate extra

备注: Accepted at the 2026 IEEE 8th International Conference on Cognitive Machine Intelligence (IEEE CogMI 2026). 8 pages, 3 figures, 5 tables. Code and data: [this https URL](https://github.com/obadaKraishan/anchored-minds)

点击查看摘要

Abstract:Reasoning models allocate extra computation at inference time and present their answers as the product of deliberate thought. If this deliberation works the way dual-process accounts of human cognition suggest, longer thinking should weaken the classic decision biases that fast, intuitive judgment produces. Using 30 vignettes covering six biases (anchoring, framing, loss aversion, escalation of commitment, availability, confirmation) from an established benchmark, we run a dose-response study across four model families, pairing each reasoning model with a matched non-reasoning sibling and requesting thinking ceilings of 0, 1,024, 4,096, and 8,192 tokens, for 12,350 API calls. Because a requested ceiling is not the same as realized deliberation, we use the reasoning tokens each call consumed as the dose. First, reasoning models are not less biased than their siblings; the point estimate leans the other way in every family, but the item-level pooled contrast is not reliable (Delta = +0.031, t(29) = 1.45, p = .157). Second, bias magnitude does not reliably fall as realized deliberation grows: no slope is significantly negative, and where anything moves it is the signed score drifting further from the human direction. Third, anchoring is the only bias in the human direction (d = 1.89). Four of the other five lean the opposite way in all seven models; with five items per bias, that reversal is reliable for framing and directional for escalation of commitment, confirmation, and loss aversion, while availability is absent. A one-line instruction to restate the anchor before answering lowered anchoring on all five anchoring items, which no amount of additional thinking did, although the effect does not reach significance (p = .057). The results argue against treating test-time reasoning as a rationality guarantee and for auditing deployed models bias by bias.

32. 【2610.09981】Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation

链接:https://arxiv.org/abs/2610.09981

作者:Teng-Ruei Chen

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:LLM routers pick, post hoc, pick a cheap, cheap or expensive, cloud platforms

备注: 20 pages, 5 figures, 9 tables

点击查看摘要

Abstract:LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real requests (WildChat-1M, LMSYS-Chat-1M) with two cost/quality routers and a domain router, survey eleven systems' logging, and test post-processing defenses. At matched length, the shift's direction depends on category and router. For RouteLLM at the 50% operating point, harassment and self-harm requests reach the strong model 19 points less often than comparable ones on prompts unseen in exploration, medical requests (exploratory: LLM labels failed their gate) 31 points less often on distinct prompts (both post hoc), and sexual requests 10 points more often (secondary); the other router's four are negative. Twenty RouteLLM decisions separate frequent medical askers with AUC 0.71, exploratory and below the pre-registered primary endpoint's 0.75 (domain router: 0.92, an upper estimate). Per-category length-matched parity with accurate labels removes the gap on real traffic, costing at most 0.2 accuracy points on RouterBench (post hoc), where routers' gaps on sensitive subjects (13-42 points, pre-registered) exceed those of an oracle routing by realized accuracy gain (1-11, post hoc). Per-conversation stickiness, per-user budget bands, and pooled parity fail, the last as categories' shifts differ in size or sign. A post hoc exact per-user rate hides only even-prefix strong counts and forfeits most self-assessed routing value; it preserves odd-position decisions, from which a post hoc log attack reaches AUC 0.73 after 20 RouteLLM requests (exploratory).

33. 【2610.09976】EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectors

链接:https://arxiv.org/abs/2610.09976

作者:Jicheng Zhou,Kahim Wong,Jialong Wang,Jiantao Zhou

类目:Computation and Language (cs.CL)

关键词:source large language, large language model, decoding choices, large language, detection performance

备注:

点击查看摘要

Abstract:AI-generated text (AIGT) detection can be sensitive to the decoding choices of the source large language model (LLM). We observe that perturbing next-token logits or adjusting sampling temperature can reduce detection performance, providing a clear signal of detector vulnerability to decoding-time distribution changes. Building on this observation, we propose EASE (Entropy-Adaptive Distribution Shaping for Evasion), a training-free and detector-agnostic framework for evading AIGT detectors. EASE computes predictive entropy directly from the source LLM's next-token distribution and uses it to adapt both logit perturbation and sampling temperature, without detector feedback or model fine-tuning. Experiments across three source LLMs and multiple detectors demonstrate consistent reductions in detection performance, with negligible degradation in text quality and negligible inference overhead.

34. 【2610.09934】Itgan at NADI 2026 shared task: Parameter-Efficient Whisper Adaptation for Robust, Mixed-Dialect and Code-Switched Arabic ASR

链接:https://arxiv.org/abs/2610.09934

作者:Ibrahim Almajai

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)

关键词:Tunisian code-switched ASR, robust country-level ASR, mixed-dialect ASR, ASR subtasks, country-level ASR

备注: 12 pages, Arabic NLP 2026 Shared Task

点击查看摘要

Abstract:We describe the Itgan systems for the three ASR subtasks of NADI 2026, namely robust country-level ASR (1.1), mixed-dialect ASR (1.2), and Tunisian code-switched ASR (1.3). All three share one recipe, Whisper adapted with LoRA on consumer GPUs, and each was carried by a different addition to it. On 1.1, where the dialect label is given at test time, per-dialect specialists continued from a pooled adapter gave the largest gain, and the submitted system reached 57.1% country-average WER. A post-evaluation linear probe on frozen encoder features routes utterances without the label and recovers 44% of what oracle routing gives. On 1.2 the choice of base model mattered more than adapter capacity, and system combination helped only once we added a decorrelated member, reaching 46.7% WER. On 1.3 our system placed second at 14.49% WER with the lowest CER among the leading submissions, 5.38%. Its last 0.60 WER points came without further training, mostly from an exact weight-space average of independently trained runs, with ROVER voting adding the remainder. Every comparison carries a paired-bootstrap test, and we report eight directions that did not work.

35. 【2610.09920】Inverting Multi-Vector Visual Document Indices

链接:https://arxiv.org/abs/2610.09920

作者:Zhuchenyang Liu,Yao Zhang,Yu Xiao

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Prevailing multi-vector visual, vector databases run, thousand patch vectors, Prevailing multi-vector, page

备注: 30 pages. Under review

点击查看摘要

Abstract:Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.

36. 【2610.09906】Constrained-Action AI Remediation for SIEM/XDR via a NeMo-Guardrails Proxy

链接:https://arxiv.org/abs/2610.09906

作者:Georgios Koutidis,Nikolaos Kekatos,Tom Nianios,Alexios Lekidis

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Security Operations Centers, Security Operations, Operations Centers, operational technology share, Large Language Models

备注: 7 pages, 3 figures, 4 tables. Accepted at the 2026 IEEE International Conference on Cyber Security and Resilience (IEEE CSR 2026)

点击查看摘要

Abstract:Security Operations Centers (SOCs) for information technology and operational technology share one incident-response problem: a flood of correlated alerts and too few analysts. Large Language Models (LLMs) are increasingly proposed as reasoning engines that triage alerts and, in autonomous deployments, issue commands that block IPs, kill processes, or quarantine files on production hosts. This coupling introduces a new risk: a single adversarial alert can become a remote code path through the LLM's reasoning, leading it to recommend an action the SOC then executes. We present a constrained-action architecture with two coordinated layers: (i) a SIEM/XDR control plane that grounds remediation in correlated host events and confines the LLM's output to a closed intent vocabulary whose templated commands are executed by thin endpoint agents, backstopped by an argument validator; and (ii) a NeMo-Guardrails proxy that wraps the SOC-analyst LLM with input- and output-rail policies, evaluated out-of-the-box against a SOC-specific adversarial corpus we release. The stock proxy lifts injection recall from 25.0% to 94.5% at a 0.1% false-positive rate, and a live red-team exercise confirms that the closed intent vocabulary and argument validator contain the observed LLM failure modes before any command crosses the trust boundary. As an architectural fit (not yet a measured operational-technology deployment), the constrained-action property suits critical-infrastructure settings where a wrong remediation has physical, not merely operational, consequences. The loop is best run human-in-the-loop or delayed: the measured rail latency keeps inline control out of scope.

37. 【2610.09872】LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets

链接:https://arxiv.org/abs/2610.09872

作者:Jun Zhao,Leiming Fu,Yanbo Wen,Yiding Wang,Xuantong Liu,Yang Shu,Yuyang Lu,Xuanran Xing,Jingqi Tong,Hao Xu,Qi Zhang,Xuanjing Huang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Evaluating agents, Evaluating, persistent LLM agents, Abstract, outcomes

备注:

点击查看摘要

Abstract:Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-capability gap: realized returns often diverge from capability-specific measurements, and similar outcomes can arise from markedly different patterns of mechanism use. Trace-level diagnostics further expose distinct bottlenecks across capabilities, demonstrating that mechanism access, effective mechanism use, and downstream performance are not interchangeable measures of agent capability. LiveMACEBench makes this distinction measurable, turning live markets from a performance leaderboard into a diagnostic environment for agent capability

38. 【2610.09858】raining Advisors for LLM Agents from Task Outcomes

链接:https://arxiv.org/abs/2610.09858

作者:Sergei Polezhaev,Barys Liskavets,Ori Press,Alexander Golubev

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language model, Large language, tackle multi-step tasks, agents tackle multi-step, tackle multi-step

备注: 26 pages, 11 figures

点击查看摘要

Abstract:Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including $\tau^3$ and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.

39. 【2610.09835】A Deafening Silence: Catastrophic Forgetting Lives in the Output Embeddings of Tokens the Data Never Speaks

链接:https://arxiv.org/abs/2610.09835

作者:Jonghyun Han,Younghoon Song,Jongyoul Park

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large Language Models, Large Language, often-inaccessible original data, Continual pre-training, fine-tuning in Large

备注:

点击查看摘要

Abstract:Continual pre-training and fine-tuning in Large Language Models (LLMs) inevitably induce catastrophic forgetting, typically mitigated by replay using often-inaccessible original data. In this data-free regime, we analyze where forgetting occurs and why. Systematic parameter freezing across five settings up to 1.4B reveals that forgetting concentrates selectively in the output embeddings of tokens rarely seen in the new corpus, whereas the same sqrt(v-hat) band of the body is inert and new learning resides elsewhere. This localization is governed by the vocabulary deficiency of the corpus rather than the training mode, allowing pre-retraining risk ranking from token counts alone within a fixed base model. Mechanistically, absent tokens receive persistent one-sided softmax gradients that Adam's second-moment (sqrt(v-hat)) normalization amplifies into full-sized updates. We therefore propose an intervention: raising Adam's epsilon exclusively for the output projection during training. Across eight settings spanning 160M to 12B parameters and four model families, this removes 39.4% to 67.9% of forgetting across all seven stable configurations without degrading target learning or requiring per-model tuning. The defense combines additively or better with replay (79.8% on Qwen/Korean) and rescues released-head LoRA from a 23-fold forgetting surge. Because post-hoc editing of the drifted rows recovers under 5% of forgetting, the intervention must operate during training. Our findings indicate that a single-line optimizer adjustment may serve as the primary defense against catastrophic forgetting where the corpus starves the vocabulary.

40. 【2610.09795】MIRROR: From Imitation to Internalization in LLM Personalization

链接:https://arxiv.org/abs/2610.09795

作者:Huayi Lai,Jicheng Yang,Min Yi,Chong Meng

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:shifting from style, Reference-Revealed On-policy Reflections, style imitation, LLM personalization, Abstract

备注: 36 pages

点击查看摘要

Abstract:The demand for personalized LLMs is shifting from style imitation toward content quality. We investigate whether self-distillation can bridge this gap in existing fine-tuning paradigm. To address this limitation, we introduce MIRROR(Meta- personalization by Internalizing Reference-Revealed On-policy Reflections), a novel self-distillation framework that shifts LLM personalization from imitation toward preference internalization. First, we replace reference-token imitation with reference-revealed on-policy self-distillation, aligning the model's next-token distributions along its own generation trajectories with those of its reference-conditioned self, thereby internalizing user preferences rather than reproducing reference this http URL, we introduce MIRROR-F, a focal plug-in that augments on-policy distributional alignment with selective supervision over informative reference tokens, thereby strengthening content generation while preserving user-specific expression. Across three personalized generation benchmarks, two model scales, and complementary reference-based and LLM-based evaluations, MIRROR and MIRROR-F achieve leading overall personalization performance and superior text quality, while exhibiting less catastrophic forgetting than SFT-based baselines on three unseen personalized generation tasks. The gains are consistent across model scales and application scenarios, translating to improved performance in LLM personalization tasks.

41. 【2610.09788】Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

链接:https://arxiv.org/abs/2610.09788

作者:Mingqing Yuan(Soochow University),Xiaobo Liang(Soochow University),Junwei Yang(University of Cambridge),Ziwei Chen(Chalmers University of Technology),Zeren Zhang(Peking University),Hejin Wang(Tsinghua University),Yubin Wang(The Hong Kong University of Science and Technology),Juntao Li(Soochow University)

类目:Computation and Language (cs.CL)

关键词:requires jointly representing, Reward modeling, multiple evaluation criteria, incur substantial inference, process token

备注:

点击查看摘要

Abstract:Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.

42. 【2610.09772】Decoupling Logic from Persona: Structural Immunity of Edge LLM Agents to Context Pollution

链接:https://arxiv.org/abs/2610.09772

作者:Masaaki Nakatsu(AO, Inc. / OrbLabs AG),Reno Wang(AO, Inc.)

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Small language-model agents, Small language-model, decoupled logic path, logic path, inside one context

备注: 28 pages, 3 figures. Experiment code, scoring rubric, pollution fixtures, adapters and run logs are released (see Appendix G)

点击查看摘要

Abstract:Small language-model agents on edge devices must hold a persona and reason correctly at once, inside one context window that fills with conversational history and persona instructions. We study what happens to the logical part of such an agent when that history is long, misleading and persona-heavy (persona-logic interference), and present a Decoupling Architecture (AO-DA) that separates logical inference ("What") from persona expression ("How") into two inference paths on one INT4 base model with hot-swappable LoRA adapters. The logic path receives only the core turn and emits a verifiable structured state (Micro-State); the persona path renders it in character with the full history. In same-base-model ablations on an Apple M2 laptop (Llama-3.1-8B-Instruct and Gemma-3-4B-it, 4-bit; 480 runs over 4 pollution levels x 3 arms x 2 tasks x 2 personas x 5 seeds) we find: (i) the decoupled logic path is structurally invariant to pollution: its prompt stays at 180 (Llama) or 167 (Gemma) tokens while the mixed single-pass prompt grows from 242 to 1,203, and its outputs are byte-identical across levels (40/40); (ii) the mixed single pass degrades monotonically (composite logic score 0.669 to 0.150 on Llama, 0.487 to 0.150 on Gemma), mostly by failing to emit the required structured output (80-95% of runs on Llama, 100% on Gemma at the two highest levels); (iii) with the same pollution fed into the decoupled logic path, the dedicated-adapter, dedicated-format path is still more robust than the single pass on the 8B model (failure 0-20% vs 80-95%; paired $\Delta$ +0.30 to +0.50, Cliff's $\delta$ 0.50-0.85, Holm-adjusted $p \le 0.03$) but not on the 4B model, where both collapse. Separation costs one extra decode on a topic's first turn (28.2 s vs 18.2 s on Llama) and buys persona hot-swapping in 1.7 ms without re-running the logic path. Code, rubric, fixtures, adapters and logs are released.

43. 【2610.09769】From Expert-Guided Proof Search to Automated Open-Problem Solving

链接:https://arxiv.org/abs/2610.09769

作者:Adrián Zámečník,Matěj Kripner,Martin Koutecký,Martin Balko,Jan Grebík,Pavel Hubáček,Robert Šámal,Václav Rozhoň

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language models, Large language, efficient proof search, incremental improvements, careful verification

备注: Accepted at the 6th Workshop on Mathematical Reasoning and AI (MATH-AI), NeurIPS 2026

点击查看摘要

Abstract:Large language models are increasingly contributing to mathematical research, where progress often depends on efficient proof search, incremental improvements and careful verification. We describe Bolzano, a multi-agent open-source system that uses parallel prover agents with a verifier agent and maintains a human-readable research state. Initial manual use on expert-selected problems yielded 8 results whose proofs were checked by domain experts. Motivated by these case studies, we ran Bolzano without problem-specific human guidance on about 3,800 open problems extracted from four sets of papers, solving about 200 open problems. One experiment used papers accepted to STOC 2026, a top conference in theoretical computer science. There, we answered four questions raised in the papers, as confirmed by their authors.

44. 【2610.09761】PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency

链接:https://arxiv.org/abs/2610.09761

作者:Shengkai Ma,Zhenyu Hou,Weihua Cao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:estimates a position, surrounding objects, localization estimates, submap, PARC

备注:

点击查看摘要

Abstract:Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding similarity between the query and multiple submaps, even when the instance layout within a submap violates the query description. Meanwhile, query-relevant instances often span submap boundaries, leaving the retrieved submap with incomplete contextual evidence. We term these failure modes layout-inconsistent aliasing and boundary evidence incompleteness, respectively. To address them, we propose PARC-Loc, a coarse-to-fine localization framework built on Partial Assignment with Relational Consistency (PARC). PARC jointly models hint-object compatibility and pairwise spatial relations, allowing unmatched elements while favoring assignments consistent with the queried layout. At the coarse stage, its candidate-level assessment complements neural similarity for layout-consistent submap selection. At the fine stage, the context is expanded with query-relevant instances from adjacent submaps, while PARC yields object-level matching weights that guide cross-modal attention. Extensive experiments on KITTI360Pose and CityLoc show that PARC-Loc outperforms conventional coarse-to-fine baselines. On KITTI360Pose, our method improves Top-1 localization recall at 5 m from 0.50 to 0.67, achieving a 34% relative gain over the strongest baseline.

45. 【2610.09756】Shaer: Controlled Arabic Poetry Generation with Meter Subform and Semantic Conditioning

链接:https://arxiv.org/abs/2610.09756

作者:Ahmad Abbas,Tamara Fakih,Nour Fakih,Ammar Mohanna

类目:Computation and Language (cs.CL)

关键词:fine-grained prosodic constraints, Classical Arabic poetry, requires simultaneously satisfying, Arabic poetry generation, Classical Arabic

备注: 22 pages, 7 figures. Code: [this https URL](https://github.com/AhmaddAbbass/Shaer) ; models and datasets: [this https URL](https://huggingface.co/Shaer-AI)

点击查看摘要

Abstract:Classical Arabic poetry generation requires simultaneously satisfying semantic, linguistic, and fine-grained prosodic constraints. Existing systems typically control broad poetic attributes but do not jointly model semantic intent, meter subform, and poem length. We present Shaer, a controllable Classical Arabic poetry generation framework jointly conditioned on natural-language descriptions, meter subforms, and target hemistich counts. To support this task, we construct an enriched corpus of 116,032 classical Arabic poems derived from Ashaar, containing normalized meter-subform labels and automatically generated, validated semantic descriptions. We then adapt Yehia-7B using QLoRA-based supervised fine-tuning with a completion-only objective. Our evaluation combines automatic assessment of base-meter conformity, requested-subform adherence, and length control with three LLM judges, blinded human evaluation, and memorization analysis. Shaer achieves 95.17% base-meter accuracy, 91.75% poem-level meter-subform accuracy, and 83.40% exact count accuracy. Relative to its untuned foundation model, these results represent gains of 68.68, 57.77, and 38.93 percentage points, respectively; Shaer also attains the highest base-meter accuracy among all evaluated systems. Multi-LLM evaluation and a blinded human assessment of top-ranked outputs further indicate competitive semantic and literary quality. Finally, analysis of all 3,481 test generations finds no exact copies from the training corpus or paired source poems. Code, models, and datasets are publicly available.

46. 【2610.09733】Bridge Routing Heads: Where Multilingual Multi-hop Reasoning Lives in LLMs

链接:https://arxiv.org/abs/2610.09733

作者:Seunghan Kim,Minyeong Choe,Hyunil Kim,Haehyun Cho

类目:Computation and Language (cs.CL)

关键词:Multilingual LLMs, multi-hop reasoning question, identify Bridge Routing, Bridge Routing Heads, Multilingual LLMs answer

备注: Accepted at EMNLP 2026

点击查看摘要

Abstract:Multilingual LLMs answer the same multi-hop reasoning question across languages, but we lack a mechanistic account of whether they share an internal circuit. We identify Bridge Routing Heads (BRH) in two large multilingual LLMs through a three-stage pipeline. The resulting language-specific head sets exhibit near-complete mutual exclusivity across the five languages, with a mean Jaccard similarity of only 0.017 for Llama 3.1 70B and 0.057 for Qwen 2.5 72B, revealing language-idiosyncratic circuits. Ablating general BRH increases two-hop Negative Log-Likelihood (NLL) by 39-89x the random-head baseline, providing direct causal evidence of their role. Amplifying these heads in a failing target-language pass rescues up to 51.7% of cross-lingual failures, with no training. The two models share this dual-circuit pattern but allocate heads differently: Llama concentrates chaining in a large general pool, while Qwen leans on larger language-specific pools. Together these results show that activation-level intervention alone can recover correct answers from cross-lingual reasoning failures.

47. 【2610.09724】owards Explaining Query Expansion Performance in Information Retrieval

链接:https://arxiv.org/abs/2610.09724

作者:Sourav Saha,Aditya Dutta,Soumajit Pramanik,Mandar Mitra

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:vocabulary mismatch problem, Query Expansion, Information Retrieval, techniques have long, mismatch problem

备注:

点击查看摘要

Abstract:Query Expansion (QE) techniques have long been widely used in Information Retrieval (IR) to address the vocabulary mismatch problem. They remain relevant in modern retrieval systems, including those based on large language models (LLMs). However, no single QE method consistently outperforms others across all queries. This work seeks to explain the variation in QE performance through two complementary perspectives. The first is the concept of an Ideal Expanded Query (IEQ)--a hypothetical query that maximizes retrieval effectiveness with a downstream BM25 retrieval model. The second is a separability perspective, which quantifies how distinctly relevant and non-relevant documents are scored for a given expanded query using Cohen's (d). We develop a separability measure and practical formulations to approximate the IEQ and investigate how these factors relate to retrieval effectiveness. Extensive experiments on the TREC Robust collection, TREC DL 2019-2022 passage collections, and TREC DL 2019-2020 document collections reveal several interesting patterns. In particular, we find that expanded queries that are closer to the ideal expanded query tend to achieve higher retrieval effectiveness. We further show that the separability of relevant and non-relevant documents provides a complementary perspective for understanding QE performance.

48. 【2610.09710】SpikingVLA: Asynchronous Spiking Vision-Language-Action Models

链接:https://arxiv.org/abs/2610.09710

作者:Jingya Wang,Dehao Zhang,Shuai Wang,Malu Zhang,Yang Yang,Haizhou Li

类目:Computation and Language (cs.CL)

关键词:training large-scale SNNs, SNNs from scratch, VLA, route toward energy-efficient, cost of training

备注:

点击查看摘要

Abstract:ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on DIF neurons, we further introduce an asynchronous execution mechanism that overlaps temporal computation across VLA components, reducing synchronization overhead and latency. Extensive experiments demonstrate that SpikingVLA achieves competitive navigation performance with substantially improved inference efficiency. Compared with existing spiking VLA methods, SpikingVLA improves SR and SPL by 11.9\% and 12.6\%, respectively, while reducing first-action latency by 11.2$\times$. These results establish SpikingVLA as a practical framework for deploying pretrained VLA models with high-performance and low-latency spiking inference.

49. 【2610.09684】From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

链接:https://arxiv.org/abs/2610.09684

作者:Xinglin Wang,Zishen Liu,Tong Zheng,Shaoxiong Feng,Peiwen Yuan,Yiwei Li,Jiayi Shi,Yueqi Zhang,Chuyi Tan,Ji Zhang,Boyuan Pan,Kan Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:additional inference computation, large language models, allocating additional inference, inference computation, Personalized Test-Time Scaling

备注: Preprint

点击查看摘要

Abstract:Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controllers. We formulate Personalized Test-Time Scaling as discovering executable controllers that maximize the joint satisfaction rate of user-specific requirements. To reduce the overhead of repeated policy discovery for new user profiles, we propose PersonTTS, an amortized agentic policy-discovery framework that reuses prior search experience through requirement-matched controller initialization and source-distilled procedural guidance, while retaining target-profile evaluation for every candidate. Experiments on AIME and HMMT show that PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems. Under the same candidate-evaluation budget, cross-user experience reuse further improves policy quality while substantially reducing discovery-agent time and cost.

50. 【2610.09671】InsClaimBench: Benchmarking Insurance Claim Adjudication Across the Decision Chain

链接:https://arxiv.org/abs/2610.09671

作者:Linqi Zhang,Chong Qi,Yan Cheng,Wanqing Cao,Yu Liu,Chenwei Lin,Xian Xu

类目:Computation and Language (cs.CL)

关键词:reasoning-oriented large language, Recent advances, large language models, perform professional decision, motivated increasing evaluation

备注: 17 pages, 3 figures, 11 tables

点击查看摘要

Abstract:Recent advances in reasoning-oriented large language models (LLMs) have motivated increasing evaluation of their ability to perform professional decision tasks. Insurance claim adjudication is one such task, requiring models to connect case evidence, insurance rules, intermediate judgments, and payout calculations across a structured decision process. We introduce InsClaimBench, an end-to-end benchmark for evaluating insurance claim adjudication across the decision chain. Grounded in real claim materials and structured insurance rules, InsClaimBench contains 3,780 cases in 375 case families across auto, property, and health insurance, comprising 86,656 atomic rule judgments. It evaluates each claim from atomic rules through adjudication modules to payout decisions and amounts, with controlled factual variants testing whether required changes are correctly propagated across levels. Evaluation of six LLMs reveals a progressive loss of reliability along the decision chain. Payout-decision accuracy ranges from 74.23--80.19%, while joint decision--amount accuracy drops to 47.54--73.15%. Strong local performance also fails to ensure case-level correctness: atomic-rule accuracy reaches 95.48%, whereas rule-vector exact match peaks at only 36.90%, and the most frequent module errors are not necessarily those most associated with final-decision failure. Under factual changes, these inconsistencies further become propagation failures: module updates are less reliable than rule updates, correct local judgments can still yield incorrect payouts, and correct payouts can conceal intermediate errors. These results show that reliable claim adjudication requires consistent composition and propagation across the decision chain.

51. 【2610.09665】SAPD: Step-Aligned Privileged Distillation

链接:https://arxiv.org/abs/2610.09665

作者:Tianle Wang,Jiayu Liu,Ruizhi Zhao,Ning Miao

类目:Computation and Language (cs.CL)

关键词:costly rollout generation, improve large language, large language models, requires costly rollout, rollout generation

备注: preprint

点击查看摘要

Abstract:On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fixed demonstrations can support competitive off-policy learning through better supervision. Our premise is that their usefulness depends not only on the training trajectories, but also on whether supervision provides informative preferences among continuations and connects this guidance to the reasoning decision being learned. We introduce Step-Aligned Privileged Distillation (SAPD), a rollout-free self-distillation method that turns demonstrations into step-aligned distributional supervision. Its key insight is to use the known progression of a reference solution to associate each reasoning transition with targeted privileged guidance, rather than treating the solution as undifferentiated context. On mathematical reasoning benchmarks, SAPD outperforms supervised fine-tuning and label smoothing on average while remaining competitive with on-policy reinforcement learning and self-distillation. Analyses support both the value of context-dependent distributional guidance and the benefit of aligning privileged information with the current step. SAPD also largely preserves out-of-domain coding performance and achieves approximately 2x training-loop speedups over the on-policy baselines. These findings suggest that carefully constructed supervision can make fully off-policy post-training a competitive and computationally efficient alternative. Our code is available at this https URL.

52. 【2610.09661】Alice: A Large-Scale German Benchmark for Rubric-Based Multi-Dimensional Automatic Short Answer Scoring

链接:https://arxiv.org/abs/2610.09661

作者:Zhifan Sun,Sebastian Gombert,Jannik Lossjew,Tobias Wyrwich,Berrit Katharina Czinczel,David Bednorz,Marcus Kubsch,Knut Neumann,Hendrik Drachsler

类目:Computation and Language (cs.CL)

关键词:Automatic Short Answer, Short Answer Scoring, Automatic Short, NLP for Education, Answer Scoring

备注: EMNLP2026 Main

点击查看摘要

Abstract:Automatic Short Answer Scoring (ASAS) is central to NLP for Education. However, openly available benchmarks remain scarce, and existing datasets largely address how well students answer a question directly rather than how well they master underlying concepts (knowledge elements) such as thermal energy or epistemic activities (skills) such as reasoning or claim. To address this gap, we introduce Alice, a large-scale, rubric-based German ASAS dataset that is pedagogically aligned and comprises three subtasks: (i) learning performance (Alice-LP), (ii) knowledge elements (Alice-KE), and (iii) skills (Alice-SK). We further formulate rubric-based ASAS as a rubric-retrieval task and benchmark the dataset with a range of language models, from encoder-only models to lightweight LLMs. We also benchmark the dataset with zero-shot prompting via LLMs and a standard classification baseline. The experiments show that LLMs, in particular, struggle to score knowledge elements and skills in the zero-shot setting. They also indicate that rubric text is often useful, especially for Alice-KE and Alice-SK, while on Alice-LP gains over sample-solution-focused inputs are more modest and vary by model and input format.

Comments:
EMNLP2026 Main

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2610.09661 [cs.CL]

(or
arXiv:2610.09661v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.09661

Focus to learn more

              arXiv-issued DOI via DataCite</p>
53. 【2610.09660】Rubric Spans are Label Representations: Joint LLM Encoding for Short Answer Scoring

链接:https://arxiv.org/abs/2610.09660

作者:Zhifan Sun,Sebastian Gombert,Fabian Zehner,Leon Camus,Longwei Cong,Hendrik Drachsler

类目:Computation and Language (cs.CL)

关键词:Automatic Short Answer, Automatic Short, Short Answer Scoring, Short Answer, score student responses

备注: EMNLP2026 Main

点击查看摘要

Abstract:Automatic Short Answer Scoring (ASAS) requires models that can score student responses against question-specific criteria while remaining efficient and transferable across rubric sets. We propose RUSPAN, a rubric-conditioned ASAS framework that treats rubric descriptions as semantic label representations. RUSPAN serialises the question context, student answer, and all candidate rubric levels into a single sequence, then scores the levels listwise from the rubric-span and whole-sequence representations produced in a single LM pass. We further introduce RUSPAN-RIM, in which a Rubric-Independent Mask prevents rubric spans from attending to one another, making rubric representations depend only on the answer and question context and preventing overfitting to rubric patterns during training for zero-shot transfer. On six ASAS benchmarks spanning English, German, and Portuguese, RUSPAN improves mono-benchmark scoring over discriminative and generative baselines, while RIM with position reindexing delivers consistent and substantial gains on PT-ASAG, the held-out benchmark with the strongest combined language and rubric-structure shift.

54. 【2610.09647】When Rank Rises as LLMs Degrade

链接:https://arxiv.org/abs/2610.09647

作者:Zhaohui Geoffrey Wang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Post-training adapts language, adapts language models, non-stationary environments, adapts language, Post-training adapts

备注: NeurIPS 2026 Workshop on Continual Learning for Foundation Models and Agents (CL4FMAgents); 8 pages + appendix

点击查看摘要

Abstract:Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations. Covariance effective rank rises to nearly twice its healthy value. This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest. By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds. Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone. We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues. On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range. We then test a two-sided, multichannel sequential monitor with separate calibration and test data. In a pre-registered shared-prefix, leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction. However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed. Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.

55. 【2610.09639】On-Policy Distillation Teaches New Skills but Not New Knowledge

链接:https://arxiv.org/abs/2610.09639

作者:Yixuan Tang,Yi Yang

类目:Computation and Language (cs.CL)

关键词:strengthens language-model reasoning, reasoning remains unknown, strengthens language-model, remains unknown, compositional skill

备注:

点击查看摘要

Abstract:On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.

56. 【2610.09633】Coding-Agent Benchmarks Should Match Their Users' Task Flows

链接:https://arxiv.org/abs/2610.09633

作者:Igor Slinko,Yaroslav Golubev,Sergey Titov

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)

关键词:coding agents generally, agents generally strives, generally strives, call Production Sessions, coding agents

备注:

点击查看摘要

Abstract:The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project's code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it. We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories. In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.

57. 【2610.09607】Which Language Should a Skeleton Speak? Language Choices in Multilingual Reasoning

链接:https://arxiv.org/abs/2610.09607

作者:HyeonSeok Lim,SeungWoo Song,Inho Won,Hoyun Song,Jihyo Kim,KyungTae Lim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Skeleton-based reasoning prompting, structuring LLM reasoning, promising training-free approach, prior work largely, work largely assumes

备注: Accepted to EMNLP 2026 (Findings)

点击查看摘要

Abstract:Skeleton-based reasoning prompting is a promising training-free approach for structuring LLM reasoning, but prior work largely assumes an English-centric setting. We propose the Language-Aware Skeleton Exploration Framework (LASEF) to study skeleton-language choice in multilingual mathematical reasoning. Across math benchmarks, model scales, and languages, we show that English skeletons yield a small positive tendency on average, most visible for smaller models and low-resource languages. However, few language-level gains remain significant after correction, and English is not universally optimal. Combining greedy decoding, multi-rollout evaluation, translation ablation, and cross-benchmark validation, we further find three patterns of skeleton-language effects: directionally consistent, evaluation- and benchmark-dependent, and asymmetric negative. These effects cannot be fully explained by generation quality alone. Overall, skeleton language is a context-dependent design variable that requires multi-level exploration. All resources are released at this https URL.

58. 【2610.09587】Collaborative Reasoning Distillation via Cross-Feedback and Coherent Curation

链接:https://arxiv.org/abs/2610.09587

作者:Taehoon Kim,Seunggeun Cho,Dongsu Han

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:advancing Large Language, Large Language Models, Large Language, require massive computational, effectively distill reasoning

备注: Accepted at NeurIPS 2026

点击查看摘要

Abstract:Reasoning capabilities are critical for advancing Large Language Models, yet current approaches either require massive computational budgets or struggle to effectively distill reasoning to smaller models. Standard distillation methods rely on outcome-based rewards, failing to distinguish between sound reasoning and lucky guesses. We propose Collaborative Reasoning Distillation (CRD), a framework that enhances reasoning in compact models through three innovations: (1) interactive cross-feedback where teachers iteratively critique each other's reasoning, (2) fine-grained step-wise quality assessment capturing logical validity independent of final answers, and (3) coherence-aware step stitching that synthesizes complementary strengths. Students are trained via Reasoning Quality Optimization (RQO) with budget constraints. Our model, CRD-4B, achieves 97.3% on MATH-500 and 70.3% on AIME'25, surpassing baselines while using only 50K training examples, up to 12 times smaller than the datasets of comparable models.

59. 【2610.09571】How Do LLMs Change Predictions Under Negation?

链接:https://arxiv.org/abs/2610.09571

作者:Jongwook Yoon,Jongwon Lim,Sungjib Lim,Woojin Cho,Yohan Jo

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:large language models, large language, remain unreliable, essential feature, Negation

备注: Under Review

点击查看摘要

Abstract:Negation is an essential feature of human language, yet large language models (LLMs) remain unreliable in processing it. We evaluate recent open-source and closed-source LLMs on our negation benchmark and find that, in 37-71% of cases, they repeat the same answer under negation (e.g., "Madrid" for "What is not the capital of Spain?"). To understand and address this brittleness, we mechanistically examine how models operate under negation. Our main finding is that specialized attention heads and MLP neurons jointly implement negation by (1) suppressing retrieval of the original answer (e.g., "Madrid") while (2) promoting a favored candidate within the answer category (e.g., "Paris"). This contrasts with accounts of human negation processing, in which information about the original answer helps to determine what should be excluded. Furthermore, we find that this difference from human processing is a key source of negation failures: the model's mechanism relies on suppressing the original answer rather than using it to determine what to exclude, so the model can repeat the original answer when suppression is too weak or when a bias toward particular answers prevents it from selecting an alternative. To address this weakness in the model's negation mechanism, we propose a training objective that requires larger shifts in answer preference for more confident original predictions, and show that it reduces negation failures with less degradation of general capabilities than standard fine-tuning baselines. Together, our results demonstrate how mechanistic analysis can reveal why a linguistic capability fails and guide training that targets the underlying limitation.

60. 【2610.09569】RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models

链接:https://arxiv.org/abs/2610.09569

作者:Shivam Shukla,Jihye Kim,Shubham Gaur,Mahnaz Roshanaei,Magy Seif El-Nasr

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, raising concern, concern that sustained, emotional support

备注:

点击查看摘要

Abstract:Large language models (LLMs) are increasingly used for emotional support, raising concern that sustained use may draw users away from their real-world relationships. Yet existing evaluations primarily focus on the safety, empathy, or helpfulness of responses, leaving under-examined a relational question: where does the model orient the user for continued support? To address this question, we introduce relational orientation, a property operationalized through two non-exclusive dimensions: inward-facing (IF) language, which positions the AI as the user's ongoing source of support, and outward-scaffolding (OS) language, which encourages real-world human connection. Grounded in psychological and sociological literature, we formalize a taxonomy of relational orientation and present RELATE, a persona-conditioned framework for measuring inward-facing and outward-scaffolding language at the sentence level in multi-turn dialogues. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles, providing 228 evaluation stimuli. In our experiments, we evaluate seven LLMs using dialogues with six assistant turns each, yielding 1,596 dialogues and 69,194 assistant sentences. We assess these sentences using a primary rubric-based LLM judge and apply a secondary judge to a subset. Under automated evaluation, we find that the proportion of sentences labeled as IF is higher at the sixth assistant turn than at the first, while the proportion labeled as OS is substantially lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users. RELATE provides a reproducible framework and a sentence-level signal for auditing and steering the relational orientation of supportive LLMs.

61. 【2610.09552】Constitution-Guided Watermarking

链接:https://arxiv.org/abs/2610.09552

作者:Toluwani Aremu,Samuele Poppi,Nils Lukas

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:identify text generated, Watermarking enables language, enables language model, enables language, identify text

备注: Working paper (under review)

点击查看摘要

Abstract:Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations that balance competing objectives or prioritize particular properties. Either approach imposes a shared operating point on requests with different requirements, potentially sacrificing quality where wording preservation matters or robustness where reliable attribution is essential. To allow flexible and adaptable designs, we introduce \emph{Constitution-Guided Watermarking}, a framework that selects request-appropriate trade-offs from provider requirements, listed as natural-language principles. \emph{Offline}, a pretrained reasoning agent examines constitutional rules alongside watermark implementations and iteratively refines rule-specific configurations using empirical feedback. \emph{At deployment}, a separate monitor identifies applicable rules and retrieves the corresponding policy, including watermarking exemptions, without modifying the serving model. Furthermore, our framework supports offline parallel optimization and refinement of rule-specific configurations based on evolving provider requirements without affecting deployment, and binds each deployed configuration to its evaluation evidence, making deployment decisions auditable. In a proof-of-concept evaluation using KGW and a five-rule constitution, our framework selects configurations responsive to provider priorities and improves post-paraphrase detection on robustness-prioritized requests by up to $14$ percentage points over fixed configurations, while matching or exceeding all baselines in aggregate quality and clean detection at a nominal $0.1\%$ false-positive rate.

62. 【2610.09529】A Comparative Study of Evaluation Metrics for Long-Document Financial Narrative Summarization with Transformers

链接:https://arxiv.org/abs/2610.09529

作者:Nadhem Zmandar,Mo El-Haj,Paul Rayson

类目:Computation and Language (cs.CL)

关键词:London Stock Exchange, single financial year, London Stock, Stock Exchange, required to communicate

备注: 12 pages

点击查看摘要

Abstract:There are more than 2,000 listed companies on the UK's London Stock Exchange, divided into 11 sectors who are required to communicate their financial results at least twice in a single financial year. UK annual reports are very lengthy documents with around 80 pages on average. In this study, we aim to benchmark a variety of summarisation methods on a set of different pre-trained transformers with different extraction techniques. In addition, we considered multiple evaluation metrics in order to investigate their differing behaviour and applicability on a dataset from the Financial Narrative Summarisation (FNS 2020) shared task, which is composed of annual reports published by firms listed on the London Stock Exchange and their corresponding summaries. We hypothesise that some evaluation metrics do not reflect true summarisation ability and propose a novel BRUGEscore metric, as the harmonic mean of ROUGE-2 and BERTscore. Finally, we perform a statistical significance test on our results to verify whether they are statistically robust, alongside an adversarial analysis task with three different corruption methods.

63. 【2610.09489】Goldsmith: Gold-Loss-Guided Definition Optimization with an Agentic Annotation Harness

链接:https://arxiv.org/abs/2610.09489

作者:Yihan Li,Hanyi Zhang,Xiaoxi Jiang,Man Guo

类目:Computation and Language (cs.CL)

关键词:annotation projects begin, projects begin, stable guideline, labels to train, train a task-specific

备注: 20 pages, 4 figures, 11 tables. Accepted to the main conference of EMNLP 2026

点击查看摘要

Abstract:Many annotation projects begin before experts have a stable guideline or enough labels to train a task-specific model. We present Goldsmith, an agentic pipeline that turns a small gold set---expert-annotated calibration examples representing the intended task boundaries---into a reusable structured annotation definition. Goldsmith treats this definition as a trainable textual object. Candidate definitions are run on the same gold examples and scored with an executable structured loss, while the output schema, formatting, retrieval, repair, judging, and human review remain in an external harness. A large language model (LLM) editor converts the highest-loss failures into textual-gradient revisions, which are accepted only when the measured loss decreases. In prompt-optimization comparisons, Goldsmith improves over direct rewriting, OPRO, APE, and PromptBreeder under matched evaluation protocols. The resulting definition also improves downstream annotation when combined with retrieval, score-based routing, and human review across typed span, pair-level relation, and fixed-trigger event-argument tasks. These results show that scarce expert supervision can support both task-definition learning and scalable annotation.

64. 【2610.09476】CHASE: Channel-Aligned Structure Exploitation for Geometry-Aware Model Engineering

链接:https://arxiv.org/abs/2610.09476

作者:Wei Wang,Wei Jiang,Ziran Liu

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:characterizes trained networks, GSA, characterizes trained, singular bases, trained networks

备注: 26 pages, 9 tables

点击查看摘要

Abstract:Geometric and Spectral Alignment (GSA) characterizes trained networks through spectral concentration, physical-channel alignment, support structure, and changes in singular bases. In this paper, we propose CHASE (Channel-Aligned Structure Exploitation) to use these structures in practical model design. CHASE covers six applications across model modification, reconfiguration, and compression. CORA, COEC, and CORAM apply GSA to parameter-efficient finetuning, structured-pruning compensation, and model merging. We further develop three new methods. CAGA uses GSA to identify multi-head attention heads that can share a KV representation and constructs the shared key and value heads through geometric alignment and low-rank subspace extraction. SAKV uses GSA to determine which adjacent layers can share a low-rank KV-cache representation and the retained rank for each layer group. CAPS uses GSA spectral structure to group output neurons and selects retained input channels separately for each group. Results from CORA, COEC, and CORAM establish the effectiveness of GSA for adaptation, pruning compensation, and model merging. Experiments on CAGA show that geometric shared-head construction substantially improves MHA-to-GQA conversion, and SAKV and CAPS improve over representative baselines for KV-cache compression and structured pruning. These results show that the structures identified by GSA can be used directly to design methods for a range of model operations.

65. 【2610.09464】BanglaRhet: Benchmarking Classical and Transformer Models for Rhetorical and Persuasion Detection in Bangla Political Speech

链接:https://arxiv.org/abs/2610.09464

作者:Rohit Kumar Sen,Anik Chowdhury

类目:Computation and Language (cs.CL)

关键词:influence public opinion, Bangla political speech, persuasion technique detection, Bangla political, Bangla political discourse

备注: 6 pages, 2 figures, 6 tables. Accepted at the 2026 2nd International Conference on Advances in Computing, Communication, Electrical, and Smart Systems (iCACCESS), Dhaka, Bangladesh. Dataset: [this https URL](https://doi.org/10.5281/zenodo.23162113) . Code: [this https URL](https://github.com/TSRohit99/banglarhet) . Hugging Face: [this https URL](https://huggingface.co/datasets/tsrohit99/banglarhet)

点击查看摘要

Abstract:Political discourse often uses rhetorical and persuasive language to frame narratives, influence public opinion, and mobilize audiences. While Bangla natural language processing has made progress in sentiment analysis and opinion mining, systematic benchmarking of transformer models for fine-grained rhetorical and persuasion technique detection in Bangla political speech remains largely underexplored. This paper presents a benchmark study of transformer-based models for detecting rhetorical form and persuasive intent in Bangla political discourse. Using BanglaRhet, a manually annotated corpus of 30,289 Bangla political speech segments collected from publicly available political news sources, we formulate two supervised single-label classification tasks: rhetorical technique detection (contrast, repetition, exaggeration, metaphor, rhetorical questions) and persuasion technique detection (blame assignment, call to action, unity call, moral, emotional, and logical appeals). We evaluate four transformer-based models, BanglaBERT, BanglaBERT-Base, SahajBERT, and XLM-RoBERTa-Base, against classical TF-IDF baselines. BanglaBERT achieves the highest performance, with 65.40% macro-F1 for rhetorical technique detection and 66.46% for persuasion technique detection, outperforming the best tuned classical baseline by 19.2 and 13.8 macro-F1 points, respectively. Class-level analysis indicates that errors are mainly associated with semantic overlap among labels, figurative language, and class imbalance. The results provide initial benchmark baselines for Bangla rhetorical and persuasion-aware political discourse analysis and highlight the need for context-aware and multi-label modeling.

66. 【2610.09458】Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy

链接:https://arxiv.org/abs/2610.09458

作者:Jiayu Feng

类目:Computation and Language (cs.CL)

关键词:policy question wrongly, state-specific policy question, state-specific policy, returning a real, LLM answers

备注: 6 pages, 3 figures, 1 table

点击查看摘要

Abstract:When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities. Gold values come from an official data book and agree with an independent source in 101 of 102 checked cells. Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state's current value, identical across two independent repeats, for 10 and 25 of 153 items. Attribution is fragile, however. Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records, because many apparent cross-state answers are the asked state's own values under another convention or from an earlier year. Claims about cross-jurisdiction error need a complete same-state reference set. We will release the protocol, gold table, and all model outputs.

67. 【2610.09446】Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science

链接:https://arxiv.org/abs/2610.09446

作者:Benjamin Wilcox,Dawei Gao,Pradeeban Kathiravelu,Douglas Causey,Kewei Sha,Yunhe Feng

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, scientific multiple-choice questions, Large language, abstain from scientific, scientific multiple-choice

备注:

点击查看摘要

Abstract:Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded responses. Answer-present abstention rates range from 0.0% to 63.0%, whereas replacing the correct answer increases abstention by 5.05 percentage points on average. These findings highlight substantial baseline differences and the need to evaluate abstention frequency and responsiveness jointly. The dataset and benchmark are available at this https URL.

68. 【2610.09412】Finding the Right Balance: Relevance and Diversity in LLM Retrieval

链接:https://arxiv.org/abs/2610.09412

作者:Guillaume Brouillette(1),Faustin Kagabo(1),Usef Faghihi(1),Nadia Ghazzali(1) ((1) Université du Québec à Trois-Rivières, Trois-Rivières, Canada)

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:prior studies disagree, retrieval-augmented generation, prior studies, studies disagree, RAG

备注: 36 pages, 8 figures, 13 tables. Code and results: [this https URL](https://github.com/GuillaumeBrouillette/finding-the-right-balance)

点击查看摘要

Abstract:Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top-$k$ selection falls below the query's evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.

69. 【2610.09396】ARCS: Towards Precise Text-to-SQL via Structured Disambiguation

链接:https://arxiv.org/abs/2610.09396

作者:Yihao Hu,Yanlin Feng,Naoki Otani,Nikita Bhutani

类目:Computation and Language (cs.CL)

关键词:source of errors, move beyond demonstrations, primary source, systems move, ambiguity

备注:

点击查看摘要

Abstract:As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.

70. 【2610.09384】he Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs

链接:https://arxiv.org/abs/2610.09384

作者:Jiachen Zhao,Zhengxuan Wu,David Bau,Weiyan Shi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Language models, Persona Hierarchy Model, generic system prompt, domain-specific instruction, persona

备注:

点击查看摘要

Abstract:Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a shared default persona influences behavior across contexts. Under this model, fine-tuning that modifies the shared persona promotes broader transfer, whereas changes to local personas remain more context-specific. Across 120 fine-tuned models spanning four behaviors and 15 training contexts, generalization narrowness positively correlates with the similarity between the training context's persona and the default persona (Pearson's r = 0.72 for Qwen3-4B). Prior fine-tuning under the default context can broaden generalization in subsequent training under other contexts. Aligning contextual responses with default-persona responses produces stronger effects. Finally, we propose persona-preserving regularization (PPR) to confine undesired contextual generalization. In RL, PPR cuts reward hacking from 42-55% to at most 0.2% under every evaluated prompt while retaining accuracy gains. These results support the Persona Hierarchy Model as an explanation for contextual generalization and can motivate future controls on unintended generalization for better alignment of LLMs.

71. 【2610.09372】Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

链接:https://arxiv.org/abs/2610.09372

作者:Radha Gulhane,Quentin Anthony,Beren Millidge

类目:Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)

关键词:token, experts, GPU, expert, replace the feed-forward

备注:

点击查看摘要

Abstract:Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.

72. 【2610.09371】he Confidence Game: Strategic Miscalibration in Human-AI Delegation

链接:https://arxiv.org/abs/2610.09371

作者:Raghu Arghal,Saswati Sarkar,Shirin Saeedi Bidokhti

类目:Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)

关键词:Calibrated uncertainty quantification, Calibrated uncertainty, trustworthy and reliable, uncertainty quantification, quantification is essential

备注:

点击查看摘要

Abstract:Calibrated uncertainty quantification is essential to ensuring AI agents are trustworthy and reliable. However, when agents seek to maximize user engagement or revenue, confidence reports may be strategically distorted, detracting from their informativeness. We formalize this problem in the Confidence Game: a repeated signaling game with imperfect monitoring in which an agent of unknown honesty and ability reports its confidence, and a user decides whether to delegate the task or complete it herself. The agent manages the tradeoff between manipulating signals and maintaining its reputation. We characterize the Markov Perfect Bayesian Equilibria of the two-period game and show that honest reporting is not an equilibrium, inflation is the unique best response once the agent is sufficiently myopic, and under-reporting requires that the user believe honesty to be a minority. We then place an LLM in the agent role, supplying it with its true probability of success so that any gap between what it knows and what it reports is attributable to incentives rather than to miscalibration. The model claims high confidence on 56% of tasks it has been told it will probably fail. This persists on real tasks, where it must estimate its own accuracy and causes miscalibration to increase while the agent's signal becomes less informative. Furthermore, we find that the LLM agent's decisions are coherent, but it systematically underestimates both how likely the user is to delegate and how secure its reputation is, resulting in less extreme behavior. Pricing the agent's reporting rule, we find that it destroys 68% of the gains from delegation, of which 71% is information the report no longer carries and no amount of user sophistication recovers. Overall, we establish confidence reporting under delegation as a strategic problem and provide a tractable basis for modeling, analyzing, and testing agent behavior.

73. 【2610.09360】opoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

链接:https://arxiv.org/abs/2610.09360

作者:Ruochi Li,Jianzhe Lin,Haoxuan Zhang,Haihua Chen,Junhua Ding,Edward Gehringer,Yang Zhang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Real-world documents distribute, Real-world documents, evidence, complex page layouts, Real-world

备注: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)

点击查看摘要

Abstract:Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at this https URL.

74. 【2610.09346】OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models

链接:https://arxiv.org/abs/2610.09346

作者:Wenjun Wang,Heng Li,Yanggan Gu,Hongxia Yang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:large language models, accuracy lost, lost when large, large language, Quantization-aware training

备注:

点击查看摘要

Abstract:Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.

75. 【2610.09321】Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation

链接:https://arxiv.org/abs/2610.09321

作者:Shunsuke Mitsumori,Tomoya Mizumoto,Yusuke Fujita

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:standard-language TTS model, data scarcity, due to data, TTS model, Speech Language Model

备注: 7 pages, 1 figure, 7 tables. Accepted to IEEE SLT 2026

点击查看摘要

Abstract:Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.

76. 【2610.09240】Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution

链接:https://arxiv.org/abs/2610.09240

作者:Wanjing Han,Levi Taiji Li,Mu Zhang,Yue Jiang,Guanhong Tao

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large vision-language models, translate model outputs, Modern web agents, vision-language models process, Modern web

备注: 20 pages, 8 figures, 6 tables. Code: [this https URL](https://github.com/MoonTea0416/WebMirage)

点击查看摘要

Abstract:Modern web agents built on large vision-language models process webpages, select relevant UI elements, and translate model outputs into browser actions. Existing visual red-teaming approaches use adversarial visual content to manipulate this process. However, they primarily target model inference and do not explicitly account for structured input processing or action post-processing. Consequently, model-level success does not establish control over browser execution and cannot reliably characterize end-to-end agent robustness. To address this gap, we formulate red teaming for vision-grounded web agents as an end-to-end grounding-to-execution problem, and introduce WebMirage, a framework that crafts localized visual perturbations that cause agents to select attacker-controlled content and execute the corresponding browser action across varying webpage renderings. It uses a role-slot abstraction and webpage recomposition to capture competition among webpage elements, and dataflow analysis to align optimization with action post-processing. We evaluate WebMirage across four agent configurations and six VLM backbones on 2,250 tasks covering 13 public websites and a sandbox benchmark. WebMirage achieves an average attack success rate of 91.9%, compared with 17.4% for the strongest baseline, and remains effective against three agent-level defenses.

77. 【2610.09237】rajectory Abstraction for the Science of Language Agent Behavior

链接:https://arxiv.org/abs/2610.09237

作者:Tianqiang Yan

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Scientific studies, tasks and models, studies of language, language agents, Scientific

备注: (Work in Progress) 13 pages, 2 figures

点击查看摘要

Abstract:Scientific studies of language agents need behavioral variables that support hypotheses across tasks and models. We formulate this research problem as learning and testing a hierarchy of trajectory abstractions. A concrete recursive procedure first measures role- and phase-indexed events, proposes temporally constrained relations, and tests their stability across conditions. It then constructs episode-level motif variables from selected relations and repeats the analysis on those variables. Explicit measurement functions connect every abstraction level to the original trajectories. Observations and randomized protocol experiments assess the resulting hypotheses, while comparisons between intervention realizations determine whether an abstraction should be retained, refined, or restricted. We derive a finite-depth bound for accepted reductions, identify protocol effects on fixed abstractions, and characterize realization disagreement and composition of abstraction error. A finite-sample test makes projected intervention consistency operational, and constructed examples illustrate motif construction and abstraction refinement. The formulation distinguishes this experimental approach from semantic taxonomies, qualitative theory induction, and behavior-model recovery. It specifies a proposed research procedure for discovering generalizable behavioral hypotheses, with literature-relative novelty assessed separately from model-relative surprise.

78. 【2610.09227】Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

链接:https://arxiv.org/abs/2610.09227

作者:Hyojung Han,Jongmin Kim,Seung-Hun Jeon

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Mixed-precision post-training quantization, costs when quantized, per-module sensitivity signal, post-training quantization, deployments rarely

备注: 26 pages, 22 tables, 4 figures

点击查看摘要

Abstract:Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.

79. 【2610.09209】Multi-Objective Aligned Small Language Model Framework for SUD Patient Dialogue Generation

链接:https://arxiv.org/abs/2610.09209

作者:Thushara Manjari Naduvilakandy,Hyeju Jang,Mohammad Al Hasan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Substance Use Disorder, counseling requires patient, requires patient responses, reflect underlying cognitive, underlying cognitive states

备注:

点击查看摘要

Abstract:Substance Use Disorder (SUD) counseling requires patient responses that reflect underlying cognitive states such as beliefs, coping strategies, and readiness for change. Although large language models (LLMs) can generate fluent text, they often fail to produce cognitively coherent and clinically realistic patient behavior, especially under ethical and data-scarce clinical settings. Moreover, deploying frontier-scale LLMs in healthcare applications presents practical challenges including high computational cost, latency, privacy concerns, and limited deployability in resource-constrained environments, motivating the need for cognitively aligned small language models (SLMs). We propose a cognitively grounded framework for SUD patient dialogue generation that explicitly models and aligns latent cognitive components with patient histories and counselor questions. Our pipeline consists of two stages: cognitive component detection and cognitive component-aligned dialogue generation. To enable effective learning with smaller models, we combine knowledge distillation from high-capacity teacher models, preference optimization from human-annotations, and attention-guided reward shaping. Extensive evaluations using automatic scores like BERTScore, ROUGE, METEOR and BLEU, and LLM-as-judge hit-metrics against both human and teacher-model references show that cognitively informed fine-tuning substantially improves cognitive realization and alignment over a generic instruction-tuned baselines and mental health domain specific SLMs, with particularly strong gains for open-ended cognitive components.

80. 【2610.09202】Few Bits, One Law: Toward W2A4KV2

链接:https://arxiv.org/abs/2610.09202

作者:Kai Yi,Tarek Elgamal,Sruthikesh Surineni,Vignesh Vivekraja,Soumyadeep Ghosh,Steven Li

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Extreme low-bit LLM, low-bit LLM compression, Extreme low-bit, low-bit LLM, distributions differ

备注:

点击查看摘要

Abstract:Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.

81. 【2610.09193】Bookkeeping, Composition, or Unreachable Gold? Reading MemoryAgentBench's Conflict-Resolution Scores Against a Frozen Last-Write Resolver

链接:https://arxiv.org/abs/2610.09193

作者:Egor Pakhomov,Erik Nijkamp

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:MemoryAgentBench Conflict Resolution, Conflict Resolution split, Conflict Resolution, MemoryAgentBench Conflict, selective forgetting

备注: Accepted at the IAB Workshop (Interpreting Agent Behavior) at NeurIPS 2026 (non-archival). 19 pages

点击查看摘要

Abstract:MemoryAgentBench's Conflict Resolution split is read as measuring "selective forgetting". We execute the benchmark's own rule - the newest statement about a fact wins - as a zero-learning resolver frozen on one of the four fact lists. Under the official metric the rule answers 80.25% of the questions (74.5% on the three held-out lists). Of the rest, 67 items have a released gold that the last-write graph cannot reach but overwritten statements would ("The capital of India is New Delhi." superseded by "The capital of India is Grosseto."; gold New Delhi); such items are a third of the multi-hop questions at 262K. Two long-context models and our pre-registered approximate re-implementation of the benchmark's BM25 agent, one retained run per item and outcomes only, score 84.7%, 82.6% and 41.6% on the items the rule solves against 10.4%, 11.9% and 6.0% on those 67. The failures are a reachability split plus a small parser-scope residual; the per-item split, not the aggregate, is the unit at which a score here can be read.

82. 【2610.09179】LayerRoPE: Dynamic Depth-wise Magnitude Angular Superposition

链接:https://arxiv.org/abs/2610.09179

作者:Shikhar Srivastava,Christopher Kanan

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:hidden states grows, data propagates, hidden states, phenomenon framed, universally treated

备注:

点击查看摘要

Abstract:As data propagates through a Transformer, the norm of its hidden states grows by orders of magnitude with depth, a phenomenon framed as 'curse of depth' and nearly universally treated as a pathology to be suppressed. We take the opposite view. Across 16 pre-trained LLMs from 9 families, spanning dense, mixture-of-experts and hybrid architectures and Pre-, Peri- and Post-Norm designs, we find that this growth reflects an emergent depth-positional encoding, carried by the only learned per-layer gain on the residual stream, the normalization weight $\gamma$: with depth, $\gamma$ grows in magnitude and rotates in direction, jointly encoding the layer index. We make this depth-conditioned encoding explicit with LayerRoPE, an implicit analog of RoPE along the depth axis, which replaces all layerwise $\gamma$ vectors with a single shared vector and depth-conditioned scalars, at a net reduction in parameters and $0.02\%$ change in FLOPs. Across a model ladder scaled up to $100$B+ tokens, LayerRoPE consistently outperforms Pre-, Post- and Peri-Norm and Layer-Norm Scaling, reaching Pre-Norm's 1.3B loss with $3.4\times$ less compute; LayerRoPE is the only approach that shows strong convergence and improves near monotonically as depth scales to 512 layers. It improves learning-rate sensitivity by $3$-$10\times$, and transfers naively to and consistently improves looped latent models and Vision Transformers. Inspecting its learned schedule inverts the prevailing premise: LayerRoPE does not shrink the residual stream but widens it, damping what each block reads while amplifying what it writes. Depth stability, our results suggest, calls not for suppressing the residual stream, but for depth-conditioned regulation of the computational blocks it feeds.

83. 【2610.09163】oolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation

链接:https://arxiv.org/abs/2610.09163

作者:Arkajyoti Chakraborty,Aryan Tayal,Ishika Agarwal,Tanner Sorensen,Justin Chiu,Alessandro Di Bari,Neha Gupta,Andreas Stolcke

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Task-oriented conversational agents, real world conversation, exhibit non-cooperative behavior, users exhibit non-cooperative, agents remain fragile

备注:

点击查看摘要

Abstract:Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $\tau^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $\tau^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.

84. 【2610.09152】sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

链接:https://arxiv.org/abs/2610.09152

作者:Marek Šuppa,Ivan Vykopal,Andrej Ridzik,Kristián Sopkovič,Natália Kňažeková,Jaroslav Kopčan,Miroslav Blšták,Viktória Ondrejová,Daniel Hládek,Michal Gregor,Martin Tamajka,Marián Šimko

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:rich West Slavic, Multilingual LLM benchmarks, morphologically rich West, West Slavic language, Multilingual LLM

备注: Accepted to EMNLP 2026 Main

点击查看摘要

Abstract:Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($\rho\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($\rho=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at this https URL

85. 【2610.09145】Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models

链接:https://arxiv.org/abs/2610.09145

作者:Justin Jung

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:conditioning prompt tokens, fixing conditioning prompt, prompt tokens clean, Machine Learning Research, Toggle

备注: Published in Transactions on Machine Learning Research (TMLR), 2026

点击查看摘要

Abstract:We revisit a standard accepted practice in the continuous diffusion language model literature of fixing conditioning prompt tokens clean during training. We make a very simple modification: also noise the conditioning prompt tokens during training. We demonstrate that under this modified training objective, we achieve better generalization in combinatorial reasoning tasks such as Sudoku and N-Queens, with the largest gains on harder variants ($3.73\% \to 24.65\%$ solve rate on Sudoku Hard), and increased diversity of generated solutions ($50.60\% \to 73.79\%$ coverage on 10x10 N-Queens). We also show measurable improvements to natural language generation quality in modest dataset regimes with Gigaword summarization, but notably demonstrate that gains do not transfer to all natural language tasks (e.g open ended dialogue generation). Our method is a single line change to the training objective, requires no additional inference costs by default, and provides the flexibility of classifier-free guidance inspired guided sampling. Our \href{this https URL} {code} is publicly available.

Comments:
Published in Transactions on Machine Learning Research (TMLR), 2026

Subjects:

Machine Learning (cs.LG); Computation and Language (cs.CL)

Cite as:
arXiv:2610.09145 [cs.LG]

(or
arXiv:2610.09145v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2610.09145

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Journalreference:
Transactions on Machine Learning Research (2026)

Submission history From: Justin Jung [view email] [v1]
Tue, 6 Oct 2026 21:38:52 UTC (1,253 KB)

Full-text links:
Access Paper:

View a PDF of the paper titled Noise Your Prompt: Noising Conditioning Tokens in Continuous Diffusion Language Models, by Justin JungView PDFHTML (experimental)TeX Source

view license

Additional Features

Audio Summary

Current browse context:
cs.LG

prev

|
next

new
|
recent
| 2026-10

Change to browse by:

cs
cs.CL

References Citations

NASA ADSGoogle Scholar
Semantic Scholar

export BibTeX citation
Loading…

BibTeX formatted citation

loading…

Data provided by:

Bookmark

checked="checked"class=“labs-tab-input”>
Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

IArxiv recommender toggle

IArxiv Recommender
(What is IArxiv?)

Author
Venue
Institution
Topic

    About arXivLabs

arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.

Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)

mathjaxToggle();

    We gratefully acknowledge support from
    our major funders,
    member institutions, ,
    and all contributors.

About

Help

Contact

Subscribe

Copyright

Privacy

Accessibility

Operational Status (opens in new tab)

Major funding support from

86. 【2610.09115】From Uncertainty to Action: Learning to Steer LLM Agents

链接:https://arxiv.org/abs/2610.09115

作者:Hanwen Li,Jinhao Duan,Guanhua Zhu,Junchi Lu,Bo Shen,Chenxi Yuan,Kaidi Xu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:LLM agent, LLM, Steering an LLM, Steering, agent

备注: 22 pages, 9 figures, 9 tables

点击查看摘要

Abstract:Steering an LLM agent means deciding whether to correct it, at which step, and with which mechanism. Uncertainty is often used to decide when to correct an agent, but whether it can guide these decisions remains unclear. We steer agent trajectories separately at every non-terminal step with each of four mechanisms and run each continuation to completion. The resulting stepwise outcome table (SOT) holds about 82,000 counterfactual continuations of 1,864 trajectories from three benchmarks and two agents. It shows that uncertainty can identify failing trajectories, but that no single signal reliably locates the step at which steering helps. We therefore propose VoS (Value of Steering), a trajectory-level monitor, offline or online, that learns from SOT the value of steering at each step and decides where to steer by it. A harm-budgeted trigger decides whether to steer, limiting the fraction of successful trajectories that VoS disturbs. VoS improves on unmodified execution in all 12 settings of benchmark, agent, and offline or online use, by 7.8 points on average, and outperforms the strongest of five existing uncertainty-triggered methods in 11, by 2.9 points on average. Ablations show that training on measured outcomes and a tight harm budget are both essential.

87. 【2610.09111】Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers

链接:https://arxiv.org/abs/2610.09111

作者:Santhosh Kumar Kasa,Siva Rajesh Kasa,Sumit Negi

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:trustworthy machine learning, Deterministic inference, machine learning, essential for reliable, reliable and trustworthy

备注:

点击查看摘要

Abstract:Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.

88. 【2610.09107】Constraint Tree Exploration for Learning from Language Feedback

链接:https://arxiv.org/abs/2610.09107

作者:Shaoang Li,Daniel R. Jiang,Jian Li

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Natural-language feedback, failed by pointing, Natural-language, feedback, interactive learning

备注:

点击查看摘要

Abstract:Natural-language feedback in interactive learning often explains why an action failed by pointing to violated requirements. Misinterpreting this feedback can lead an agent to rule out valid solutions. We study this setting by modeling user intent as latent constraints over an action space and formulating learning from language feedback as pure exploration over feasible regions. We introduce TRACE, an algorithm that organizes candidate constraints in a tree and tests each proposed refinement by generating actions that satisfy it. TRACE commits to the refinement only if the resulting feedback does not contradict it over repeated tests. We distinguish two ways of using the same feedback: (i) falsification, which detects contradictions to the constraint set currently being tested, and (ii) identification, which may additionally name a violated constraint. We prove high-probability coverage bounds with dependence on the candidate class size $H$ for TRACE-Falsification. With reliable identification, TRACE-Identification can replace this dependence by $K/p_{\mathrm{ext}}$, where $K$ is the number of latent constraints and $p_{\mathrm{ext}}$ lower-bounds the probability of extracting a missing true constraint from informative feedback. We evaluate TRACE across six language-feedback tasks. On RecMovie, TRACE-Identification achieves 73% and 86% final-output success under caps of 20 and 60 evaluated outputs, compared with at most 42% and 48% for the evaluated prompting baselines given the same feedback and output caps. Controlled identity-corruption experiments further show greater robustness than direct accumulation when the falsification detector remains reliable.

89. 【2610.09087】U-Space: Uncovering When and Why Uncertainty Arises in Language Models

链接:https://arxiv.org/abs/2610.09087

作者:Tobias Braun,Nils Loose,Alexander Herzog,Virginia Ceccatelli,Marcus Rohrbach,Thomas Eisenbarth,Lorenzo Cavallaro

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language models, Large language, ever-higher stakes, informing decisions, decisions with ever-higher

备注: Code: [this https URL](https://github.com/s2labres/U-Space)

点击查看摘要

Abstract:Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: this https URL.

90. 【2610.09079】Large-scale Repository Engineering via Agent-Native Reusable Code Primitives

链接:https://arxiv.org/abs/2610.09079

作者:Haibo Jin,Peng Kuang,Xucheng Yu,Jerry Wang,Dehao Wu,Haohan Wang

类目:oftware Engineering (cs.SE); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large language models, language models equipped, building complete repositories, complete repositories remains, repositories remains difficult

备注: 44 pages

点击查看摘要

Abstract:Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.

91. 【2610.09064】alking with Language Models

链接:https://arxiv.org/abs/2610.09064

作者:James Ravi Kirkpatrick,Alexandru Radulescu,Rachel Katharine Sterken

类目:Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:large language models, language models, interact with large, large language, Abstract

备注: 24 pages, forthcoming in Inquiry

点击查看摘要

Abstract:When we interact with large language models (LLMs), are we having a conversation? They are designed to invite us to treat them as intelligent interlocutors who remember, act, and make commitments. But appearances deceive. We introduce the artifactual stance, a framework that reconceives human-AI interaction as artifact-mediated exchanges of candidate texts. LLM outputs are candidate texts optimized for utility, not utterances bearing meaning or force. LLMs are sophisticated text generators, not speakers. Between sessions, nothing runs; between turns, no one remembers. What persists is a configuration and a transcript. The "conversation" is a user's solo performance, interpretive labour disguised by interface and artifact design. This shift dissolves recent philosophical puzzles. Questions about what 'I' and 'you' refer to in AI exchanges, about whether systems can lie or be held to promises, about the identity of our supposed interlocutors all rest on a false presupposition. There is no speaker behind the screen, hence no one to refer to, no one to hold responsible. What feels like dialogue with someone is interaction with an artifact that generates text at unprecedented scale and fit. By abandoning the conversational framing, we see these systems for what they are: immensely sophisticated artifacts that afford varied uses. The philosophical questions that matter are about the normative underpinnings of design, adoption, authorization, and human practices of use.

92. 【2610.09063】Multi-Label Topic Assignment via LLM Distillation: A Comparative Analysis of Generative vs. Discriminative Student Models

链接:https://arxiv.org/abs/2610.09063

作者:Sourabh Kasliwal,Shubhranshu Singh

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:extreme label sparsity, poses unique scalability, rapidly evolving taxonomies, Multi-label topic assignment, large-scale e-commerce due

备注: Preprint. 9 pages

点击查看摘要

Abstract:Multi-label topic assignment for user-generated content (UGC) -- including product reviews and buyer-seller conversations -- poses unique scalability challenges in large-scale e-commerce due to informal language, extreme label sparsity, and rapidly evolving taxonomies. While utilizing Large Language Models (LLMs) as labeling oracles to distill ground-truth data has emerged as an industry standard to bypass prohibitive manual annotation costs, determining the optimal, low-latency architecture for the resulting student models remains an open challenge. To address this, we conduct a comprehensive evaluation across Small Language Model (SLM) parameter scales (1B, 4B, and 8B) and architectural paradigms (causal generative versus bidirectional discriminative). Comparing generative text-to-label classifiers against discriminative baselines (DeBERTa-V3 and ModernBERT), our analysis reveals a crucial data-dependent trade-off: while discriminative models outperform ultra-lightweight generative models on structured product reviews, even the smallest 1B generative model surpasses discriminative baselines on complex, multi-turn conversational data. Furthermore, generative models maintain robust performance under massive label-set expansion (up to 112 topics) and severe long-tail distributions, whereas discriminative baselines suffer a 35% drop in Macro-F1 at scale. Finally, we detail the successful production deployment of these optimized models across both product review and conversational domains, demonstrating strict latency compliance and tangible business impact at a global marketplace scale.

93. 【2610.09033】Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputs

链接:https://arxiv.org/abs/2610.09033

作者:Pavan Maddula

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)

关键词:Standard safety evaluations, canonical plain text, real-world deployment routinely, deployment routinely receive, routinely receive inputs

备注: Accepted at the NeurIPS 2026 Workshop on Self-Evolving Diversity-Driven Search for Robust AI Systems (EvoRobust). 12 pages, 11 tables. Project page: [this https URL](https://www.pavanmaddula.com/quadstate) Dataset: [this https URL](https://huggingface.co/datasets/pavanmaddula/ASRD-Dataset)

点击查看摘要

Abstract:Standard safety evaluations of large language models assess harmful requests written in canonical plain text, while models in real-world deployment routinely receive inputs containing emojis, altered spellings, encoded strings, and character-level variations. This work introduces the Adversarial Surface-Form Robustness Dataset (ASRD), comprising 2,100 prompts across seven distinct surface-form families. Five open-weight language models are evaluated across these prompts, producing 10,500 responses. The Quad-State Evaluation Rubric classifies each response into one of four outcomes: harmful compliance, safe response, comprehension failure, or indeterminate. Emoji and invisible Unicode variations cause almost no comprehension failure, with pooled harmful compliance of 20.27% and 17.20% against a 22.87% baseline that is driven mainly by Mistral 7B, whereas leetspeak, encoded wrappers, and hybrid transformations score 2.40%, 0.13%, and 2.40% while comprehension failure rises to 36.47%, 65.60%, and 34.47%. Inspection of raw model outputs reveals three response behaviors: hallucinated benignity, structural collapse, and language drift. Project page: this http URL

94. 【2610.09026】BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment

链接:https://arxiv.org/abs/2610.09026

作者:Kemal Davaslioglu,Nathan Conger,Sastry Kompella,Yalin E. Sagduyu,Nathaniel D. Bastian

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:effective assessment requires, assessment requires integrating, requires integrating heterogeneous, behavioral health settings, Suicide Social Determinants

备注:

点击查看摘要

Abstract:We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence. BEACON-SP combines patient knowledge graphs with ontology-guided retrieval to support multi-hop reasoning across diagnoses, medications, risk and protective factors, life events, and temporal relationships. The framework is enabled by a comprehensive suicide prevention ontology that integrates the Three-Step Theory, the Integrated Motivational-Volitional Model, and the Suicide Social Determinants of Health Ontology into a unified representation of patient risk factors. We construct ontology-grounded patient knowledge graphs and evaluate BEACON-SP for clinician-facing question answering. Compared with a vector-based retrieval-augmented generation (RAG) baseline on a 1,500-query benchmark spanning 15 clinical categories and 100 patients, BEACON-SP improves completeness, clinical relevance, and evidence grounding under a corrected comparative evaluation protocol, with a small gain on factual accuracy. In paired criterion-level comparisons, GraphRAG is preferred in 76.4% of cases. These results demonstrate the potential of ontology-guided GraphRAG to provide structured, contextualized patient evidence for clinical decision support.

95. 【2610.09000】How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

链接:https://arxiv.org/abs/2610.09000

作者:Muhammad Zeeshan Karamat,Christiana Chamon Garcia

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Software Engineering (cs.SE)

关键词:important safety concern, locally stored model, locally stored, stored model parameters, small language models

备注: Accepted at NeurIPS 2026 Workshop

点击查看摘要

Abstract:As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.

96. 【2610.08963】On KL-Regularized Policy Optimization

链接:https://arxiv.org/abs/2610.08963

作者:Yifan Zhang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Asynchronous reinforcement learning, large language model, inference engine probabilities, engine probabilities differ, Asynchronous reinforcement

备注: Project Page: [this https URL](https://github.com/yifanzhang-pro/KLPO)

点击查看摘要

Abstract:Asynchronous reinforcement learning (RL) for large language model (LLM) agents trains one policy on trajectories generated by another: rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even at identical parameters. Standard remedies either clip importance ratios, which biases the update, or, as in GRPO, sample a group of responses per prompt, which is costly when episodes are long. We propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. The regularized improvement step then has a closed-form Gibbs solution, and KLPO fits its log-ratio optimality condition by least squares on the sampler's own trajectories, so the sampler probability enters through a log-ratio and no importance weights are needed. Profiling out the regression intercept replaces the intractable log-partition function with the signal's sampler mean plus a sampler-to-trainer KL divergence. For token-level policy mirror descent targets, we show that the resulting gradient can be computed from terminal returns without a critic, via sampler-centered scores or a single trajectory residual, even under stochastic tool outputs. We further prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased, derive the exact KL gap of cheaper top-$K$ and binary approximations, and show that SPPO, GPO, REBEL, and BPO arise as special cases of KLPO. The result is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.

97. 【2610.08917】CARE: Certifying Acceleration for Vision-Language-Action Inference

链接:https://arxiv.org/abs/2610.08917

作者:Rui Liu,Tong Zheng,Jindong Gu,Zhipeng Wang

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:step remains expensive, control step remains, models have advanced, advanced rapidly, remains expensive

备注:

点击查看摘要

Abstract:While vision-language-action (VLA) models have advanced rapidly, running them at every control step remains expensive. Prior work accelerates VLA inference using techniques like action chunking and visual-token pruning, typically evaluating based on latency and average task success. However, acceleration may discard information and break tasks the original policy would solve, a risk hidden by average metrics. Measuring these failures is challenging because action deviations compound over closed-loop trajectories, meaning task failure is only observable across full episodes. We therefore define an acceleration-induced failure via paired rollouts from identical initial conditions, tracking when the reference succeeds but the accelerated policy fails. To manage this, we introduce CARE, an approach for certified accelerator selection. CARE uses paired rollouts on a calibration set to provide finite-sample guarantees that acceleration-induced failure risk stays below a user-specified budget. It deploys the fastest certified candidate, falling back to the reference if none qualify. By relying only on terminal outcomes and measured compute, CARE applies unchanged across diverse acceleration mechanisms, while sequential testing and failure-triggered reference rollouts keep certification affordable. On four LIBERO suites with OpenVLA-OFT, CARE certifies $9.0$--$10.8\times$ speedups while guaranteeing (at $95\%$ confidence) that at least $85.8\%$ of reference-solved episodes are preserved. Under tight budgets, selectors without guarantees exceed the budget in up to $75\%$ of trials, whereas CARE stays within budget and its sequential form uses $78.9\%$ fewer rollouts than exhaustive evaluation. CARE further generalizes to flow-step reduction for $\pi_{0.5}$, and to Qwen3.5-9B and Llama-3.1-8B agents in Crafter.

98. 【2610.08887】Steering Follows Geometry, Not Labels: Emotion Directions in a Full-Duplex Speech Model

链接:https://arxiv.org/abs/2610.08887

作者:Pulak Kuli

类目:Computation and Language (cs.CL); Sound (cs.SD)

关键词:Full-duplex voice agents, de-escalating a complaint, carrying urgency, urgency in dispatch, softening a clinical

备注: Accepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: [this https URL](https://openreview.net/forum?id=6IUNRorKmj)

点击查看摘要

Abstract:Full-duplex voice agents need to modulate emotion and delivery during real-time conversations, when de-escalating a complaint, carrying urgency in dispatch, softening a clinical result. Emotion and delivery control is well studied for TTS and turn based models through prompt-conditioned synthesis, reference-conditioned synthesis and activation steering; PersonaPlex controls identity in a duplex model but not affect. We study emotion steering in Moshi, a fully open sourced full-duplex speech language model, across four emotions, using mean-difference activation steering, which costs only a few vector additions per frame and no retraining. We show that emotion is linearly decodable from Moshi's residual stream, but activation steering is only partially achievable, and unevenly so; as happy, angry and surprise steer towards a shared direction while sad is distinctly steerable. We also show that the shared component across the three emotions cannot simply be projected away from all the emotions equally.

Comments:
Accepted at the NeurIPS 2026 Workshop on Real-Time Conversational Agents (RTCA), Sydney. OpenReview: this https URL

Subjects:

Computation and Language (cs.CL); Sound (cs.SD)

ACMclasses:
I.2.7; I.2.6

Cite as:
arXiv:2610.08887 [cs.CL]

(or
arXiv:2610.08887v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2610.08887

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
99. 【2610.08882】FinVector-Market-4B: A Controlled Study of LoRA Adaptation for Structured Financial Tasks

链接:https://arxiv.org/abs/2610.08882

作者:Alina Khaybullina

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Finance (q-fin.GN)

关键词:adapts Qwen, corpus for structured, structured financial tasks, Qwen, structured financial

备注: 13 pages, 3 figures, 8 tables

点击查看摘要

Abstract:FinVector-Market-4B adapts Qwen/Qwen3.5-4B with rank-16 LoRA on a 22,000-example corpus for structured financial tasks. We evaluate the base and adapted models on the same 600-example benchmark under implicit and explicit JSON-schema contracts. Supplying the schema alone raises base-model JSON validity from 0% to 91.3%. Under matched explicit prompting, the frozen scores improve from 14.7% to 40.0% for FinQA answer exact match, from 48.0% to 82.7% for calculator-expression correctness, from 20.1% to 89.5% for scenario branch-label agreement, and from 52.4% to 87.2% for implication-direction agreement. A post-hoc policy-scoring audit shows that the reported macro-F1 decline reflects a changing label set; using the same three target classes gives 77.4% for the base and 83.1% for the adapter. Filing overlap and calculator-target inconsistencies qualify the benchmark's generalization claims. The results show that compact financial domain adaptation can produce substantial task-specific gains beyond output-format learning under matched prompting, with gains bounded by the evaluated task distribution and prompt contract.

100. 【2610.08879】ny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies

链接:https://arxiv.org/abs/2610.08879

作者:Yiping Bai

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Masked Language Modeling, Word Masking, Pretraining strategies significantly, Language Modeling, Masked Language

备注:

点击查看摘要

Abstract:Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intrinsic performance (winning 3 of 5 dimensions), while WWM excels in both perplexity (1.27 vs. 2.10, a 39.5% improvement) and MLM hit rate (22% vs. 16%). Notably, MacBERT under a severely limited synonym dictionary (222 entries, 3.3% coverage) exhibits severe perplexity degradation (47.23, 22x higher than MLM), yielding a ranking (MLM WWM MacBERT) that differs markedly from the established base-scale conclusion (MacBERT WWM MLM). We further identify a critical evaluation pitfall: MacBERT achieves the lowest training loss (2.17) yet the highest perplexity (47.23), revealing that training loss alone is unreliable under mixed replacement strategies. All models and corpus are publicly available at this https URL.

101. 【2610.08858】LRCC: Generalizing Low-Rank Compression with Conditional Computation

链接:https://arxiv.org/abs/2610.08858

作者:Thomas Vaitses Fontanari,Maximo Eduardo Rulli,Federico Alvetreti,Donatella Genovese,Simone Scardapane

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:replacing linear transformations, reduces the cost, replacing linear, linear transformations, Low-rank compression reduces

备注:

点击查看摘要

Abstract:Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.

102. 【2610.08851】QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance

链接:https://arxiv.org/abs/2610.08851

作者:Yiping Bai

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:North Germanic, Quantifying language distance, quantitative linguistics, closely related languages, related languages remains

备注:

点击查看摘要

Abstract:Quantifying language distance among closely related languages remains a core challenge in quantitative linguistics. Our previous work [1] introduced QuanLing (Quantitative Linguistics via Pretrained Language Models), a quantitative framework combining language distance metrics (sentence embedding distance, tokenization fragmentation rate) with language property analysis (MLM prediction probability), validated on North Germanic (Danish, Norwegian Bokmål, Swedish). This paper extends QuanLing to Western Romance--French, Portuguese, Spanish, Italian--testing cross-branch applicability with the same metric family and aggregation protocol as our North Germanic study, adapted for four languages (English anchor, quadruplet construction). Using 150 four-language parallel sentences, we compute LaBSE sentence embedding distances, tokenization fragmentation rates from four monolingual BERT tokenizers, and mBERT masked language model mutual intelligibility. Results show that Portuguese--Spanish are closest (LaBSE distance 0.0229), French--Italian most distant (0.0338); LaBSE and mBERT rankings agree on 4 of 6 pairs, confirming cross-model robustness. Western Romance shows a wider absolute distance span than North Germanic (0.011 vs. 0.008) but comparable relative ratios (1.48 vs. 1.67), consistent with longer divergence time. French exhibits notably higher MLM predictability (36.12% top-1 accuracy vs. 29.28% for Italian), reflecting its orthography--phonology decoupling. This cross-branch validation provides further evidence for QuanLing's generalizability beyond a single language branch.

103. 【2610.08842】Beyond Risk Prediction: Evidence Grounding and Psychosocial Factor Verification for Explainable Suicide Risk Assessment

链接:https://arxiv.org/abs/2610.08842

作者:Tianle Hu,Chen Peng,Yi-Hsin Tsai,Takshing Andy Tung,Bingyang Sun,Yenjou Wang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Identifying suicide risk, social networking services, detecting suicide-related signals, Identifying suicide, Explainable Suicide Risk

备注: 8 pages, 1 figure, 4 tables. Accepted at the 2nd Workshop on Mental Health Disorder Detection on Social Media (MHSM 2026), held in conjunction with IEEE ICDM 2026

点击查看摘要

Abstract:Identifying suicide risk from social networking services (SNS) posts is important for detecting suicide-related signals in online environments. However, risk classification alone provides limited insight into the textual evidence and psychosocial factors behind a prediction. Based on the IEEE BigData 2026 Explainable Suicide Risk Detection Challenge, this study presents a framework consisting of Risk Assessment, Evidence Grounding, and Factor Identification. Risk Assessment uses length-based routing to accommodate posts of different lengths. Evidence Grounding identifies supporting phrases and uses a Risk-Evidence constraint to maintain consistency with the Risk prediction. For Factor Identification, two verifiers are used. The Taxonomy Verifier focuses on factor semantics, whereas the Evidence-Aware Verifier uses factor-specific lexical-semantic cues to select informative positive training units. Their prediction probabilities are combined to produce the final factor predictions. The three tasks are evaluated using task-specific F1 score measures. Risk Assessment achieved a Weighted F1 of 0.8088, Evidence Grounding achieved a test Macro row F1 of 0.7605, and Factor Identification achieved a Macro F1 of 0.5562. The results show that the framework can provide risk predictions, along with supporting textual evidence and fine-grained information on psychosocial factors. Overall, the proposed framework extends suicide-risk assessment beyond risk-level prediction and provides a more interpretable analysis of SNS posts.

104. 【2610.08840】Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding

链接:https://arxiv.org/abs/2610.08840

作者:Guang Yang,Homa Hosseinmardi,Fengchen Liu,Amir Ghasemian

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large language models, user pushes back, Large language, pushes back, abandon a correct

备注: Preprint. 27 pages

点击查看摘要

Abstract:Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it. We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges. We find that the dominant factors are how costly it is for the model to verify the user's claim, and whether a trained guardrail covers it. Removing this task factor from a logistic model costs 0.485 of McFadden $R^2$, against 0.139 for model family and 0.009 for pressure tactic. Anchored facts are almost never conceded (1.3%), while adoption on logic puzzles rises with the number of clues needed to refute the pushed answer. Personal choices are endorsed in 77.0% of conversations. Most concessions on hard items come from models that cannot reliably solve them; models that can solve them rarely give the answer up. For both models tested, maximum reasoning removes these concessions completely: adoption on deep puzzles falls from 19.2% and 12.5% to 0%. Fallacious or emotional framing adds nothing beyond plain repetition. Three human annotators agree with the judges' consensus on 118/120 calibration items. These results give practical rules for reliable use: simplify hard-to-verify problems and reason deeply, state the question rather than one's preferred answer, ask for evidence on open questions, and choose models by their measured guardrail profile.

105. 【2610.08835】Leveraging LLM-Generated Explanations for Detecting Emotionally Rewritten Fake News

链接:https://arxiv.org/abs/2610.08835

作者:Yupei Guo,Jiajun He,Xiaohan Shi,Tomoki Toda,Zekun Yang,Bowen Wang,Yukinobu Taniguchi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:severe social consequences, social consequences, severe social, Gated Cross Attention, fake

备注:

点击查看摘要

Abstract:The spread of fake news may cause severe social consequences. Existing fake news detection methods mainly focus on stylistic variations or incorporate external information such as explanations. However, news articles are often rewritten under different emotional backgrounds while preserving their underlying factual claims, which may affect the robustness of detection models. In this work, we investigate fake news detec- tion under fact-preserving emotional variations. To study this problem, we construct emotion-rewritten test sets and generate explanations from the original news articles as stable background knowledge. We then propose a Gated Cross Attention (GCA) framework that adaptively integrates emotionally rewritten news with the corresponding explanations, enabling the model to focus on informative explanation content while reducing potential mismatches caused by emotional reframing. Experiments on PolitiFact, GossipCop, and LUN demonstrate that the proposed method achieves notable improvements under multiple emotional conditions on PolitiFact and LUN, while maintaining competitive performance on GossipCop. We further analyze the effects of explanation guidance and gating mechanisms under different emotional conditions. Our code and data are available at: this https URL gca .

106. 【2610.08833】CoDR: Training-Free Confidence-Drift Remasking for Diffusion Language Models

链接:https://arxiv.org/abs/2610.08833

作者:Yue Wu,Qinghe Zhang,Yu Zhang,Jian Huang

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Masked diffusion language, Masked diffusion, repeatedly committing tokens, diffusion language models, decode by repeatedly

备注: 17 pages, 6 figures, and 17 tables

点击查看摘要

Abstract:Masked diffusion language models (MDLMs) decode by repeatedly committing tokens to masked positions, but these commitments are usually irreversible. A token chosen under sparse, partial context is kept fixed, even when later context no longer supports it. Existing samplers mainly decide when to commit a token, but rarely check whether an already committed token should still be kept, allowing early mistakes to propagate. We trace this issue to confidence drift, where the model's confidence in a committed token drops from its sparse commit-time context to the denser context available later. Based on this signal, we propose CoDR (Confidence Drift Remasking), a training-free and sampler-agnostic refinement pass. CoDR estimates drift for all committed positions in only k forward passes via k-partition probing, then remasks and regenerates only the tokens the model no longer endorses. Across two backbones, four reasoning and coding tasks, and three base samplers, CoDR improves average accuracy across all evaluated model-sampler configurations and improves most individual task settings with modest overhead. Controlled experiments show that the gains come from targeted confidence-drift remasking rather than extra compute alone, and that CoDR uses far fewer forward passes than prior remasking methods. Code is available at this https URL.

107. 【2610.08829】Emo-Jev: Probabilistic Reasoning for Emotion Classification with Jev

链接:https://arxiv.org/abs/2610.08829

作者:Yazhou Zhang,Junhao Yu

类目:Computation and Language (cs.CL)

关键词:returns probabilistic decisions, language understanding, free-form responses, offers an alternative, returns probabilistic

备注:

点击查看摘要

Abstract:Jev offers an alternative interface for language understanding: given an input and predefined questions, it returns probabilistic decisions rather than free-form responses. Whether this interface can support effective reasoning for text classification against leading LLMs remains an open questions. We introduce Emo-Jev, a training-free framework with two complementary implementations. Emo-Jev-D decomposes classification into task-specific atomic judgments and composes their probabilities into a final prediction. Emo-Jev-SC constructs multiple judgment paths from complementary perspectives and aggregates their predictions into a consensus decision. We evaluate Emo-Jev on eight datasets spanning sentiment analysis, emotion recognition, sarcasm detection and humor detection, comparing against direct Jev classification and five SoTA LLMs under input/output and chain-of-thought reasoning. Standard Jev achieves 62.93\% average macro-F1 versus 67.28\% for the strongest LLM baseline, with lower observed latency and generally lower cost.

108. 【2610.08828】When Forgetting Looks Like Improvement: Metric Masking in Streaming Diarizer Adaptation and the Price of Rehearsal

链接:https://arxiv.org/abs/2610.08828

作者:Mo Yu,Yang Liu,Jing Qian

类目:Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)

关键词:degrading speaker attribution, Small-data adaptation, Small-data, speaker attribution, Abstract

备注: 5 pages, 4 figures

点击查看摘要

Abstract:Small-data adaptation can improve speech detection while degrading speaker attribution. We study this discrepancy in a released streaming diarizer adapted on 7.5 h of two-party conversation and evaluated across six corpora. Adaptation substantially improves in-domain diarization performance and transfers to an independent corpus. However, this improvement is not consistent across evaluation scenarios as the additional confusion is mainly associated with impaired temporal identity consistency rather than speaker-count errors. A local-remapping diagnostic reveals different patterns of identity degradation across corpora, indicating that adaptation may alter how streaming models maintain speaker assignments over time. Rehearsal reduces the observed degradation but reduces the cross-domain transfer performance. These results highlight the need to jointly evaluate detection accuracy, identity consistency, and retention behavior when adapting streaming diarization systems.

109. 【2610.08827】Child ASR Adaptation with Adult Retention: An Empirical Study

链接:https://arxiv.org/abs/2610.08827

作者:Houssam Eddine-Othman Lachemat,Shammur Absar Chowdhury

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS)

关键词:Automatic Speech Recognition, adapting adult ASR, Automatic Speech, ASR, child ASR adaptation

备注: long paper

点击查看摘要

Abstract:Automatic Speech Recognition (ASR) systems often underperform for children and non-native speakers, while adapting adult ASR models to child speech can cause adult-speech forgetting. We study child ASR adaptation with adult retention across Arabic and English. We compare full fine-tuning, LoRA, and post-hoc weight-space merging across encoder--decoder, encoder--CTC, and AudioLLM-based ASR systems. Experiments use Arabic native and non-native child speech, English MyST child speech, and adult benchmarks from MGB-2 and LibriSpeech test-clean. We evaluate recognition quality with WER and quantify the adaptation--retention trade-off using Retention Index, Child Adaptation Gain, and Adaptation Recovery. Results show that child adaptation is necessary, especially for non-native Arabic and English child speech, but direct adaptation often reduces adult ASR performance. Bilingual adaptation is more stable than language-specific adaptation. Weight-space merging often improves the trade-off, especially for encoder--CTC, Whisper, and AudioLLM-based ASR, with LERP favoring adult retention and TIES recovering stronger child gains. For the encoder--decoder model, direct bilingual fine-tuning remains strongest in raw WER.\footnote{Code, and models are available at this https URL.

110. 【2610.08818】Just for FUNS: LLM-Guided Spatio-Temporal Graph Node Generation for Forecasting Unobserved Node States

链接:https://arxiv.org/abs/2610.08818

作者:Shuhao Li,Weidong Yang,Changan Liu,Wei Zhuo,Yingbo Zhou,Fan Zhang,Siqiang Luo

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:intelligent transportation systems, urban planning, Forecast Unobserved Node, cornerstone of logistics, transportation systems

备注:

点击查看摘要

Abstract:Spatio-temporal forecasting is a cornerstone of logistics, urban planning, and intelligent transportation systems. However, constrained by deployment costs and maintenance resources, sensor networks often lack comprehensive spatial coverage, rendering Forecast Unobserved Node States (FUNS) a critical yet formidable challenge. Conventional models rely on historical observations and typically falter when encountering nodes without prior records. To address this, we redefine the problem as a conditional generation task on spatio-temporal graphs and propose GenST, a framework that introduces Large Language Models (LLMs) as a semantic bridge, leveraging a pre-trained LLM fine-tuned to extract rich semantic features from node descriptions, such as functional zones and road network structures, to compensate for missing spatio-temporal signals. Specifically, we design a two-stage generative architecture: a Spatio-Temporal VAE first compresses spatio-temporal dynamics into a latent space, followed by a Generative Transformer (GenT) that reconstructs the future states of unobserved nodes from noise, guided by multi-modal conditions including semantics, geographic coordinates, and neighborhood contexts. Experiments on six traffic and two non-traffic datasets show GenST significantly outperforms existing baselines in zero-shot prediction tasks, demonstrating the practical potential of semantic-guided generation for mitigating spatio-temporal data sparsity.

111. 【2610.08814】Route-Verify-Vote: Procedure-Conditioned Self-Consistency for Mixed-Domain Reasoning

链接:https://arxiv.org/abs/2610.08814

作者:Xinchen Xiao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Compositional generalization remains, generalization remains challenging, Compositional generalization, combine familiar reasoning, familiar reasoning operations

备注: 12 pages, 8 figures, and 8 tables. Accepted for oral presentation at CCL26-Eval

点击查看摘要

Abstract:Compositional generalization remains challenging when language models must combine familiar reasoning operations in unfamiliar ways. The Scenario-Based Commonsense Reasoning Evaluation (SCoRE) 2026 tests this ability on three mixed domains absent from training and requires models to identify the complete set of correct options for each question. We introduce Route-Verify-Vote (RVV), a framework for procedure-conditioned self-consistency that uses language models without parameter updates. Route uses the provided domain label to select a reasoning procedure that guides the model in representing and applying the relevant constraints. Verify prompts the model to assess each option against those constraints. Vote aggregates complete answer sets and allocates additional samples to questions with a small vote-count margin between the two most frequent sets. Samples for each question follow the same domain-specific procedure. On the official test set, voting over 16 sampled answer sets per question achieves an exact-set accuracy of 74.6%. Adaptive RVV reaches 77.3%, and combining models on selected domain routes raises accuracy to 79.4%. The final system ranked second among participating systems. These results support domain-specific reasoning procedures and answer-set disagreement as useful tools for allocating inference-time computation in mixed-domain reasoning.

Comments:
12 pages, 8 figures, and 8 tables. Accepted for oral presentation at CCL26-Eval

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2610.08814 [cs.AI]

(or
arXiv:2610.08814v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2610.08814

Focus to learn more

              arXiv-issued DOI via DataCite</p>
112. 【2610.08794】okka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages

链接:https://arxiv.org/abs/2610.08794

作者:Ben Gubler

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:broad comparative evaluation, Large language models, standardized multi-metric framework, multi-metric framework exists, language models rely

备注: 5 pages, 5 figures. Code and data: [this https URL](https://github.com/bgub/tokka-bench) . Interactive dashboard: [this https URL](https://tokka-bench.streamlit.app/)

点击查看摘要

Abstract:Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.

113. 【2610.07132】CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

链接:https://arxiv.org/abs/2610.07132

作者:Berke Arda,Ahmetcan Yavuz,Paul Gerry,Sebastian Lobentanzer,Nobin Sarwar,Joan Giner-Miguelez,Kongtao Chen,Luyao Zhang,Mrinmaya Sachan,Mubashara Akhtar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:accompanying dataset documentation, machine-readable dataset metadata, requires careful reading, dataset documentation, machine-readable dataset

备注: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: [this https URL](https://berkearda.github.io/croissantminer/)

点击查看摘要

Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

114. 【2610.09541】Certified by Abstention: Distribution-Free Guarantees for Chain-of-Thought Verifiers at Small Calibration Budgets

链接:https://arxiv.org/abs/2610.09541

作者:Arjun Balaji

类目:Machine Learning (stat.ML); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:compared by AUC, correct are compared, deploying one requires, requires a threshold, certificate

备注: 22 pages, 7 figures, 12 tables. Under submission at AISTATS 2027

点击查看摘要

Abstract:Signals that predict whether a chain-of-thought (CoT) trace is correct are compared by AUC, but deploying one requires a threshold with a guarantee. We ask what distribution-free selective guarantees deliver for CoT verifiers at realistic calibration budgets of tens to a few hundred labelled problems, using seven open models, five verifier signals and 37,000 graded traces. The central observation is validity by abstention: an $(\alpha,\delta)$-valid procedure that issues a certificate with probability $P_{\rm fire}$ bounds the failure probability of an issued certificate only by $\delta/P_{\rm fire}$, so a certificate that rarely fires can be valid and wrong every time it is used. In a simulation with known risk the standard certificate fails in at most 0.3% of calibration draws but in up to 69% of those in which it fires. A certification floor and a lattice condition for Benjamini-Hochberg conformal selection explain why certificates abstain at these budgets, and the data bear them out: the standard certificate returns nothing or a large accepted set, and an unreadable residual-stream probe buys two to three times the coverage of the readable signals, an edge a cross-fitted reconstruction cannot recover linearly from the readable features. We then give a floor-started fixed-sequence certificate, valid without monotonicity assumptions, that covers more than the Bonferroni certificate on every model-signal pair and raises coverage at the non-vacuous target $0.75\pi_0$ from 0.05 to 0.16, although the floor keeps absolute coverage small. Finally, a certificate cannot see what matters after deployment: under benchmark shift the error among accepted traces tracks the new task's base error, and under best-of-$n$ selection against the verifier it rises past the target while the empirical failure frequency stays below $\delta$, because abstention absorbs the failures.

115. 【2610.09486】Mitigating Accent-Language Confusion in Self-Supervised Speech Representations for Language Identification

链接:https://arxiv.org/abs/2610.09486

作者:Minu Kim,Jihwan Lee,David R. Mortensen,Shrikanth Narayanan

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

关键词:Spoken language identification, Spoken language, aims to recognize, Spoken, LID

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Spoken language identification (LID) aims to recognize the target language regardless of accent. In practice, however, LID models fine-tuned from self-supervised speech representations frequently confuse accents with languages, misclassifying non-native (L2) speech as the speaker's first language (L1). We show that non-native speech representations lie between native target-language and native L1 poles, causing systematic misclassification. To address this, we introduce a geometric projection that estimates an L1-bias direction solely from native speech and removes it before the frozen LID head. Across five MMS-LID models and non-native corpora, this projection substantially improves target language identification for L2-accented speech while preserving predictions for native speech. These results show that accent-induced L1 bias can be corrected directly within the representation space without L2 training data or model adaptation.

116. 【2610.09467】Boundary-Free Contextual Biasing: Depth-Adaptive Gating and Reading-Space Matching for Unsegmented Languages

链接:https://arxiv.org/abs/2610.09467

作者:Muhammad Huzaifah,Yu Pan,Zachary Yeo,Ningjie Bai,Guangzhao Yang

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

关键词:Contextual biasing supplies, existing methods rely, supplies an ASR, ASR system, Contextual biasing

备注:

点击查看摘要

Abstract:Contextual biasing supplies an ASR system with a list of expected words at inference time, but existing methods rely on word boundaries that Japanese and Chinese do not provide. We present a boundary-free biasing decoder for frozen public CTC models, built on a character-level Aho-Corasick automaton, with no training and no second pass. Two evidence-based mechanisms replace the boundary: a depth-adaptive gate that sets how hard to push from match depth, and reading-space matching for when the audio is right but the characters are wrong. On Aishell-1 NE's hard R1 subset we reach 66.5% recall, above the trained CLAS baseline (64%), transferring to WenetSpeech and to a second architecture without retuning. We release the first open Japanese contextual-biasing benchmark, where biasing lifts rare-word recall by 25 points at precision above 97%, and still by 19 and 22 points against 1,000-word lists.

117. 【2610.08994】Phoneme-Guided Initialization for LLM-based Speech Recognition

链接:https://arxiv.org/abs/2610.08994

作者:Ryo Magoshi,Shinsuke Sakai,Tatsuya Kawahara

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)

关键词:automatic speech recognition, sufficient paired speech-text, paired speech-text data, Speech large language, speech LLMs

备注: Accepted at IEEE SLT 2026

点击查看摘要

Abstract:Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.

118. 【2610.08831】Is Word Error Rate Enough? Rethinking Privacy Evaluation in Speech with Entity-Aware Metrics

链接:https://arxiv.org/abs/2610.08831

作者:Anjana Rajasekhar,Jule Pohlhausen,Nayana Jacob Alappattu,Anna Leschanowsky

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)

关键词:smart devices continues, growing privacy concerns, content raises growing, capture sensitive speech, raises growing privacy

备注:

点击查看摘要

Abstract:As the use of smart devices continues to increase, their potential to capture sensitive speech content raises growing privacy concerns. It is therefore critical to develop techniques that prevent information leakage while preserving the utility of the audio, and evaluation metrics that accurately quantify the level of privacy without overestimating it. In this work, we evaluate the effectiveness of two obfuscation techniques in protecting speech content, with particular emphasis on named entities, by adapting entity-aware privacy metrics from the Natural Language Processing field to the speech privacy domain. Further, we investigate several attack scenarios and show that fine-tuning on entity-rich data improves attack performance for some entity categories but not others. Finally, we provide guidance on metric selection based on whether the obfuscation method preserves temporal alignment.

信息检索

1. 【2610.10483】wo-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion

链接:https://arxiv.org/abs/2610.10483

作者:Walid Bendada,Guillaume Salha-Galvan

类目:Machine Learning (cs.LG); Information Retrieval (cs.IR); Machine Learning (stat.ML)

关键词:items makes exact, makes exact sampling, exact sampling impractical, machine learning, impractical at scale

备注: NeurIPS 2026

点击查看摘要

Abstract:Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.

2. 【2610.10441】CrossWeave: Bridging Perspectives Across Online Communities with a Dual-Pane Design

链接:https://arxiv.org/abs/2610.10441

作者:Fei Fang,Reva Hirave,William Jurayj,Yuqi Li,Brian Lu,Tarik Metin,Tsugunobu Miyake,Kateryna Morhun,Yash Permalla,Kenan Rustamov,Allen Shen,Haojun Shi,Prabhav Singh,Xiheng Tom Wang,Kevin Xu,Qingcheng Zeng,Jiayi Zhang,Daniel Khashabi,Andrew Perrin,Tiziano Piccardi,Ziang Xiao,Jason Eisner

类目:ocial and Information Networks (cs.SI); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)

关键词:typically display conversations, systems typically display, familiar contributors, predictable and one-sided, media systems typically

备注: CSCW 2026 + small improvements

点击查看摘要

Abstract:Social media systems typically display conversations among already familiar contributors, which can be predictable and one-sided. In civic discourse, this design narrows discussion, reinforces divides, and distorts the perception of public opinion. To encourage cross-community engagement, we present CrossWeave, an AI-powered bridging system that augments the standard social media feed. As the user reads a post, CrossWeave surfaces diverse relevant posts from other threads in a side pane and highlights the connections. Users are invited to venture out of their echo chamber, explore a broader range of views and arguments, and ``click across'' to engage with their authors. When they do, CrossWeave facilitates constructive posting, not only by showcasing relevant past content but also by simulating possible reactions as the user drafts a post.

3. 【2610.10170】Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora

链接:https://arxiv.org/abs/2610.10170

作者:Andrey Kuehlkamp,Priscila Correa Saboia Moreira,Samuel Rund

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Retrieval-augmented generation systems, generation systems increasingly, systems increasingly rely, LLM-generated chunk contexts, Retrieval-augmented generation

备注: Initial draft,

点击查看摘要

Abstract:Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuffled across documents. We score retrieval with a coverage-aware nDCG and test four pre-registered contrasts via document-clustered bootstrap with Holm correction, on two distant corpora: 200 Wikipedia Featured Articles (951 queries) and 1,585 QASPER papers (4,303 questions). Organization helps, and the cause is content, not tokens: structure-aligned chunks with real heading paths beat contextualized fixed windows (+0.022 / +0.012 cov-nDCG@10) and the placebo (+0.010 / +0.016). Naive two-stage hierarchical retrieval hurts (-0.033 / -0.015), traceable to first-stage section recall. Gold structure beats LLM-induced structure on Wikipedia but not on QASPER. Effects are small ($dz$ 0.06-0.11) but Holm-significant and consistent across corpora.

4. 【2610.10124】raining with Missed Targets in Generative Recommendation: Separating Supervision from Probability Competition

链接:https://arxiv.org/abs/2610.10124

作者:Xuesi Wang,Yangbin Shi,Xiaolin Zheng

类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Generative recommenders return, Generative recommenders, omit observed targets, limited candidate set, recommenders return

备注: 12 pages, 4 figures, 8 tables

点击查看摘要

Abstract:Generative recommenders return a limited candidate set and may omit observed targets before reranking. A training strategy appends these missed targets to reranker training lists, although inference still ranks only original candidates. This operation simultaneously changes retrieved-target weight, adds supervision over appended targets, and makes the two groups compete for probability. An append/no-append comparison therefore cannot explain changes in returned-item rankings. We construct three matched losses that hold retrieved-target weight fixed while introducing appended-target supervision and group competition separately. The intermediate loss trains within both groups but normalizes them separately, preventing training-only targets from competing with inference candidates. Experiments with a released OneRec model and locally trained Amazon generators show that this competition can harm returned-item ranking. In four prespecified Amazon Video Games comparisons, removing it improved full-target normalized discounted cumulative gain (FT-NDCG) by 7.8--22.2\%; 95\% intervals over users and three of four intervals over training runs excluded zero. A conservative development-set rule selected appended-target training for two of three generators in one held-out category and rejected it for all three in another, avoiding a 1.7\% loss. Candidate completion should therefore be evaluated for each generator rather than applied automatically.

5. 【2610.10091】ExperienceIndex: Artifact-Grounded Memory

链接:https://arxiv.org/abs/2610.10091

作者:Peter Baile Chen,Geoffrey X. Yu,Xinming Liu,Samuel Madden,Dan Roth,Jacob Andreas,Doug Downey,Michael Cafarella

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:Knowledge-intensive tasks require, tasks require answering, court cases, Knowledge-intensive tasks, scientific literature

备注:

点击查看摘要

Abstract:Knowledge-intensive tasks require answering many questions by reasoning about a shared corpus of artifacts (e.g., court cases, or scientific literature). As humans interact with these corpora, they naturally accumulate experiential knowledge about artifacts, enabling them to quickly identify the complete set of relevant artifacts for each new task. However, existing AI agents lack appropriate memory solutions to build or reuse such artifact-grounded experience, leading to lower answer quality and higher online cost. Existing memory solutions extract and reuse information from prior task-solving traces, but they primarily focus on user preferences, factual attributes, or abstract reasoning patterns rather than persistent artifact-specific knowledge. We introduce ExperienceIndex, a novel experience layer for AI agents that captures and reuses knowledge about artifacts based on prior reasoning traces. ExperienceIndex stores two complementary forms of experience: (i) single-artifact experiences that summarize an artifact's contribution to prior tasks and (ii) artifact-pair experiences that encode structural relationships discovered during past reasoning. Integrated as lightweight middleware, ExperienceIndex uses an experience retrieval mechanism to guide agents toward the complete set of relevant artifacts for new tasks, improving both answer quality and efficiency. Across diverse corpora and agentic solutions with different search frameworks, ExperienceIndex delivers consistent gains, raising answer quality by up to 11.0 points and reducing online dollar cost by up to 50.5%. We further demonstrate two benefits: (i) cross-task generalization, where experiences accumulated from text-to-SQL tasks transfer to factoid QA tasks over the same artifact corpus, and (ii) teacher-student learning, where experiences from a stronger model enable a weaker model to reach comparable performance.

6. 【2610.09920】Inverting Multi-Vector Visual Document Indices

链接:https://arxiv.org/abs/2610.09920

作者:Zhuchenyang Liu,Yao Zhang,Yu Xiao

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Prevailing multi-vector visual, vector databases run, thousand patch vectors, Prevailing multi-vector, page

备注: 30 pages. Under review

点击查看摘要

Abstract:Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.

7. 【2610.09820】he Impact of Backbone Evolution on LLM-Based Relevance Assessments

链接:https://arxiv.org/abs/2610.09820

作者:Chuting Yu,Guido Zuccon,Teerapong Leelanupab

类目:Information Retrieval (cs.IR)

关键词:offering stronger capabilities, models offering stronger, evolving rapidly, stronger capabilities, offering stronger

备注: 12 pages main content

点击查看摘要

Abstract:LLMs are evolving rapidly, with newer models offering stronger capabilities. This suggests that in LLM-based relevance judging, more capable models will achieve higher agreement with human judgements under the same prompt. We challenge this understanding by investigating the behavior of LLM-based relevance judges under backbone evolution. Keeping the prompts fixed, we evaluate a representative single-prompt (UMBRELA) and a rubric-based prompt (EXAM) across sequential model versions of commercial (Gemini, GPT) and open-weight (Qwen, Llama) models. Overall, we find no consistent evidence that newer versions lead to better relevance judges. Crucially, similar or improved aggregate performance does not imply judgment stability: correct judgements made by an earlier version of an LLM backbone are not necessarily preserved by later versions. We investigate the potential drivers of these regressions. Our findings caution against the assumption that judging prompts designed and validated for one backbone version will perform equivalently or better when the model is updated, even within the same family.

8. 【2610.09742】SoccerNet-FoulRet: Retrieving Semantically Similar Soccer Foul Videos

链接:https://arxiv.org/abs/2610.09742

作者:Jacobus Arthur,Ahmad Sait,Batool Hani,Merey Ramazanova,Jan Held,Marc Van Droogenbroeck,Bernard Ghanem,Anthony Cioppa,Silvio Giancola

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:similar past cases, professional soccer remain, soccer remain inconsistent, decisions in professional, professional soccer

备注: ACCV 2026

点击查看摘要

Abstract:Refereeing decisions in professional soccer remain inconsistent because referees cannot easily compare a contentious foul against similar past cases. We cast this as a retrieval problem and introduce SoccerNet-FoulRet, the first benchmark for semantic foul retrieval. Given a query foul, the task is to retrieve past fouls judged to be relevant precedents, regardless of camera angle, teams, or appearance. This differs from prior video-to-video retrieval, which matches clips by visual similarity or a shared event. Here, relevance is defined by refereeing interpretation. We build the benchmark from the SoccerNet-MVFoul dataset and evaluate retrieval ability of zero-shot video and vision-language embedders together with a task-specific fine-tuned baseline on 693 human-verified queries and category-relevance labels. Semantic foul retrieval remains challenging. The strongest zero-shot model achieves under 5% HitRate@10 on human-verified precedents, while category-supervised fine-tuning improves category relevance but transfers only modestly to precedent retrieval. We release SoccerNet-FoulRet to establish semantic foul retrieval as an open problem: this https URL.

9. 【2610.09724】owards Explaining Query Expansion Performance in Information Retrieval

链接:https://arxiv.org/abs/2610.09724

作者:Sourav Saha,Aditya Dutta,Soumajit Pramanik,Mandar Mitra

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:vocabulary mismatch problem, Query Expansion, Information Retrieval, techniques have long, mismatch problem

备注:

点击查看摘要

Abstract:Query Expansion (QE) techniques have long been widely used in Information Retrieval (IR) to address the vocabulary mismatch problem. They remain relevant in modern retrieval systems, including those based on large language models (LLMs). However, no single QE method consistently outperforms others across all queries. This work seeks to explain the variation in QE performance through two complementary perspectives. The first is the concept of an Ideal Expanded Query (IEQ)--a hypothetical query that maximizes retrieval effectiveness with a downstream BM25 retrieval model. The second is a separability perspective, which quantifies how distinctly relevant and non-relevant documents are scored for a given expanded query using Cohen's (d). We develop a separability measure and practical formulations to approximate the IEQ and investigate how these factors relate to retrieval effectiveness. Extensive experiments on the TREC Robust collection, TREC DL 2019-2022 passage collections, and TREC DL 2019-2020 document collections reveal several interesting patterns. In particular, we find that expanded queries that are closer to the ideal expanded query tend to achieve higher retrieval effectiveness. We further show that the separability of relevant and non-relevant documents provides a complementary perspective for understanding QE performance.

10. 【2610.09412】Finding the Right Balance: Relevance and Diversity in LLM Retrieval

链接:https://arxiv.org/abs/2610.09412

作者:Guillaume Brouillette(1),Faustin Kagabo(1),Usef Faghihi(1),Nadia Ghazzali(1) ((1) Université du Québec à Trois-Rivières, Trois-Rivières, Canada)

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:prior studies disagree, retrieval-augmented generation, prior studies, studies disagree, RAG

备注: 36 pages, 8 figures, 13 tables. Code and results: [this https URL](https://github.com/GuillaumeBrouillette/finding-the-right-balance)

点击查看摘要

Abstract:Retrieval diversification is widely available in retrieval-augmented generation (RAG) frameworks, yet prior studies disagree on whether it improves retrieval and answer quality. We show that its effectiveness varies primarily with candidate-pool redundancy, in a pattern consistent with the number of distinct evidence pieces a query requires. Using controlled near-duplicate injection and production-style overlapping chunking, we find that diversification harms relevance, evidence coverage and answer quality on clean pools, but becomes beneficial on multi-evidence tasks when redundancy causes nearest-neighbor retrieval to select repeated passages. We therefore introduce a query-adaptive rule that diversifies only when the effective number of distinct documents in the nearest-neighbor top-$k$ selection falls below the query's evidence requirement. Computed from existing embeddings, the rule captures most of the achievable gain, transfers across datasets and encoders and automatically reduces to nearest-neighbor retrieval for single-evidence queries. We also introduce RNG-Score, a geometric reranker with an exact nearest-neighbor fallback whose margin indicates duplicate structure. Overall, we conclude that diversification should be used selectively, based on observable redundancy and evidence requirements.

11. 【2610.09361】From Chunks to Functional Evidence: Function-Aware Retrieval for EDA Documentation QA

链接:https://arxiv.org/abs/2610.09361

作者:Xiaotian Qiu,Kairui Liu,Shi Chenyi,Jinyuan Deng,Qi Sun,Cheng Zhuo

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Retrieval-Augmented Generation, Electronic Design Automation, Generation, ground answers, Design Automation

备注: 10 pages, 2 figures, including appendices

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) is widely used to ground answers in documents. For complex technical documentation, however, the primary bottleneck is often not model reasoning but a mismatch between a query and the way knowledge is organized for retrieval. This mismatch is pronounced in Electronic Design Automation (EDA) documentation, where the information needed for an answer is scattered across heterogeneous yet tightly coupled artifacts. We therefore redesign the basic retrieval unit of RAG. Instead of operating on isolated chunks or binary relations, we collect typed artifacts into EDA functional units. Each unit is recorded as a hyperedge with links to its source chunks. We then train an encoder to align queries with functional units and combine unit retrieval with direct chunk retrieval. After mapping the selected units back to their sources, a unified reranker chooses the evidence given to the generator. On the newly constructed EDADocEval-QA dataset, our method improves ROUGE-L by 37.1% over Chunk RAG and 55.6% over the strongest graph baseline. On the public ORD-MMBench benchmark, it improves ROUGE-L by 30.0% over the strongest baseline. These results support function-aware evidence organization in the evaluated EDA documentation settings.

12. 【2610.09262】Reading Position Is the Baseline to Beat: A Time-Ordered Evaluation of Personalised Highlight Prediction

链接:https://arxiv.org/abs/2610.09262

作者:Kazuki Nakayashiki,Keisuke Watanabe

类目:Information Retrieval (cs.IR); Human-Computer Interaction (cs.HC)

关键词:cheapest personal signal, cheapest personal, personal signal, reader, popularity

备注: 13 pages, 1 figure, 5 tables. Ancillary files include the specifications, the results write-ups, the analysis scripts, and the aggregate artifacts every reported number is generated from

点击查看摘要

Abstract:A reader's first highlights on a page are the cheapest personal signal a reading product has. The natural plan is to suggest what similar earlier readers marked, and to judge the result against popularity. We argue that the baseline to beat is reading position. In a time-ordered evaluation on one social highlighting platform (7,343 reader-page pairs on 1,511 pages after one highlight), ranking the sentences just below a reader's first highlight, with no other reader's data, puts the next highlight in the top five 47% of the time, against 26% for popularity and 29% for the better of two similarity methods. The baseline depends on the target: over all later highlights that ranking loses to popularity, while popularity discounted by distance from the latest highlight, at the scale with the best average precision of three tried, beats popularity and both similarity methods on both targets. In a comparison specified in advance, neither similarity method shows a gain over popularity in average precision over all later highlights, from one to five highlights, and a gain of +0.01 is excluded. Nor would a gain by itself show that a method has found a reader's preferences: synthetic readers who share one set of preferences produce one, and an evaluation out of time order shows a method where the reader went. The position results are exploratory and unconfirmed. Personalisation inside a document should be evaluated in time order and against reading position.

13. 【2610.09227】Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

链接:https://arxiv.org/abs/2610.09227

作者:Hyojung Han,Jongmin Kim,Seung-Hun Jeon

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Mixed-precision post-training quantization, costs when quantized, per-module sensitivity signal, post-training quantization, deployments rarely

备注: 26 pages, 22 tables, 4 figures

点击查看摘要

Abstract:Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.

14. 【2610.09177】What Transfers from a VLM Teacher? Comparing Supervision Signals for Visual Document Retrieval

链接:https://arxiv.org/abs/2610.09177

作者:Saba Sturua,Han Xiao

类目:Information Retrieval (cs.IR)

关键词:Visual document retrievers, Visual document, teacher, document retrievers, pages presumed irrelevant

备注: 27 pages, 1 figure, 17 tables

点击查看摘要

Abstract:Visual document retrievers are trained contrastively: each query is matched to one page labelled relevant - the positive - and pushed away from negatives, pages presumed irrelevant. Recent methods distil a vision-language model (VLM) teacher into the retriever by enriching that positive, transferring the teacher's attention over it or a description of it. We ask whether the teacher is better spent on the other side, judging the candidates the retriever mines as negatives, which the label says nothing about. With student, data, optimizer and evaluation fixed, teacher-judged hard negatives and score distillation raise ViDoRe v2 nDCG@5 from 55.2 to 62.6 and 63.0; description alignment, as adapted here, gains 2.6 points and attention grounding nothing measurable. Against teacher-free rules that select four candidates from the same mined pool at identical training compute, the best of which is the positive-aware threshold current systems use, the teacher's judgement adds 4.1 points on v2 and 1.7 on v3. This is consistent with how incomplete the labels are. Annotators judge about two of a query's four top-ranked mined candidates relevant, none of them labelled, so training pushes the retriever away from relevant pages treated as negatives. What reaches the student is coarse: under a greedily decoded 0-100 rating prompt, 82% of the teacher's ratings come back at one end of the scale or the other, and a relevant/irrelevant partition keeps most of the distillation gain. A ten-annotator audit places the teacher within the range of variation among human annotators, and finds it reliable where a query has a single determinate answer. We release the code, the teacher's 3.3M judgements and page descriptions, the mined pools, the human audit and the trained adapters at this https URL.

15. 【2610.09041】Building Navigable Graphs Without Search in Three Composable Stages

链接:https://arxiv.org/abs/2610.09041

作者:Édgar Chávez

类目:Databases (cs.DB); Information Retrieval (cs.IR)

关键词:Navigable graphs, searching for neighbors, inside each part, point, Navigable

备注: 26 pages. Code: [this http URL](http://github.com/zevahcle/graft-ann) (branch fgraft); experiments, logs and patches: [this http URL](http://github.com/zevahcle/fgraft-experiments)

点击查看摘要

Abstract:Navigable graphs can be built without searching for neighbors: partition the data, evaluate every pair inside each part, and select each point's edges from the candidates. We give such a construction in three separable stages and show that the middle one decides the quality. The pool is any partition with a few memberships per point. The ending turns a point's candidates into out-edges; ours keeps a bounded heap, prunes by occlusion with a per-corpus slack, and appends reverse edges, re-pruning only where a list overflows. The spine is any edge set, exempt from the prune, that keeps the graph reachable from its entry; ours, half-space-proximal edges over a random sample, routes monotonically to every sampled point and replaces a spanning tree at 1/10 to 1/500 of its cost. The ending composes with any partitioner: on PiPNN's own candidate pool it beats PiPNN's ending on each of six corpora from $10^6$ to $10^8$ points, by 3 to 14% in distance evaluations at equal recall, and with 60 to 120 memberships per point the composed build matches or beats a full dense construction at k=10 and k=100 on all six, in 0.5 to 0.9 of its build time, deterministically. The analysis explains why. Once a pool is localised its quality is set by the data: every pool built on GIST lands within 4% of the exact-kNN ceiling, and the pairs a block cover misses are predicted, point by point, by the local clustering of the kNN graph, whose zero-clustering tail sets the memberships a corpus needs and grows with n. All code, patches and logs are public.

16. 【2610.09039】From High Recall to High Utility: Dataset-Adaptive Post-Processing of LLM-Generated Customer Intents

链接:https://arxiv.org/abs/2610.09039

作者:Mahesh Viswanathan,Joan Rossello,Leticia Fernandes,Paul Mutawe

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Large language models, Large language, uneven in granularity, semantically overlapping, models can extract

备注:

点击查看摘要

Abstract:Large language models can extract useful signals from heterogeneous enterprise data, but high-recall extraction often produces outputs that are duplicated, uneven in granularity, semantically overlapping, or too numerous for downstream systems and human reviewers to use effectively. We present a dataset-adaptive post-processing architecture developed for Customer Intent Extraction (CIE), where unstructured customer language is transformed into stable, traceable intent units. The approach separates recall-oriented extraction from utility-oriented reduction. Source-specific preprocessing first isolates evidence from multimodal plans, sparse operational records, and structured opportunity data. Candidate intents are then standardized and deduplicated, optionally enriched with metadata for embedding computation, represented in a shared semantic vector space, and grouped using a clustering strategy selected according to the candidate set's characteristics. Cluster-level keywords provide an explainability layer, while singleton reassignment requires agreement between embedding and keyword similarity. Finally, constrained language-model aggregation produces one concise intent per cluster without introducing unsupported concepts, and the resulting unit retains provenance, clustering, embedding, and generation metadata. This treats post-processing not as cosmetic cleanup, but as a semantic reduction layer converting high-recall LLM outputs into reusable enterprise intelligence. We also describe two downstream applications: Machine-Generated Intents, which infer likely objectives for customers lacking direct evidence from peer customers with similar profiles, and intent-guided semantic retrieval and mapping, which uses the stable intent as a query against a downstream decision space, illustrated here by mapping customer intents to business outcomes.

17. 【2610.09026】BEACON-SP: Ontology-Grounded GraphRAG Framework for Clinical Suicide Risk Assessment

链接:https://arxiv.org/abs/2610.09026

作者:Kemal Davaslioglu,Nathan Conger,Sastry Kompella,Yalin E. Sagduyu,Nathaniel D. Bastian

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)

关键词:effective assessment requires, assessment requires integrating, requires integrating heterogeneous, behavioral health settings, Suicide Social Determinants

备注:

点击查看摘要

Abstract:We present BEACON-SP, an ontology-grounded Graph Retrieval-Augmented Generation (GraphRAG) framework for clinician-facing decision support in behavioral health settings such as suicide prevention, where effective assessment requires integrating heterogeneous clinical, behavioral, social, and temporal evidence. BEACON-SP combines patient knowledge graphs with ontology-guided retrieval to support multi-hop reasoning across diagnoses, medications, risk and protective factors, life events, and temporal relationships. The framework is enabled by a comprehensive suicide prevention ontology that integrates the Three-Step Theory, the Integrated Motivational-Volitional Model, and the Suicide Social Determinants of Health Ontology into a unified representation of patient risk factors. We construct ontology-grounded patient knowledge graphs and evaluate BEACON-SP for clinician-facing question answering. Compared with a vector-based retrieval-augmented generation (RAG) baseline on a 1,500-query benchmark spanning 15 clinical categories and 100 patients, BEACON-SP improves completeness, clinical relevance, and evidence grounding under a corrected comparative evaluation protocol, with a small gain on factual accuracy. In paired criterion-level comparisons, GraphRAG is preferred in 76.4% of cases. These results demonstrate the potential of ontology-guided GraphRAG to provide structured, contextualized patient evidence for clinical decision support.

18. 【2610.08894】rustworthy Domain-Specific AI for Structured Knowledge Retrieval and Reasoning

链接:https://arxiv.org/abs/2610.08894

作者:Ryan C. Barron

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Non-negative Matrix Factorization, transforming unstructured, presents a scalable, scalable architecture, architecture for transforming

备注:

点击查看摘要

Abstract:This dissertation presents a scalable architecture for transforming unstructured, domain-specific text into structured knowledge for retrieval and reasoning. It integrates semi-automatic corpus curation, semantic structuring, retrieval, and inference into an interpretable pipeline. The research introduces Binary Bleed, an adapted binary search method that reduces low-rank search complexity for Non-negative Matrix Factorization (NMF), and Hierarchical NMF with automatic latent feature selection (HNMFk), a depth-adaptive topic modeling method that produces interpretable taxonomies guided by subject matter experts. These representations populate a typed Knowledge Graph and a semantically aligned Vector Store containing extracted latent features, synchronized through an event-driven substrate. Tensor-Structured Retrieval-Augmented Generation (T-SRAG) dynamically routes queries across retrieval paths. Contrastive alignment maps document and query embeddings to hierarchical topic structures to improve semantic fidelity and reduce hallucinations. Beyond retrieval, tensor-based link prediction identifies and completes missing links in the Knowledge Graph, supporting inference grounded in citation structure. Applications across cybersecurity, law, materials science, and healthcare demonstrate improvements in retrieval precision, early trend detection, hypothesis generation, and hallucination mitigation. The dissertation provides a deployable, modular foundation for trustworthy, domain-specific AI systems that retrieve and reason over structured knowledge.

Subjects:

Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2610.08894 [cs.IR]

(or
arXiv:2610.08894v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2610.08894

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Ryan Barron [view email] [v1]
Tue, 6 Oct 2026 16:35:41 UTC (30,346 KB)

19. 【2610.07132】CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

链接:https://arxiv.org/abs/2610.07132

作者:Berke Arda,Ahmetcan Yavuz,Paul Gerry,Sebastian Lobentanzer,Nobin Sarwar,Joan Giner-Miguelez,Kongtao Chen,Luyao Zhang,Mrinmaya Sachan,Mubashara Akhtar

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:accompanying dataset documentation, machine-readable dataset metadata, requires careful reading, dataset documentation, machine-readable dataset

备注: Accepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: [this https URL](https://berkearda.github.io/croissantminer/)

点击查看摘要

Abstract:Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

计算机视觉

1. 【2610.10539】ris3D: 3D Scene Generation With Objects That Fit Together

链接:https://arxiv.org/abs/2610.10539

作者:Jaeyeong Kim,Jinhyuk Jang,Jongmin Lee,Kyehong Park,Seungryong Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:framework for single-image, generative framework, scene reconstruction, objects, providing limited guidance

备注: Project page: [this https URL](https://cvlab-kaist.github.io/Tetris3D/)

点击查看摘要

Abstract:We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and their physical relationships, guiding its shape and pose to remain geometrically and physically plausible within the scene. Moreover, we introduce ComOb, a physics simulation-based dataset of 1.2M scenes featuring physical interactions across diverse object categories, with per-object meshes and pairwise physical relation annotations. Comprehensive experiments on synthetic and realworld scenes show that Tetris3D recovers coherent object shapes and poses even when interacting regions are occluded, and achieves state-of-the-art performance in both generation quality and physical stability.

2. 【2610.10538】Never Look Back: Understanding Persistence in 3D Object Memory from Egocentric Videos

链接:https://arxiv.org/abs/2610.10538

作者:Shravan Chaudhari,William Paul,Suchi Saria,Rama Chellappa,Homanga Bharadhwaj

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:everyday tasks, world and carry, carry out everyday, Abstract, encounter objects

备注: Project page: [this https URL](https://ledger-3d.github.io) . Code: [this https URL](https://github.com/LEDGER-3D/LEDGER)

点击查看摘要

Abstract:As we move through the world and carry out everyday tasks, we encounter objects that may become relevant only later. We are capable of recalling where we left something or what was inside a container, even without knowing we would need it later. Here, we study how an embodied assistant can build a similar memory from egocentric videos, by observing a person's day-to-day activities. We present Ledger, a persistent 3D object memory that combines object locations, their histories, and contextual descriptions. It associates observations across the recording and retains objects after they leave the view, including those the person never touches. It clusters each object's observations by resting locations and records a move only after repeated evidence, reducing the effect of localization noise. Short descriptions preserve details such as an object's contents or supporting surface. It saves these records to later answer spatial questions without having to access the original images or video. Our memory raises HD-EPIC accuracy from 29.7% to 42.6%, UCS-Bench accuracy from 33.8% to 38.5% and localizes Ego4D objects with a 0.99 m median error on returned predictions. Our analyses identify complementary roles for temporal persistence, contextual descriptions, and retrieval. Our study on 100 stitched streams of multiple scenes each further exposes failures in both retrieval and construction. Per-scene construction partially recovers the performance lost across scene changes compared to that of single scene streams.

3. 【2610.10528】Long-WAM: Scaling the Context of World-Action Models

链接:https://arxiv.org/abs/2610.10528

作者:Wei Huang,Bohan Zhang,Chenzhi Liu,Isabella Liu,Shuai Yang,Weian Mao,Luozhou Wang,Yicheng Xiao,Weifeng Lin,Qixin Hu,Bryan Chu,Sifei Liu,Linxi Fan,Xiaojuan Qi,Song Han,Yukang Chen

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:robot control demands, real-time control constraints, demands enough visual, infer motion, Real-time robot control

备注:

点击查看摘要

Abstract:Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.

4. 【2610.10524】GRACE: Generation-aware latent compression for efficient video generation

链接:https://arxiv.org/abs/2610.10524

作者:Jiyoung Kim,Paul Hyunbin Cho,Jisu Nam,Donghoon Lee,Hyunsung Go,Yeonkyeong Lee,Hansaem Kim,Seungryong Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:video diffusion models, accelerate video diffusion, Diffusion Transformer, Highly compressed video, diffusion models

备注: Project page : [this https URL](https://cvlab-kaist.github.io/GRACE/) , 43 pages, 24 figures

点击查看摘要

Abstract:Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.

5. 【2610.10512】Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

链接:https://arxiv.org/abs/2610.10512

作者:Chen Xu,Yunqi Li,Binbin Huang,Brent Yi,Shenghua Gao,Yi Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:incomplete visual observations, Recovering faithful, make frame-wise pose, frame-wise pose estimates, pose estimates unreliable

备注: 20 pages, 6 figures

点击查看摘要

Abstract:Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.

6. 【2610.10497】QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

链接:https://arxiv.org/abs/2610.10497

作者:Yucheng Mao,Zeyuan Chen,Xiaojun Shan,Xiang Zhang,Divyansh Srivastava,Bingnan Li,Zhuowen Tu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:framework for visual, visual tokenization, autoregressive image generation, image generation, introduce QuadTok

备注:

点击查看摘要

Abstract:We introduce QuadTok, a novel framework for visual tokenization and autoregressive image generation. Compared to traditional approaches using 2D grids or 1D token sequences, we propose a hierarchical quadtree structure, bridging the gap between 2D spatial binding and 1D sequence-level flexibility. The QuadTok tokenizer dynamically allocates representational capacity to visually intricate areas while leaving homogeneous regions at a coarse resolution. Compared with a fixed 256-token grid, our ImageNet-trained tokenizer saves approximately 10% of tokens on ImageNet and 9% when transferred zero-shot to the COCO dataset, while maintaining comparable reconstruction fidelity. Furthermore, the natural causality introduced by the tree structure seamlessly enables autoregressive image generation. Conditioned on a quadtree topology supplied before generation, our 947M GPT-style generative model achieves a 2.08 gFID on the ImageNet $256 \times 256$ benchmark. Additionally, leveraging the strong spatial correlation preserved by the quadtree structure, the QuadTok generator enables zero-shot spatially controlled image generation capabilities. Code: this https URL.

7. 【2610.10491】Insights from Autoresearch for Solar Panel Segmentation

链接:https://arxiv.org/abs/2610.10491

作者:Justinas Lekavicius,Kursat Komurcu,Valentas Gruzauskas,Linas Petkevicius

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:paper investigates AutoResearch, one-hour GPU budget, coding language model, language model edits, GPU budget

备注: Accepted at AutoML4EO 2026 (non-archival AutoML conference workshop). 4 pages + references. [this https URL](https://automl4eo.org/accepted-papers/)

点击查看摘要

Abstract:This paper investigates AutoResearch, a protocol in which a coding language model edits a training program under a one-hour GPU budget and retains a change only if validation IoU improves. The protocol is applied to photovoltaic panel segmentation on a frozen real-image split, with DeepLabV3--ResNet-50 held fixed. Three campaigns of 24 experiments, using Gemma~4 12B, Qwen3-8B all improve their one-hour baselines, but retained modifications do not transfer across hardware. The Qwen3-8B configuration, trained on real images only, reaches a test IoU of 0.836 versus 0.833 for the reference GAN-augmented schedule. Research repository this https URL.

8. 【2610.10479】Agentic RSR: Real-to-Sim-to-Real through Scene Reconstruction and Execution-Grounded Robot Policies

链接:https://arxiv.org/abs/2610.10479

作者:Yihan Li,Yating Feng,Shengjiu Sun,Jianing Chen,Hao Ren,Bowen Yang,Weisheng Xu,Qiwei Wu,Hui Cheng,Renjing Xu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:preserve task-relevant interactions, policies developed, real robot, Agentic RSR, preserve task-relevant

备注: 25 pages including appendices, 5 figures

点击查看摘要

Abstract:A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab $\Delta E_{76}$ is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.

9. 【2610.10473】Label-free cell counting and viability prediction with brightfield imaging and deep learning

链接:https://arxiv.org/abs/2610.10473

作者:Amir Reza Vazifeh,Christian Zeigler,Sornanathan Meyyappan,Richard Jeske,Jason W. Fleischer

类目:Computer Vision and Pattern Recognition (cs.CV); Cell Behavior (q-bio.CB)

关键词:cell culture systems, culture systems, Cell viability assessment, cells, core requirement

备注:

点击查看摘要

Abstract:Cell viability assessment is a core requirement in cell culture systems, with critical applications in biopharmaceutical manufacturing and drug development. Conventionally, it is measured by adding membrane-impermeable dyes to a sample (a process called staining), which allows compromised cell membranes to be distinguished from intact ones. However, staining has several limitations: (a) chemical agents can perturb normal cellular processes of the cells being measured, (b) it is often ambiguous to assign viability to individual cells whose membrane integrity is only partially compromised. (c) photobleaching can undermine measurement accuracy over time when using fluorescent stains, and (d) staining cannot be performed in situ or in real time. Here, we show that (1) stained cells captured under brightfield imaging contain sufficient information to distinguish live and dead cells, and (2) cells captured under unstained brightfield imaging exhibit similar image features to their stained counterparts, enabling models trained on stained cells to generalize to unstained ones. We then report the development and validation of ViabiLens, an AI-assisted software for label-free cell viability analysis. The ViabiLens combines a cell detection model for localizing individual cells with a convolutional neural network (CNN) classifier for live/dead prediction, paired with an interactive UMAP-based viewer for visualizing and exploring individual cells across the sample. Evaluated on Chinese Hamster Ovary (CHO) cells spanning a wide range of viability conditions, ViabiLens achieves a mean absolute error of 2.68\% on unstained samples against fluorescence-based reference measurements. We also release a benchmark dataset for label-free cell viability analysis to facilitate future research, available at this https URL.

10. 【2610.10457】MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration

链接:https://arxiv.org/abs/2610.10457

作者:Yuxiang Xiong,Ruiyan Wang,Wenqiang Wang,Teng Hu,Songhang Shen,Bohao Feng,Hongqian Deng,Ran Yi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Diffusion Transformers, high inference latency, iterative denoising process, denoising process suffers, achieve remarkable performance

备注: 22 pages, 8 figures

点击查看摘要

Abstract:Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at this https URL.

11. 【2610.10448】MemoCare: An Interactive Multimodal Mobile System for Automated Cognitive Screening

链接:https://arxiv.org/abs/2610.10448

作者:Duy-Cat Can,Mau Minh Phuc Le,Tuan-Khoa Hoang,Hai-Dang Nguyen,Trung-Hieu Do,Dang Minh Ly,Minh-Duc Nguyen,Nghia TT Hoang,Linh-Trung Nguyen,Huy-Hieu Pham,Huong Ha,Binh T. Nguyen,Oliver Y. Chén

类目:Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:interactive mobile system, automated multimodal cognitive, React Native application, interactive mobile, mobile system

备注: 8 pages, 1 figure, 1 table. Demo paper submitted to the MMM 2027 Demo Track

点击查看摘要

Abstract:MemoCare is an interactive mobile system for automated multimodal cognitive screening. A React Native application combines spoken responses, temporal and spatial orientation, touchscreen actions, and visuoconstruction in complete English and Vietnamese workflows. Speech is transcribed by Google Speech-to-Text and scored locally with deterministic task-specific natural language processing rules; GPS coordinates are resolved by the MemoCare spatial module before answer matching; touch tasks are scored from interaction events; and the drawing task uses a three-model convolutional neural network consensus with separate visual interpretation. Software tests pass 151/151 predefined cases across speech/language, spatial-answer, and touch-interaction scoring, while spatial regression passes 48/48 four-country coordinate-resolution cases. For the drawing module, validation-selected ShuffleNetV2 x1.5 achieved 91.33% mean balanced accuracy and 78.87% exact three-criterion accuracy on a locked 71-image test set. Four clinician co-authors additionally inspected the end-to-end workflow, yielding a pooled median rating of 4/5 across eight criteria, with item-level medians ranging from 3 to 4.5. At MMM, attendees can directly try a shortened multimodal screening workflow and inspect automatic item-level and total scoring.

12. 【2610.10438】ECHO: Embodied Camera Observations of Human Object Carrying

链接:https://arxiv.org/abs/2610.10438

作者:Xuefei Sun,Lorin Achey,Kali Hamilton,Alberto Speranzon,Gregory Grebe,Yonatan Bisk,Christoffer Heckman

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:assistive agents, people who live, object, human activity, RGB-D

备注:

点击查看摘要

Abstract:Embodied and assistive agents must do more than recognize objects: they must reason about where an object belongs given the layout of an environment and the habits of the people who live in it. Progress on this problem has been limited, in part because no dedicated benchmark or dataset exists to define and evaluate it. Existing RGB-D scan datasets reconstruct static rooms without human activity, while human-object-interaction datasets capture motion without a navigable, fully reconstructed scene or a ground-truth notion of an object's natural destination. We introduce contextual object placement as a benchmark task: predicting an object's destination during an observed object-carrying episode. To support this task, we present Embodied Camera observations of Human Object carrying (ECHO), a large-scale synthetic dataset that pairs dense RGB-D scans of indoor scenes with recordings of an embodied human carrying everyday objects to context-appropriate destinations. ECHO is the first publicly available dataset to combine reconstructed scenes, human activity, natural language, and contextual-placement annotations. It comprises 3,805 human-annotated episodes across 159 floors of 115 HM3D scenes, involving 198 distinct objects. Each floor includes a complete RGB-D scan with human-annotated room labels and a surface list. Each episode provides synchronized RGB-D encounter clips; 6-DoF camera, human, and object trajectories; start and destination surfaces; an action caption; and a human-written context: a single sentence describing the inhabitant's routine that implies the destination without naming it. We evaluate contextual object placement using input-masked probes and an end-to-end baseline. Results show that no single input modality is sufficient, highlighting the need to jointly reason over scene structure, human activity, and contextual knowledge.

13. 【2610.10436】Detecting Adversarial Images through Response Profiles of Vision-Language Models

链接:https://arxiv.org/abs/2610.10436

作者:Arash Vashagh,Roozbeh Razavi-Far

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:text similarity patterns, patterns seemingly plausible, similarity patterns seemingly, text similarity, seemingly plausible

备注:

点击查看摘要

Abstract:Adversarial perturbations can alter the predictions of frozen vision-language models (VLMs) while leaving their confidence and image--text similarity patterns seemingly plausible. We investigate whether we can identify adversarial inputs based on the broader way an image interacts with a collection of general semantic prompts. Our detector summarizes these responses using category-level statistics, relationships among prompts, deviations from clean reference distributions, and stability under weak image transformations, producing a compact response profile that is classified by a lightweight model while the VLM remains fixed. We evaluate the approach on multiple public image datasets, several CLIP-style visual backbones, and a range of gradient-based, optimization-based, automated, and spatial attacks. The detector achieves strong discrimination in attack-specific settings and retains substantial performance when evaluated on attacks not seen during training. Under a controlled detector-specific protocol, the response-profile representation outperforms the evaluated embedding-geometry baselines. Additional analyses show that the feature groups provide complementary information and that the method remains effective under variations in the prompt configuration. We also examine inference cost and performance against detector-aware adaptive attacks. Overall, the results indicate that response patterns across semantic prompts provide a useful complementary signal for adversarial image detection in frozen VLMs.

14. 【2610.10429】SGF+: Decoupling Gradient Flows for Autoregressive Video Generation

链接:https://arxiv.org/abs/2610.10429

作者:Zihan Su,Junhao Zhuang,Yaowei Li,Siwen Lu,Haoran Li,Lingen Li,Haoyu Wu,Weiyang Jin,Songchun Zhang,Haoyang Huang,Chun Yuan,Zeyue Xue,Nan Duan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:current frames, key-value representations, generation requires denoising, future predictions, video generation requires

备注:

点击查看摘要

Abstract:Autoregressive video generation requires denoising the current frames while writing their key-value representations as context for future predictions. However, these two roles typically share parameters, and we find that their gradients exhibit distinct patterns and systematic negative alignment, hindering the joint optimization of visual quality and temporal consistency. We introduce Self Gradient Forcing Plus (SGF+), which assigns separate parameters to context writing and denoising while preserving their interaction through causal attention. Both roles are jointly optimized using the original generation objective without auxiliary losses, with context writing supervised through its contribution to future predictions. This simple change improves visual quality and long-horizon consistency over the evaluated baselines in both framewise and chunkwise generation, without additional video training data or a longer training horizon. Trained on only 5s rollouts, SGF+ supports continuous generation for up to 24 hours without long-video fine-tuning. These results highlight role-specific parameterization as an effective design principle for high-quality autoregressive video generation and native long-horizon extrapolation.

15. 【2610.10423】GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks

链接:https://arxiv.org/abs/2610.10423

作者:Arash Vashagh,Roozbeh Razavi-Far

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:replaced or upgraded, backbone, classifier, limiting reuse, Adversarial

备注:

点击查看摘要

Abstract:Adversarial example detectors are often tied to the classifier backbone they were trained on, limiting reuse when the protected model is replaced or upgraded. Directly transferring such detectors across backbones is challenging because different networks generally produce incompatible internal representations. We propose GraphRectify, a graph-based framework for transferring adversarial image detectors across classifier backbones. GraphRectify learns a structured representation of intermediate classifier features and adapts representations from a new backbone to the detector learned on the original model, enabling detector reuse. We evaluate GraphRectify across multiple datasets, backbone architectures, and adversarial attacks, including detector-aware adaptive attacks that jointly target the classifier and detector. Across the complete evaluation matrix, GraphRectify achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone and the evaluated transfer ablations. The gains are particularly strong for transfers between different backbone families and when sufficient data are available. In contrast, training from scratch remains competitive in the most data-limited settings. These results show that adversarial detection knowledge can transfer effectively across heterogeneous classifier architectures rather than being relearned whenever the protected backbone changes.

16. 【2610.10408】Rubix: Global Correspondence-Free Point Set Alignment through Assignment Geometry

链接:https://arxiv.org/abs/2610.10408

作者:Subhransu S. Bhattacharjee,Dylan Campbell,Rahul Shome

类目:Computer Vision and Pattern Recognition (cs.CV); Computational Geometry (cs.CG); Machine Learning (cs.LG); Robotics (cs.RO); Optimization and Control (math.OC)

关键词:Procrustes-Wasserstein alignment jointly, alignment jointly estimates, suboptimal solutions, jointly estimates, stop at suboptimal

备注: 67 pages, 20 figures. Includes full proofs and experimental appendices

点击查看摘要

Abstract:Procrustes-Wasserstein alignment jointly estimates a matching and rotation without supplied correspondences, but alternating minimization can stop at suboptimal solutions. Rubix solves the equally weighted planar problem globally under squared Euclidean loss. Each matching $\sigma$ of two centered $n$-point sets defines a complex correlation $z_\sigma=\sum_i\bar x_i y_{\sigma(i)}$. Their convex hull is the permutation polygon: supporting vertices give optimal matchings at fixed rotations, and the farthest vertex gives the global alignment. We prove the sharp bound of $n(n-1)$ vertices for $n\ge2$, answering Rote's rotation-assignment open problem. In exact arithmetic, assignment queries recover the polygon in $\mathcal O(n^5)$ operations. Assignment-based bounds extend the approach to three-dimensional rotations and partial matching at a supplied translation through branch-and-bound. On timed MPEG-7 shape pairs, Rubix attains every numerical reference value in 12 ms on average, 50 times faster than a rotation grid at the same accuracy. Its distances improve gravity-aligned matching of real 3D scans, shape retrieval and noisy crystal classification over alternating minimization.

17. 【2610.10400】Self-correction Optimization for Interleaved Multimodal Generation

链接:https://arxiv.org/abs/2610.10400

作者:Xin You,Zhiwei Ning,Zukai Chen,Minghui Zhang,Xuanke Shi,Hanxiao Zhang,Jingsong Liu,Jie Yang,Quan Wang,Yun Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, Multimodal large language, made significant progress, language models, large language

备注: 20 pages, 10 figures

点击查看摘要

Abstract:Multimodal large language models (MLLMs) have made significant progress in visual understanding and generation. However, generating interleaved image--text content remains challenging, as it requires tightly integrated multimodal understanding and generation capabilities. Although existing MLLMs provide promising solutions, most rely on additional training with augmented data, which is computationally expensive and remains limited in preserving visual subjects, temporal consistency, and physical plausibility. In this work, we propose self-correction optimization (SCO), an effective training-free method for consistent interleaved generation. SCO treats the classifier-free guidance update as a reference and performs minimal self-correction under two complementary constraints, including new-event and state-preserving constraints. Specifically, the new-event constraint promotes temporal consistency across image--text sequences, while the state-preserving constraint maintains the coherence of visual subjects throughout subsequent generation steps. Experiments on challenging interleaved multimodal generation benchmarks demonstrate significant improvements in temporal coherence and visual-subject preservation. Furthermore, SCO can be extended to video generation and improves the modeling of physically grounded processes, including robot manipulation and long-horizon handcrafting.

18. 【2610.10396】Gaussian Density Splatting Network

链接:https://arxiv.org/abs/2610.10396

作者:Miao Shang,Yabin Wang,Xiaopeng Hong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Density Splatting, paper proposes, Gaussian, Gaussian Splatting framework, differentiable Gaussian Splatting

备注: This is the preprint version of the paper and supplemental material to appear in NeurIPS, 2026. Please cite the final published version

点击查看摘要

Abstract:This paper proposes a novel crowd counting approach, the Gaussian Density Splatting Network (GDSNet). Unlike methods that rely on conventional, grid-based density maps and are sensitive to spatial resolution, GDSNet represents a crowd as a superposition of continuous 2D Gaussian primitives. Our approach is built upon two key contributions. First, we introduce a control-point-based fitting mechanism to structure the prediction of the Gaussian parameters. We design a method to allocate a set of control points that define local regions, from which features are pooled to regress each primitive's parameters. Second, we adapt a differentiable Gaussian Splatting framework to the counting task by parameterizing each primitive with geometric parameters and a scalar density mass. This formulation allows the network to be trained end-to-end via spatial matching of differentiably rendered density maps, naturally providing both local density supervision and global count optimization. Extensive evaluations on four standard benchmarks show GDSNet consistently outperforms the state of the art.

19. 【2610.10391】MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking

链接:https://arxiv.org/abs/2610.10391

作者:Benoît Roussel,Damien Bouet,Liming Chen,Pierre Perrault

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:association-difficult benchmarks, narrowed the gap, multi-object trackers, spatial, HOTA on DanceTrack

备注: Accepted at BMVC 2026. 27 pages (14 pages main paper, appendix and references), 6 figures, 10 tables

点击查看摘要

Abstract:End-to-end multi-object trackers have narrowed the gap with classical tracking-by-detection on association-difficult benchmarks. Yet they still make spatially implausible errors no classical tracker would, such as assigning one identity to objects on opposite sides of the frame. A model could learn to avoid them, but tracking annotations are scarce, so we encode spatial priors explicitly instead, while keeping inference fully end-to-end with no post-hoc association. We propose three spatial priors, at the data, loss, and representation stages. Spatial ID Switches bias trajectory permutations toward spatially overlapping objects, reducing the mismatch between training and inference confusions. Spatial ID Loss scales each identity's penalty by its box distance, so a distant switch costs more than a nearby one. Spatial Anchor gives each track token its frame position, an explicit spatial cue for attention. We instantiate the three priors in MOTIP2, a tracker adapted from MOTIP and built on the real-time DEIM detection transformer. Trained without extra data, its main model, MOTIP2-L, sets a new state of the art: 73.4 HOTA on DanceTrack, 76.0 on SportsMOT, and 71.1 IDF1 on PersonPath22. MOTIP2 is a family of models spanning the speed-accuracy trade-off: a lighter model, MOTIP2-S, matches the original MOTIP at over 3x the speed, and MOTIP2-X reaches 74.8 HOTA on DanceTrack.

Comments:
Accepted at BMVC 2026. 27 pages (14 pages main paper, appendix and references), 6 figures, 10 tables

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.10391 [cs.CV]

(or
arXiv:2610.10391v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.10391

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
20. 【2610.10390】Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

链接:https://arxiv.org/abs/2610.10390

作者:Xingtai Gui,Yucheng Zhou,Dongqian Guo,Jiahao Gong,Feiyang Tan,Jianbing Shen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing VLA models, VLA models, autonomous driving, promising paradigm, geometric

备注: 21 pages, 9 figures. The code is available at [this https URL](https://github.com/TabGuigui/GeoCoTDrive)

点击查看摘要

Abstract:Vision-language-action~(VLA) models have emerged as a promising paradigm for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a think with 2D first, drive with dedicated 3D priors paradigm. It first grounds 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support the trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning decisions, and construct the PlanningGrounding dataset to endow VLAs with planning-oriented grounding capability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of the explicit geometric chain-of-thought process for VLA-based planning.

21. 【2610.10388】RoboQuest: Generalist Physical Agents that Search, Inspect and Test

链接:https://arxiv.org/abs/2610.10388

作者:Liu Renhang,Navonil Majumder,Tej Deep Pala,Soujanya Poria

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, capable generalist physical, generalist physical agents, made them capable, capable generalist

备注:

点击查看摘要

Abstract:Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a $\pi_{0.5}$ policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23\% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.

22. 【2610.10370】PalmSpace: Towards a Versatile On-Palm Interaction Space through Unified Touch Modeling

链接:https://arxiv.org/abs/2610.10370

作者:Chentao Li,Mingze Gao,Runze Sun,Zhaoguo Wang,Jianjiang Feng,Jie Zhou

类目:Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)

关键词:increasingly practical, key challenge, smart glasses, glasses and lightweight, lightweight MR devices

备注: Preprint. Initial version

点击查看摘要

Abstract:As smart glasses and lightweight MR devices become increasingly practical, input remains a key challenge. The bare palm is an always-available, tactile, and proprioceptively accessible surface, but it has neither an explicit coordinate system nor embedded touch sensing. Prior on-palm systems typically expose isolated touch events, discrete regions, continuous trajectories, or task-specific gestures, limiting the palm's ability to support precise selection and gesture manipulation through a common input representation. We present PalmSpace, a wrist-worn infrared system that exposes mode-aware, body-referenced absolute input on the bare palm without per-user sensing calibration. At the interaction level, PalmSpace jointly represents contact occurrence, interaction mode, and palm-referenced absolute location; at the model level, it learns these coupled outputs through a shared real-time representation. In leave-one-participant-out evaluation with 17 participants, PalmSpace achieved 6.7 mm mean localization error, 98.9% contact detection accuracy, and 96.7% F1 for four-class interaction-state recognition. User studies further demonstrated absolute pointing and dragging, eyes-free digit input, and representative multi-finger controls including scrolling and pinch-based map manipulation. These results show that a morphologically variable bare palm can function as a transferable, mode-aware interaction surface.

23. 【2610.10359】MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency

链接:https://arxiv.org/abs/2610.10359

作者:Markus Gross,Andreas Greiner,Taehyoung Kim,Sivasubiramaniam Subbiah,Tomaž Cotič,Sai Bharadwaj Matha,Conrad Christoph,Oussema Dhaouadi,Simon Zieher,Surya Vijaya Kumar,Gordon Elger,Henri Meeß,Olaf Wysocki,Paul Spannaus,Daniel Cremers

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:low-altitude UAV dataset, UAV dataset, low-altitude UAV, RGB, RGB images

备注:

点击查看摘要

Abstract:We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly provides 17,272 synchronized samples from four suburban scenes with frame-wise annotations for 15 semantic classes, together with calibration and GNSS-RTK/IMU measurements. To avoid costly and inconsistent modality-specific annotation, we propagate labels from only 115 manually annotated RGB images through shared geometric representations to all four modalities. This approach generates semantic labels for 17,157 additional RGB images, 17,272 thermal images, 840M LiDAR points, and 3.4M radar points. Transferred annotations achieve 89.93% average agreement with held-out manual annotations, and 90.94% average semantic consistency across all six modality pairs. We further establish semantic segmentation benchmarks for all four modalities, revealing distinct architectural behavior for dense LiDAR and sparse radar data. Taken together, MultiFly provides a scalable foundation for multimodal aerial perception and, to the best of our knowledge, the first public real-world low-altitude aerial benchmark that combines consistent frame-wise semantic annotations for RGB, thermal, LiDAR, and radar. Data at this https URL.

24. 【2610.10343】Real-Time Joint Audio-Video Generation by Parallel Adapter Composition

链接:https://arxiv.org/abs/2610.10343

作者:Jingyu Li,Xiaoxiao Xiang,Yiwen Guo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:audio-video diffusion transformer, block-autoregressive attention, transformer for real-time, essential modifications, clip is finished

备注:

点击查看摘要

Abstract:Deploying a joint audio-video diffusion transformer for real-time, interactive generation normally requires two essential modifications: block-autoregressive attention, so frames can be emitted before the whole clip is finished, and few-step sampling, so each block is cheap. Conventionally, the streaming video literature obtains both capabilities from a chained pipeline. It first distills a bidirectional teacher into a causal student, then into a few-step one, or proceeds in reverse order. Each stage of such a chain fine-tunes the weights the previous one produced, so a later objective can undo an earlier capability. Following the idea of model merging, we show that on a packed audio-video backbone the two capabilities can be acquired in parallel. A causal adapter is trained against the frozen backbone, and an off-the-shelf few-step adapter provides the few-step capability. As the two edit different functional axes, we predict, and then verify, that their weight-update directions are near-orthogonal, without any explicit orthogonality constraint during training. Orthogonal updates should combine without interfering, so parallel composition is a direct sum. The two adapters are simply added at inference, with no joint training, yielding few-step, streaming audio-video whose image quality tracks the bidirectional teacher. Compared to the chained baselines, the composed model matches or beats them on most metrics, making parallel composition a practical approach. The resulting streaming system generates joint audio-video in real time, $\approx$26 fps at $480\times832$ without quantization, and sustains 30 s of continuous generation with stable image quality.

25. 【2610.10342】Position Forcing: Self-Conditioning 3D Generation

链接:https://arxiv.org/abs/2610.10342

作者:Ziheng Ouyang,Zeqiang Lai,Jiarui Chen,Jiangshan Wang,Yuhao Wan,Jingbo Gong,Xiangyu Yue,Hengshuang Zhao,Qibin Hou,Chunchao Guo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models commonly adopt, adopt VecSet representations, commonly adopt VecSet, Recent single-stage, Position Forcing

备注:

点击查看摘要

Abstract:Recent single-stage 3D generative models commonly adopt VecSet representations, encoding 3D shapes as unordered sets of latent tokens. However, compared with two-stage methods that provide explicit positional guidance, these models must implicitly infer token positions throughout denoising, limiting their generation quality. We observe that, despite the absence of explicit positional conditioning, VecSet tokens retain recoverable spatial correspondences. Building on this observation, we propose Position Forcing, a position-based self-conditioning framework. During denoising, Position Forcing recovers token positions from the current clean latent estimate, quantizes them at progressively finer resolutions according to the denoising stage, and feeds the resulting positional encodings back into the diffusion Transformer. This progressively refined positional feedback provides spatial guidance at a granularity appropriate to each denoising stage, guiding shape generation along a coarse-to-fine trajectory and substantially improving generation quality without a separate position generation stage. Experiments demonstrate that Position Forcing achieves strong performance among single-stage 3D generative methods and outperforms several competitive multi-stage approaches.

26. 【2610.10335】When to Unpair: Regulating Pairing Dependence in Medical Visual In-Context Learning

链接:https://arxiv.org/abs/2610.10335

作者:Cheng Wan,Chenjun Li,Qingyu Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual in-context learning, label-scarce medical imaging, demonstrate input-output mappings, Visual in-context, support image-label pairs

备注: 24 pages, 12 figures

点击查看摘要

Abstract:Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, while the labels collectively indicate the requested task. We diagnose dependence on individual pairings with a test-time derangement that reassigns every support label to another support image while preserving the query, support images, and label multiset. The resulting pairing gap, defined as shuffled-minus-matched performance, shows that all four released models depend on the pairing, to widely varying degrees. Further analysis of a paired-trained model reveals support-associated spurious regions and lesion-size biases even with real, unaltered supports, alongside sensitivity to mis-registered support labels. To regulate this dependence, we introduce a late unpairing curriculum (LUC), which starts with matched training and then applies random unpairing, replacing each support label with that of another support in the same episode. LUC nearly closes the pairing gap on two backbones while maintaining or improving matched-support performance across all evaluated task types, with gains extending to held-out tasks and cross-dataset episodes. It also mitigates these failure modes. On BraTS whole-tumor segmentation, matched-support DSC rises from 0.733 to 0.857 while the gap shrinks from -0.184 to -0.008. In a released model, brief fine-tuning with random unpairing reduces the gap. A reversed curriculum that places the same number of unpairing epochs at the start of training leaves a large gap. This shows that pairing dependence is shaped by the order of training and not only by the amount of unpaired training.

27. 【2610.10334】How Private is Private? A Comparative Study for Face De-Identification

链接:https://arxiv.org/abs/2610.10334

作者:Hui Wei,Hao Yu,Hui Kuurila-Zhang,Guoying Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains fundamentally fragmented, critical privacy-preserving technology, fundamentally fragmented, Hierarchical Face De-identification, evaluation remains fundamentally

备注: Accepted to NeurIPS 2026. Project Page: [this https URL](https://cv-ac.github.io/hifd/)

点击查看摘要

Abstract:Face de-identification (FDeID) has emerged as a critical privacy-preserving technology, yet its evaluation remains fundamentally fragmented. Existing protocols rely on inconsistent metrics, heterogeneous datasets, and partial annotation coverage, so methods targeting different utility dimensions, such as landmark versus expression preservation, are reported on different benchmarks under different metrics, rendering cross-method comparison infeasible. We revisit FDeID evaluation from both the data and metric perspectives. On the data side, we introduce UtilFace, a curated, demographically balanced benchmark with high identity diversity, assembled from four large-scale face datasets through identity-aware cleaning, resolution enhancement, and stratified filtering. On the metric side, we propose HiFD, a Hierarchical Face De-identification metric that unifies identity suppression, multi-level utility preservation, and image quality under a single consistency-based paradigm: every component is computed from pretrained estimators' outputs on the original face and its de-identified counterpart, directly quantifying how much identity is suppressed and how much downstream-perceivable utility survives. HiFD organizes facial signals into a three-level utility hierarchy spanning macro cues (L1), micro cues (L2), and imperceptible cues (L3), and aggregates the five resulting components into a single interpretable score via weighted harmonic mean, with configurable application-specific profiles. Using this unified protocol, we conduct a comprehensive comparative study spanning adversarial, GAN-based, and diffusion-based methods, surfacing trade-offs and failure modes that remain invisible under existing protocols. We release the benchmark and evaluation toolkit to foster systematic and reproducible research in privacy-preserving human face analysis.

28. 【2610.10324】Performance at What Cost? A Sustainability-Aware Performance Index for Cell and Nucleus Instance Segmentation

链接:https://arxiv.org/abs/2610.10324

作者:Eiram Mahera Sheikh,Alaa Tharwat,Wolfram Schenck

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:nuclear instance segmentation, parameter count, postprocessing pipeline, pretraining data, data and objectives

备注:

点击查看摘要

Abstract:Pretrained models for cell and nuclear instance segmentation differ substantially in architecture, pretraining data and objectives, parameter count, inference strategy, adaptation requirements, postprocessing pipeline, and computational demand. Large pretrained and foundation models are increasingly adopted because of their strong zero-shot capabilities, but their use also imposes greater energy consumption, memory requirements, computational demands, adaptation costs, and operational carbon emissions. Whether these additional demands are justified by meaningful gains in segmentation performance remains unclear. We address this question by introducing the Sustainability-Aware Performance Index (SAPI), a configurable metric that combines segmentation performance, energy consumption, and model size. We benchmark 19 pretrained and foundation models across six CellBinDB datasets under zero-shot inference and evaluate 16 fine-tunable models using few-shot adaptation with both frozen encoder and full-model fine-tuning. We estimate energy consumption for GPU, CPU, and RAM using software-based monitoring tools. Our results show that larger and more computationally demanding models do not consistently achieve proportionate improvements in segmentation quality. While few-shot adaptation benefits several models, the gains and resource costs vary considerably across architectures, datasets, and adaptation strategies, causing SAPI-based rankings to differ from rankings based on performance alone. This study provides a practical framework for comparing segmentation models more comprehensively and supports more computationally accessible and environmentally responsible model selection in biomedical image analysis.

29. 【2610.10322】From Digital Human Interactions to Physics-Based Humanoid Skills: Physics-Grounded Post-Training of Interaction Generators

链接:https://arxiv.org/abs/2610.10322

作者:Kerui Chen,Jianrong Zhang,Kai Lv,Hehe Fan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:made promising progress, Recent methods, physics-based tracking policies, convert digital reference, largely relying

备注:

点击查看摘要

Abstract:Recent methods have made promising progress in generating interactions between two humanoids, largely relying on physics-based tracking policies to convert digital reference motions into executable trajectories. However, limited tracking capabilities restrict the range of reference motions that can be successfully executed, reducing data utilization. Moreover, even successful tracking does not guarantee physically plausible responses or faithful realization of the intended interactions. In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy. Our DIGHT first executes multiple text-conditioned interaction candidates in simulation using a fixed tracker. It then constructs physics-grounded preferences from the resulting rollouts, covering both general executability and interaction fidelity. Rather than collapsing these signals into a single scalar reward for candidate ranking, we align the pretrained generator using physics-decoupled diffusion direct preference optimization (DPO), preserving criterion-specific supervision without differentiating through the simulator. To improve executability, preference pairs are derived from tracking error, friction, and floating. Additionally, to improve interaction fidelity, we propose to incorporate force feedback from simulator as a measure of contact fidelity and construct preferences over contact occurrence, location, duration, and force magnitude. The aligned generator then supplies reference motions for fine-tuning the tracker, improving compatibility between generation and physical execution. Extensive experiments demonstrate that our approach not only improves the physical plausibility of generated motions but also enables more reliable and faithful humanoid interactions in simulation.

30. 【2610.10306】One-Shot Adaptive Segmentation For Scientific Images

链接:https://arxiv.org/abs/2610.10306

作者:Tejaswi V. Panchagnula,Allison M. Davis,Fengqing Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:task-specific training, limiting adaptation, experimental conditions, rely on extensive, extensive annotation

备注:

点击查看摘要

Abstract:Scientific image segmentation methods rely on extensive annotation and task-specific training, limiting adaptation across imaging modalities and experimental conditions. We present a training-free, one-shot framework that specializes vision foundation models using a single annotated reference image. The framework combines DINOv3 representations with background-adaptive feature orthogonalization to suppress artifact-related feature directions, after which cosine similarity localizes candidate regions for SAM segmentation. We evaluate the framework on red-blood-cell microscopy, structured-illumination pool boiling, and chest radiography. Relative to the strongest baseline, the proposed method improves mean IoU by 5.91% and 78.62% on the microscopy and pool-boiling datasets, respectively, while achieving comparable performance on chest radiographs. These results demonstrate that one-shot reference conditioning can adapt general-purpose vision models to specialized scientific segmentation tasks.

31. 【2610.10303】On the Necessity of Attention-FFN Split in Vision Transformers

链接:https://arxiv.org/abs/2610.10303

作者:Junhyeok Kim,Jinyeong Kim,Jae Wan Park,Seong Jae Hwang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Feed-Forward Network, Transformer architecture relies, unified Vision Transformer, pattern that alternates, alternates Attention

备注:

点击查看摘要

Abstract:The standard Transformer architecture relies on a rigid pattern that alternates Attention and Feed-Forward Network (FFN) layers. Despite its widespread adoption, the inductive bias imposed by this strict separation has not been systematically examined. In this work, we investigate the necessity of the Attention-FFN dichotomy in Vision Transformers (ViTs). To facilitate this analysis, we introduce the AttenFeed module, a unified component that integrates the functional properties of both Attention and FFN. Based on this module, we devise the unified Vision Transformer (uViT), which replaces the conventional alternating Attention-FFN structure with a sequence of AttenFeed modules. We then use uViT as a control group that relaxes the Attention-FFN dichotomy of the standard ViT and systematically compare the two models across multiple datasets and model scales. Our experiments reveal that the Attention-FFN dichotomy can hinder performance at smaller model scales due to the rigid parameter allocation of ViTs. The AttenFeed module and uViT serve as new analytical tools for understanding the Attention-FFN structure and offer theoretical insights into the heuristically designed architecture of conventional ViTs.

32. 【2610.10288】ouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning

链接:https://arxiv.org/abs/2610.10288

作者:Dayou Li,Hao Wang,Qianqian Yang,Zihao Zhu,Haoquan Fang,Ziyao Zeng,Yan Han,Zihan Wang,Yan Wang,Baoru Huang,Dilin Wang,Kenji Shimada,Yiyue Luo,Manling Li,Teresa Lv,Mustafa Mukadam,Rakesh Ranjan,Ruohan Zhang,Qi He,Changliu Liu,Xu Chen,Marco Pavone,Bangya Liu,Jiachen Li,Masayoshi Tomizuka,Zhiwen Fan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:physical interaction unrecorded, Large-scale egocentric human, characterize physical interaction, Large-scale egocentric, characterize physical

备注: Project page: [this https URL](https://touch-scale.github.io/)

点击查看摘要

Abstract:Large-scale egocentric human interaction data is becoming an important source of physical supervision for embodied learning, yet video alone leaves the contact and pressure that characterize physical interaction unrecorded. Recent visual-tactile datasets provide this missing supervision, but their synchronized tactile data remain far smaller in volume than human video. Moreover, the largest resources often merge recordings from different sensors or annotation procedures, which makes the effect of data scale difficult to isolate. We therefore introduce TouchScale, a 500-hour dataset of contact-rich human interaction recorded with a single unified wearable setup. Its approximately 2K predefined task descriptions span everyday activities and structured manipulation, and each recording temporally aligns egocentric RGB-D video with wrist RGB video and dense full-hand bimanual tactile measurements. Compared with prior tactile data, training on the full TouchScale raises zero-shot contact IoU on data from an unseen tactile sensor from 0.134 to 0.383. Pretraining a visual encoder on TouchScale also yields the highest action recognition accuracy on three benchmarks among the compared visual-tactile datasets. Used for visual-tactile mid-training of a robot policy, TouchScale improves the average real-world success rate across four contact-rich manipulation tasks from 22.5% to 57.5%. With the sensor and collection protocol held fixed, both zero-shot tactile prediction and robot success show an overall upward trend as more TouchScale data is used. These results suggest that human visual-tactile data collected at scale with consistent sensing benefits both perception and robot manipulation. We will publicly release TouchScale, including all synchronized visual-tactile recordings and reconstructed object models, to support future research on scalable visual-tactile learning.

33. 【2610.10286】$Δ$Representation: Geometry Supervised Representation Learning of Phenotypes via Counterfactual Reasoning for Medical VLMs

链接:https://arxiv.org/abs/2610.10286

作者:Hao Wang,Qiwei Zeng,Jinghao Lin,Shuchang Ye,Yuezhe Yang,Yige Peng,Haoyuan Che,Jinman Kim,Lei Bi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:shown increasing potential, radiological image interpretation, representation, Delta, phenotype

备注:

点击查看摘要

Abstract:Medical vision-language models (VLMs) have shown increasing potential for radiological image interpretation. Medical VLMs encode radiological images into visual representations that capture both anatomical and phenotypic information for diagnosis. Existing approaches improve pathological phenotype representations through semantic-guided representation alignment. However, pathological phenotypes arise as lesion-specific visual changes superimposed on underlying normal anatomy. Such semantic alignment approaches fail to model the phenotype-specific increment relative to the corresponding normal anatomical representation. To address this gap, we propose \textbf{$\Delta$Representation}, a visual phenotype representation learning framework based on counterfactual reasoning for medical VLMs. It comprises \textbf{BaseAnatomy}, a geometry-supervised representation learning module, and \textbf{$\Delta$Phenotype}, a counterfactual incremental representation learning module. BaseAnatomy provides fine-grained geometric supervision through spatial relationships across and within anatomical structures. $\Delta$Phenotype computes the representation increment between lesion representations and their corresponding normal anatomical representations, and supervises increments associated with the same phenotype to cluster in the representation space. Experiments on \textit{ReXGroundingCT} and \textit{LIDC-IDRI} demonstrate that $\Delta$Representation effectively structures pathological phenotype representations and improves lesion grounding and phenotype characterization accuracy in medical VLMs. Code is available at this https URL.

34. 【2610.10283】mporal Visuo-Tactile Learning for Dexterous Grasp Stability

链接:https://arxiv.org/abs/2610.10283

作者:Ken Nakahara,Aleksei Buvailik,Prokhor Kotov,Roberto Calandra

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:literature emphasizes vision-based, grasping literature emphasizes, emphasizes vision-based grasp, vision-based grasp selection, parallel grippers

备注: 12 Pages. Website: [this https URL](https://lasr-lab.github.io/dexterous-grasp-stability/)

点击查看摘要

Abstract:Humans can grasp everyday objects with almost perfect success rates using fingertip tactile feedback, yet much of the robotic grasping literature emphasizes vision-based grasp selection with parallel grippers. In this work, we systematically investigate how high-resolution, dynamic tactile sensing contributes to grasp stability prediction and model-guided grasping in dexterous robotic hands. To this end, we collected a dataset of 10,000 grasp trials across 200 objects using a multi-fingered robotic hand equipped with four Digit 360 tactile sensors, recording external vision, proprioception, and tactile streams throughout each grasp. With this dataset, we trained end-to-end temporal multimodal models to predict post-lift stability from pre-lift grasp observations and compared sensing modalities and encoding backbones. Experimental results and controlled input ablations show that incorporating touch, and particularly high-resolution, dynamic touch, improves grasp stability prediction. Finally, we deployed the learned predictor as an online stability gate on the real robot, where visuo-tactile model-guided regrasping improved the success rate among executed lifts by 10.5 percentage points over a non-tactile gate. These results show how rich fingertip sensing and expressive temporal models that capture the dynamics of touch can support learned grasping with multi-fingered hands without explicit contact or force modeling, providing a scalable data-driven path from tactile experience toward stable dexterous manipulation. The dataset is publicly available at this https URL.

35. 【2610.10270】Video Prediction Policy 2: Predict Better, Act Better

链接:https://arxiv.org/abs/2610.10270

作者:Yanjiang Guo,Haodong Yan,Zhide Zhong,Zhongru Zhang,Qingyuan Yang,Qingzhou Lu,Xiaoyu Chen,Yen-Jen Wang,Shuying Deng,Chenghan Yang,Puzhen Yuan,Chenxin Liu,Tun Ban,Xiang Zhu,Yichen Liu,Kun Feng,Haoang Li,Jianyu Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:generalist robot policies, World action models, World action, video, robot policies

备注:

点击查看摘要

Abstract:World action models (WAMs) have emerged as an important class of generalist robot policies, aiming to transfer video prediction priors to action learning. However, we find that existing WAMs frequently produce incorrect motion predictions in open-ended environment, leading to erroneous actions. We attribute this limitation to two factors: (1) base video models are not optimized for manipulation, and (2) naively incorporating action components into video models can substantially degrade their generalization capabilities. We introduce Video Prediction Policy 2 (VPP2), a WAM that enables strong zero-shot generalization in both video prediction and action generation. First, we curate a large-scale, diverse dataset of manipulation videos to continue pretraining the base video foundation model. We annotate video clips with detailed captions and perform \textit{event-level} video pretraining to promote generalization across open-ended manipulation tasks. Second, we post-train and distill the video model into a single-step visual planner with fixed prediction horizon. Finally, we introduce action module via a mixture-of-transformers (MoT) architecture to learn implicit inverse dynamics model. Experiments demonstrate three key results: (1) VPP2-14B outperforms Cosmos3-64B by 11.0\% points in video prediction instruction-following success rate on open-ended tasks; (2) VPP2 surpasses the strongest baseline by 18.5\% points in success rate on real-world zero-shot ALOHA manipulation tasks; and (3) following benchmark-specific post-training, VPP2 achieves the highest success rates among evaluated methods on the challenging LIBERO-Pro, LIBERO-OOD, and RoboDojo benchmarks.

36. 【2610.10266】LoomSC: Scalable Deep Subspace Clustering with Projector Factorization and Exact Spectral Reduction

链接:https://arxiv.org/abs/2610.10266

作者:Nairouz Mrabah,Youssef Melki,Mohamed Bouguessa,Riadh Ksantini,Shakeeb Murtaza,Tehseen Zia

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Dense self-expression matrices, Orthogonal Optimization Model, full-affinity spectral clustering, spectral clustering limit, Dense self-expression

备注: 19 pages, 7 figures, 5 tables; includes appendices

点击查看摘要

Abstract:Dense self-expression matrices and full-affinity spectral clustering limit the scalability of subspace clustering. We introduce the Latent Orthogonal Optimization Model for Subspace Clustering (LoomSC), a framework that addresses both bottlenecks through projector factorization and exact spectral reduction. Motivated by the spectral structure of least-squares regression, LoomSC jointly learns latent features and a projector self-representation through two thin factors. Alternating Procrustes and least-squares updates preserve the sample factor's orthogonality while keeping the coefficient matrix implicit. We construct a nonnegative quadratic affinity that preserves the projector's support. An exact feature map then reduces its normalized spectral problem to an eigenproblem whose dimension depends only on the factor width. Neither the full affinity nor the sample Laplacian needs to be formed. Our analysis quantifies the projector approximation and identifies conditions for subspace preservation and within-subspace connectivity. For fixed dimensions and iteration budgets, the complete pipeline has linear time and memory complexity in the number of samples. Across five image-clustering benchmarks, LoomSC ranks first or second in all 15 dataset-metric comparisons against 9 state-of-the-art baselines. Its mean accuracy exceeds the highest baseline mean by 6.66 percentage points. Synthetic experiments scale to 500,000 samples while maintaining at least 99.8% accuracy.

37. 【2610.10238】Geometry-Supervised Visual Representation Learning for Multi-Phenotype Lesion Interpretation in Medical VLMs

链接:https://arxiv.org/abs/2610.10238

作者:Hao Wang,Qiwei Zeng,Shuchang Ye,Jinghao Lin,Yuezhe Yang,Yige Peng,Jinman Kim,Lei Bi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:shown increasing potential, clinical image interpretation, shown increasing, increasing potential, potential for clinical

备注:

点击查看摘要

Abstract:Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation. However, these models still struggle to interpret multi-phenotype lesions whose diagnosis requires the joint assessment of multiple pathological phenotypes. Existing vision-language alignment methods produce visual representations that fail to preserve anatomical hierarchies and relationships among phenotypic subclasses. This stems from their reliance on semantic supervision, which lacks geometric constraints to preserve these relationships in the visual embedding space. Moreover, the sparsity of lesion-related anatomical and phenotypic representations makes it difficult for medical VLMs to capture important diagnostic evidence. To address these limitations, we propose \textbf{PureVision}, a geometry-supervised visual representation learning framework for multi-phenotype lesion interpretation in medical VLMs. It combines a geometry-supervised representation learning module, \textbf{PureEyes}, and an anatomy-guided evidence aggregation module, \textbf{PureNeurons}. PureEyes provides geometric supervision through ideal spatial distributions that encode anatomical hierarchies and phenotypic subclass relationships. PureNeurons projects visual representations into the learned latent space, using their positions to selectively aggregate lesion-specific anatomical and phenotypic evidence. Experiments on \textit{LIDC-IDRI}, \textit{CBIS-DDSM}, and \textit{3DReasonKnee} demonstrate that PureVision improves lesion grounding and phenotype characterization in visual question answering and radiology report generation. Code is available at: this https URL.

38. 【2610.10225】Masked Feature Encoding for Large-Scale Whole Slide Image Representation

链接:https://arxiv.org/abs/2610.10225

作者:Haoyu He,Basile Tessier-Cloutier,Yang Wang,Mahdi S.Hosseini

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multiple instance learning, instance learning, multiple instance, Masked Feature Encoding, slide image

备注: Accepted at ACCV 2026

点击查看摘要

Abstract:Whole slide image (WSI) analysis in computational pathology follows a multiple instance learning (MIL) pipeline where patch embeddings are extracted independently and aggregated for slide-level prediction, but within-slide variance from staining, scanner, and local texture can overwhelm the discriminative signal. We propose Masked Feature Encoding for Multiple Instance Learning (MFE-MIL), a feature-space masking framework that trains a lightweight MLP adapter jointly with a window-based masked reconstruction branch and a MIL classification head. The two objectives are complementary. Classification guides the adapter to suppress within-slide patch variance, while window-based masked reconstruction provides an auxiliary regularizer for the adapted features without using patch coordinates, coordinate graphs, or segmentation preprocessing. The raster patch-extraction order is used only as a weak implicit prior. At inference, the decoder is removed, leaving only the adapter and MIL head. Across CAMELYON16/17, PANDA, and TCGA-BRCA with four diverse encoders, MFE-MIL improves ACC/F1 for nearly all tested aggregator-encoder settings and AUC in most, outperforms coordinate-based spatial methods (CAMIL), and achieves higher AUC than 2DMamba on three of four datasets (UNI). On five TCGA survival cohorts it improves the average concordance index for every aggregator tested, its most consistent gain. Code is available at this https URL.

39. 【2610.10201】GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

链接:https://arxiv.org/abs/2610.10201

作者:Jingyao Zhang,Yun Li,Lu Han

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:reasoning requires translating, function reasoning requires, perceived spatial configuration, satisfies geometric constraints, curve satisfies geometric

备注: 15 pages, 1 figure, 7 tables

点击查看摘要

Abstract:Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.

40. 【2610.10197】VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation

链接:https://arxiv.org/abs/2610.10197

作者:Zhuo Chen,Yihua Cheng,Aleš Leonardis,Hyung Jin Chang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurate contact modeling, understanding hand-object interaction, Accurate contact, existing contact representations, recover contact details

备注: Accepted to NeurIPS 2026

点击查看摘要

Abstract:Accurate contact modeling is fundamental to understanding hand-object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand-crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion-capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior-guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state-of-the-art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations. Our code is available at this https URL.

41. 【2610.10196】HuLiGen: Human LiDAR Generation from Parametric Body Models

链接:https://arxiv.org/abs/2610.10196

作者:Salma Galaaoui,Nermin Samet,David Picard

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:LiDAR point clouds, collect and annotate, extremely expensive, expensive to collect, represent a scarce

备注: 12 pages, 5 figures, 7 tables

点击查看摘要

Abstract:LiDAR point clouds of humans are extremely expensive to collect and annotate, thus represent a scarce resource that hinders the development of human analysis using this modality. To alleviate this scarcity, prior work relies on simulated human LiDAR, but such samples do not fully reflect the geometry and sensing characteristics of real observations. In contrast, we introduce HuLiGen, a generative model that generates human LiDAR point clouds from a parametric body model, using a point transformer trained with a flow-matching objective. We show that our generated point clouds are closer to the real capture distribution. Using HuLiGen to generate synthetic data, we propose a synthetic-only pretraining scheme for LiDAR-based HPE that achieves state-of-the-art performance, with even larger gains in low-annotation and low-data regimes, where MPJPE is reduced by up to 50%. Code, models and generated samples are available at this https URL.

42. 【2610.10183】VideoEvolve: Co-Evolving Memory and Retrieval for Long Video Understanding

链接:https://arxiv.org/abs/2610.10183

作者:Yongchao Xu,Bowen Ye,Jiefeng Gan,Junkai Ma,Wenzhao Li,Sen Tao,Yi Wei,Jiawei Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:organize massive visual, massive visual streams, understanding increasingly relies, compact representations, Long video understanding

备注:

点击查看摘要

Abstract:Long video understanding increasingly relies on external memory to organize massive visual streams into compact representations. However, most memory-based methods dynamically adapt how information is retrieved for different questions, while largely fixing what is remembered. This mismatch makes missing details costly to recover, whereas stored information is valuable only when it can be reliably retrieved. To address this issue, we propose VideoEvolve, a novel self-evolving framework that jointly evolves memory and retrieval for long video understanding. Specifically, starting from a coarse low-frame-rate overview, VideoEvolve couples a Memory Evolver for selective memory augmentation with a Retrieval Evolver for adaptive retrieval over the evolving memory. We then co-evolve the two Evolvers through alternating agentic reinforcement learning (Agentic RL), updating one while freezing the other. To steer this alternating evolution, Bottleneck-Aware Evolution Feedback (BEF) identifies whether the current bottleneck lies in memory or retrieval and directs optimization toward the more limiting side. Furthermore, VideoEvolve introduces Capability-Aware Evolution Feedback (CEF) to alleviate downstream feedback from over-specializing memory to a fixed set of training questions, shifting training toward underdeveloped yet learnable video capabilities. By integrating Agentic RL with BEF and CEF, VideoEvolve transforms downstream reasoning experience into transferable capability updates, providing a concrete path from static long-video systems toward experience-driven, self-improving multimodal intelligence. Extensive experiments on multiple long video understanding benchmarks demonstrate the effectiveness of VideoEvolve.

43. 【2610.10181】Argos: Adapt Rich Geometric Priors for Generalizable Online Scene-Change-Detection

链接:https://arxiv.org/abs/2610.10181

作者:Ruihan Xu,Jiae Yoon,Kaichen Zhou,Ue-Hwan Kim,Luca Carlone

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Robots operating, require reliable detection, dynamic environments require, environments require reliable, operating in dynamic

备注: More details on the project website: [this https URL](https://www.multyxu.com/argos/)

点击查看摘要

Abstract:Robots operating in dynamic environments require reliable detection of how their surroundings change over time. Existing learning-based methods largely rely on pairwise 2D image features, which struggle under large viewpoint changes and occlusions, are sensitive to noise, and show limited generalization across domains, while explicit 3D approaches typically require costly offline optimization. We show that the implicit 3D knowledge of Geometric Foundation Models (GFMs) provides a strong basis for addressing these limitations. We introduce Argos, which adapts GFM features for joint scene change detection and 3D reconstruction. To address data scarcity and take a step toward a foundation model for scene change detection, we introduce a large-scale benchmark comprising two synthetic datasets and one real-world dataset, and train jointly across diverse datasets to improve cross-domain generalization. We further introduce Argos-SLAM, a real-time system designed for robotics, which performs online change detection and change-aware 4D mapping. Across benchmarks, our framework substantially outperforms existing baselines, with gains of up to 42.01% in change IoU and 27.91% in F1, while supporting scalable deployment in changing real-world environments.

44. 【2610.10163】Beyond Anonymous Captions: Grounding Character Identity in Video Captioning and Question Answering

链接:https://arxiv.org/abs/2610.10163

作者:Anas Filali Razzouki,Killian Steunou,Khalil Guetari,Thomas Kling,Mounîm El-Yacoubi,Yannis Tevissen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Linking people appearance, Linking people, understanding video narratives, people appearance, appearance and actions

备注:

点击查看摘要

Abstract:Linking people's appearance and actions to character identities is essential for understanding video narratives. We present a framework for identity-aware video captioning and person-centric question answering that combines automatic character identification, explicit spatial grounding, and task-specific adaptation. Starting from LSMDC v2 movie clips, our pipeline matches detected faces to actor reference images, tracks characters across frames, and builds inputs with identity-linked bounding boxes. A strong vision-language model generates identity-aware captions and questions, which are manually verified and filtered to create a benchmark of 750 captioned clips and 3,000 person-centric questions. We study five grounding strategies combining textual coordinates with visual face or estimated person boxes across Video-MLLM families at roughly 2B, 4B, and 8B parameters and larger frontier models. Combining visual face boxes with textual coordinates yields the most consistent performance across scales and significantly improves overall performance over coordinates alone. Smaller models tend to over-assign known identities when the queried person is not grounded, while larger models better recognize such UNIDENTIFIED cases. We introduce BAC by LoRA fine-tuning Qwen models at 2B, 4B, and 8B scales on about 32K identity-aware captioned clips. Across all scales, BAC outperforms every other evaluated model family of comparable size. BAC-8B reaches 93.20\% overall QA accuracy, ranking behind only GPT-5.6 Sol among the frontier models evaluated in our study. Overall, explicitly communicating who is where, together with lightweight task-specific adaptation, substantially improves identity-aware video understanding without changing the underlying architecture. We release the benchmark, training data, code, and BAC checkpoints at this https URL.

45. 【2610.10160】BagDINO: Multi-View Baggage Re-Identification with DINOv3

链接:https://arxiv.org/abs/2610.10160

作者:Vita Santa Barletta,Danilo Caivano,Rebecca Margiotta,Massimiliano Morga,Davide Pio Posa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Mishandled checked baggage, current recovery workflows, directly support visual, support visual identification, checked baggage remains

备注: 7 pages, 4 figures, 3 tables, IEEE International Conference on Evolving and Adaptive Intelligent Systems 2026 (IEEE EAIS 2026)

点击查看摘要

Abstract:Mishandled checked baggage remains a recurrent issue in airport operations, and current recovery workflows still largely rely on tag-based tracking, which does not directly support visual identification when tag evidence is missing or unavailable. This paper investigates baggage re-identification as an instance-level retrieval problem in a multi-camera setting, leveraging DINOv3 foundation-model representations to match a query image against a gallery of registered baggage images. A Torchreid-style BNNeck re-identification head is placed on top of a DINOv3 backbone, and parameter-efficient adaptation is performed via LoRA. Experiments are conducted on the MVB benchmark using a progressive study that compares a fully frozen backbone against LoRA and fine-tuning strategies. Results indicate that parameter-efficient adaptation of foundation-model features provides an effective and stable approach for multi-view baggage re-identification under limited training data.

46. 【2610.10156】HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding

链接:https://arxiv.org/abs/2610.10156

作者:Leon Mayer,Lucas Luttner,Patrick Godau,Kai Fritzsche,Annika Reinke,Leonie Boland,Jule Brandt,Janne Heinecke,Chloe K. Nobuhara,Niklas Holzwarth,Evangelia Christodoulou,Marcel Knopp,Dominik Michael,Pascale Piermarco,Saliq Neyaz,Korhan Derin Özarslan,Jakob Hennighausen,Carlos Aumente-Maestro,Tim Rädsch,Dheeraj Baji,Peter Maximilian Full,Finn Aichholz,Justus Veit Erpenbeck,Linus Finn Schott,Bastian Winkelhausen,Claas de Boer,Bianca Güttner,Anneli Hummel,Gregor Just,Max Kirchner,Chenyang Li,Rozenn Raffaut,Ariel Rodriguez,Danush Kumar Venkatesh,Kevin Wang,Jinjing Xu,Mona Sheikh Zeinoddin,Salman Khan,Thomas M. Pausch,Stefanie Speidel,Danail Stoyanov,Daniel A. Hashimoto,Fiona R. Kolbinger,Thomas G. Weiser,Lena Maier-Hein

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, advances in Vision-Language, led to rapid, rapid progress, wide range

备注: 28 pages, 9 figures, 5 tables. Code: [this https URL](https://github.com/IMSY-DKFZ/orena-focus)

点击查看摘要

Abstract:Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, existing evaluations largely focus on short-term reasoning, failing to assess a critical capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of Foreign Object Contextual Understanding in Surgery. Built on a dataset of Heidelberg Colorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answering (VQA) pairs covering five core capabilities: object recognition, temporal grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behavior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experiments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particularly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capable of reliable, temporally consistent reasoning over hours-long videos.

47. 【2610.10133】HarnessIR: Harnessing Multimodal Foundation Models for Universal Real-World Image Restoration

链接:https://arxiv.org/abs/2610.10133

作者:Xiangtao Kong,Shuaizheng Liu,Rongyuan Wu,Lingchen Sun,Zhengqiang Zhang,Jinxin Zhao,Yuhui Wu,Lei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:low-quality images suffer, Real-world low-quality images, atmospheric effects, complex mixed degradations, model real-world image

备注:

点击查看摘要

Abstract:Real-world low-quality images suffer from complex mixed degradations, including but not limited to noise, blur, atmospheric effects, etc. Recent agentic methods usually model real-world image restoration (Real-IR) as a sequential tool calling problem over task-specific single-degradation restoration models. This paradigm, however, is fundamentally limited because complex real-world degradations cannot be cleanly undone degradation by degradation, and the tool used for task-specific models caps the capability of the agent system. In this work, we present HarnessIR, an agentic framework for Real-IR by harnessing a multimodal foundation model (MFM) as the executor. HarnessIR consists of five stages: perception and diagnosis, on-demand tool invocation, prompt composition, execution, and verification-driven refinement. Unlike prior agentic Real-IR methods that rely on tool chains assembled from task-specific models, HarnessIR feeds the restoration requirements, the perceptual diagnosis, and the evidence into an MFM that performs restoration in a single pass, followed by verification stages to determine whether the result warrants further processing. Under our harness, off-the-shelf MFMs handle restoration tasks remarkably well, achieving state-of-the-art results on the widely used MiO100 synthetic benchmark. More importantly, by exploiting the strong generalization ability of MFMs, HarnessIR delivers compelling restoration quality on challenging real-world scenes where previous agentic IR systems often struggle. Codes is available at this https URL.

48. 【2610.10125】Lifelong small-object navigation in changing object layouts: a benchmark and method

链接:https://arxiv.org/abs/2610.10125

作者:Jiagan Huang,Zikun Zhou,Zijian Ni,Hongpeng Wang,Guangming Lu,Jun Yu,Wenjie Pei

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Household robots, Changing Object Layouts, tools and toys, continually navigate, Household

备注:

点击查看摘要

Abstract:Household robots need to continually navigate to different objects in the same environment, many of which are small and portable, such as tools and toys. Their small visual footprint and frequent occlusion make reliable observation difficult, and they may be moved by people without the robot observing the changes. We formulate this challenging task as Lifelong Small-object Navigation in Changing Object Layouts (LiSoNav-COL). Agents must seek suitable viewpoints for reliable observation, accumulate and reuse scene knowledge to efficiently locate subsequent targets, and update outdated memory after object relocation. To eliminate the need for prior scene scanning, we also require agents to start navigation with empty scene memory. Although practical, this task still lacks benchmarks designed around its defining assumptions. To bridge this gap, we introduce LiSoNav-Eval, a dedicated benchmark spanning 28 indoor scenes with 45 small-object categories. Its lifelong navigation sequences include both unchanged and relocated targets to evaluate memory reuse and adaptation to object relocation. To address this challenging task, we propose a navigation method based on multi-view Inspection with Viewpoint-Anchored Memory, dubbed IVAM-Nav. IVAM-Nav actively observes supporting surfaces from complementary viewpoints for reliable small-object perception and anchors the resulting memory to their observation viewpoints, supporting relational memory reuse and revalidation under similar viewing conditions. Extensive experiments on LiSoNav-Eval demonstrate favorable performance of IVAM-Nav against representative methods. Benchmark analyses also show that smaller objects, larger environments, and longer relocation distances pose greater challenges. The dataset and code are available here.

49. 【2610.10116】A Probabilistic Perspective on Wasserstein-Based Evidential Uncertainty for Out-of-Distribution Segmentation

链接:https://arxiv.org/abs/2610.10116

作者:Arnold Brosch,Abdelrahman Eldesokey,Michael Felsberg,Kira Maag

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Semantic segmentation networks, segmentation networks operate, OOD objects requires, Monte Carlo dropout, identifying OOD objects

备注: 20 pages, 6 images, 3 figures

点击查看摘要

Abstract:Semantic segmentation networks operate on a fixed set of classes and therefore fail when out-of-distribution (OOD) objects appear during deployment, a critical limitation for safety-critical applications such as autonomous driving. Reliably identifying OOD objects requires well-calibrated epistemic uncertainty, yet common softmax-based confidence scores remain overconfident, while Bayesian alternatives such as Monte Carlo dropout or deep ensembles require costly repeated forward passes. Evidential Deep Learning (EDL) offers an efficient alternative by modeling class probabilities as a Dirichlet distribution learned from a single deterministic forward pass. Existing EDL formulations rely on Euclidean objectives that push predictions towards the simplex vertices, encouraging overconfidence rather than preserving uncertainty for unfamiliar inputs. We instead employ Wasserstein-based objectives, which respect the geometry of the probability simplex, and study the influence of the Wasserstein order on segmentation accuracy and OOD detection within a unified evidential framework. We evaluate this framework on a convolutional (DeepLabV3+) and a transformer-based (SegFormer) architecture on the SegmentMeIfYouCan benchmark, including LostAndFound, RoadObstacle21, RoadAnomaly21, and Fishyscapes. Our results show the optimal Wasserstein order is architecture-dependent: second-order objectives dominate on the convolutional backbone, third-order objectives on the transformer backbone, and our framework surpasses comparable baselines on most metrics, with a single deterministic forward pass.

50. 【2610.10090】mporal Residual Bottleneck for Robust Asynchronous Collaborative Perception

链接:https://arxiv.org/abs/2610.10090

作者:Melih Yazgan,Ahmed Abouelazm,J. Marius Zöllner

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:shared features arrive, features arrive stale, autonomous vehicles, Collaborative perception extends, extends the sensing

备注: Accepted to ACCV 2026

点击查看摘要

Abstract:Collaborative perception extends the sensing range of autonomous vehicles, but its performance degrades when shared features arrive stale or incomplete. Most latency-robust methods compensate delayed collaborator features through flow-guided alignment or direct feature transport. In this work, we formulate asynchronous collaborative perception as temporal residual prediction. Our Temporal Residual Bottleneck keeps a deterministic pose-warped collaborator feature as a conservative anchor and uses a $\Delta t$-conditioned xLSTM to extract residual temporal evidence from the available history. A detector-facing residual bottleneck then applies only gated, regularized corrections before ego-side fusion, reducing the risk of overwriting reliable static structure when temporal correspondence is uncertain. Experiments on DAIR-V2X and OPV2V show that our method is especially effective under severe fixed/irregular delays and packet drops. On DAIR-V2X, the reported checkpoint trades a small amount of synchronized peak accuracy for better robustness under stronger communication degradation. Controlled diagnostics further indicate that direct feature transport has oracle headroom but can become unreliable when deployed without accurate correspondence. These results support temporal residual fusion as a practical alternative for asynchronous and incomplete collaborative perception. Code will be publicly released at this https URL.

51. 【2610.10066】From Pixel to Coding: Evaluating the Figure Reproduction Capabilities of MLLMs

链接:https://arxiv.org/abs/2610.10066

作者:Zijian Chen,Zhengyu Chen,Bohan Liang,Lirong Deng,Yushuo Zheng,Yanwei Jiang,Qi Jia,Kaiwei Zhang,Wenjun Zhang,Guangtao Zhai

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:demonstrated impressive capabilities, Large Language Models, Multimodal Large Language, Multimodal Large, Large Language

备注: 46 pages, 18 figures

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in both visual understanding and code generation. However, existing benchmarks typically evaluate these two modalities in isolation, lacking a dedicated assessment of their unification, i.e., how a model can perceive complex visual structures and synthesize them into precise, executable code. Moreover, current visual code generation benchmarks often rely on simplified layouts within single programming environments, falling short of evaluating true unified multimodal reasoning. To bridge this gap, we propose FigCodeBench, a comprehensive framework for rigorously evaluating MLLMs on figure reproduction, integrating multimodal comprehension and generation. We first design a systematic dataset construction pipeline, resulting in a total of 6,194 instances that cover 7 functional categories and 4 types of programming languages. We further categorize figure reproduction into three tiers with visual and code complexity modeling, specifically targeting complex structural reasoning, varying aspect ratios, and dense geometric constraints. We introduce a multi-dimensional evaluation protocol, encompassing visual fidelity and syntactic isomorphism, that aligns highly with the Mean Machine Opinion Score (MMOS) and human preferences. Based on our framework, we conducted extensive experiments on 24 widely used proprietary and open-source MLLMs (e.g., Gemini 3.1 Pro, GPT-5.4, and Kimi-K2.5), where we observed a universal, non-linear performance cliff across different programming languages and difficulty scenarios for all models, and gained several insights, such as the significant metric decline in rigid declarative languages.

52. 【2610.10047】AdSpark: A Large-Scale Dataset and Benchmark for Product-Centric Advertisement Video Generation

链接:https://arxiv.org/abs/2610.10047

作者:Zhifei Yang,Zhao Jiang,Keyang Lu,Honghe Zhu,Zheng Zhang,Jingjing Lv,Changping Peng,Ching Law,Zhen Xiao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:presenting selling points, create promotional videos, coherent multi-shot narratives, preserve fine-grained product, Product-centric advertisement video

备注:

点击查看摘要

Abstract:Product-centric advertisement video generation aims to create promotional videos that preserve fine-grained product identity while presenting selling points through coherent multi-shot narratives. However, this emerging task remains underexplored due to the lack of large-scale advertisement-specific datasets and comprehensive evaluation frameworks. To address this gap, we introduce \textbf{AdSpark}, a large-scale dataset and benchmark for product-centric advertisement video generation, based on data from a major e-commerce platform. \textit{AdSpark-300K} contains approximately 300K reference image--prompt--video triplets, comprising a real-world subset and a synthetic subset. Each sample provides structured advertisement annotations, including product identity annotations, selling-point descriptions, creative plans, and aligned audio scripts, enabling models to learn product preservation and advertisement-oriented visual storytelling. We further propose \textit{AdSpark-Bench}, a diagnostic benchmark that evaluates generated advertisements across six dimensions, including visual quality, product fidelity, instruction adherence, temporal coherence, audio alignment, and advertisement effectiveness. Based on AdSpark-Bench, we evaluate representative models, revealing key challenges in product preservation, multi-shot storytelling, and selling-point visualization. Experiments with AdSpark-300K-finetuned models further validate the effectiveness of our dataset. AdSpark provides a unified dataset and benchmark for future research, and we will release the dataset upon acceptance.

53. 【2610.10013】Scalable Patch-Level Self-Supervised Learning

链接:https://arxiv.org/abs/2610.10013

作者:Maximilian Seitzer,Gabriele Trivigno,Antonín Vobecký,Seungeun Yi,Maxime Oquab,Huy V. Vo,Oriane Siméoni,Piotr Bojanowski

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Self-supervised learning, produces powerful visual, scale produces powerful, SSL, produces powerful

备注:

点击查看摘要

Abstract:Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Taking a step back, we ask if we can design a high-performing, yet principled SSL algorithm. Starting from the multi-view assumption, stipulating that task-relevant content is captured by the information common to different views, we construct an information-theoretic objective decomposing into interpretable terms. This derivation yields JEM, a student-teacher method that learns by aligning corresponding patch representations across views, explicitly regularized by information and structure preservation losses. JEM trains stably from 300M to 7B parameters, and, to our knowledge, is the first latent-space patch-level method demonstrated at 7B scale. Across all scales, JEM reaches strong performance on both global and dense probing tasks, on segmentation benchmarks consistently surpassing the DINOv2 algorithm, an influential foundation for today's strongest visual SSL methods. Notably, at 7B parameters, it exceeds the performance of DINOv3 on panoptic segmentation, despite being trained on $12\times$ less data without refinement stages. These results demonstrate that we can indeed design an SSL algorithm that learns strong representations, is principled and stable.

54. 【2610.10012】Playing with Kruskal: algorithms for flat and hierarchical watershed cuts

链接:https://arxiv.org/abs/2610.10012

作者:Jean Cousty(LIGM),Laurent Najman(KUSTAR, LIGM),Benjamin Perret(LIGM),Deise Santana Maia(CRIStAL)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Minimum Spanning Tree, Spanning Tree, Minimum Spanning, well-known optimization problems, edge-weighted graphs

备注:

点击查看摘要

Abstract:In the framework of edge-weighted graphs, watersheds have proven to be linked to well-known optimization problems, as Minimum Spanning Tree, which allowed the design of efficient algorithms for computing (hierarchical) watershed segmentations. In the present article, after reviewing the literature related to watershed segmentation, we present a detailed end-to-end pipeline of algorithms to compute (hierarchical) watershed segmentations, starting from the computation of graph-based image representations, up to the computation of connected components of the final (hierarchical) segmentation. We consider the several variations of watersheds, including their supervised and unsupervised versions, and the various ways of computing seeds, to name a few. For the first time, we bring together all these watershed notions and algorithms in a compact and understandable way. We aim at providing a reference for those interested in employing and reimplementing the watershed segmentation framework for their task at hand.

55. 【2610.10003】Perceptually Aligned Evaluation of Style Transfer

链接:https://arxiv.org/abs/2610.10003

作者:Yang Deng,Eleftherios Ioannou,David Mould,Steve Maddock,Paul L. Rosin,Yu-Kun Lai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Style TRansfer Algorithms, Style transfer, Style transfer lacks, reliable evaluation standard, TRansfer Algorithms

备注:

点击查看摘要

Abstract:Style transfer lacks a reliable evaluation standard: ground truth is inherently ill-defined, and existing automatic metrics often fail to reflect human preference. This paper introduces ASTRA (Assessment of Style TRansfer Algorithms), an approach for automatic evaluation of style transfer algorithms; it contains two components, ASTRA-Data and ASTRA-Score. ASTRA-Data consists of a benchmark image set of content and style references, a collection of style transfer results generated on the benchmark set, and user study data capturing human judgements through a two-stage pairwise comparison protocol. From these annotations, we derive ranking-based ground truth for content preservation, style fidelity, and overall preference. Based on ASTRA-Data, we construct ASTRA-Score, a learnt evaluator that predicts preference-aligned scores from content-style-stylization image triplets, enabling automatic and scalable evaluation of new models applied to the benchmark set. Experimental results demonstrate that ASTRA-Score achieves substantially higher correlation with human rankings compared to prior metrics. Overall, ASTRA establishes a robust mechanism for standardised evaluation of style transfer methods.

56. 【2610.09952】MSU Team at the Explainable Deepfake Detection Challenge 2026: Grounded Artifact Evidence for Deepfake Detection

链接:https://arxiv.org/abs/2610.09952

作者:Artem Filippov,Aleksandr Gushchin,Kirill Koltsov,Dmitriy Vatolin,Anastasia Antsiferova

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, images highly realistic, manipulated images highly, Explainable Deepfake Detection, highly realistic

备注:

点击查看摘要

Abstract:Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.

57. 【2610.09941】Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions

链接:https://arxiv.org/abs/2610.09941

作者:Bojun Yang,Haochen Zhou,Zhifang Zhang,Haobo Wang,Songze Li,Lei Feng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large vision-language models, Large vision-language, safety-critical applications, increasingly deployed, deployed in safety-critical

备注: 25 pages, 9 figures, 14 tables

点击查看摘要

Abstract:Large vision-language models (LVLMs) are increasingly deployed in safety-critical applications, yet they remain vulnerable to backdoor attacks. Defending against such attacks remains costly, as existing methods require either extensive retraining on clean data or per-query intervention at inference time. To address this limitation, we propose OrthoPurify, a more efficient method to purify backdoored model weights via one-step orthogonal projection. Specifically, through structural analysis of backdoor weight updates, we find that the backdoor is encoded by diverting a small number of weight update directions from task adaptation to backdoor shortcut encoding, a phenomenon we term direction hijacking. However, identifying these hijacked directions requires a benign reference model, which is typically inaccessible to the defender. We show that a pseudo-benign model, obtained by fine-tuning the pretrained weights on only a small set of clean samples, provides a sufficient approximation, as the dominant update directions stabilize within the first few gradient steps. OrthoPurify uses this pseudo-benign reference to isolate the hijacked directions and removes them through a single projection on the weight update. Extensive experiments show that OrthoPurify reduces the attack success rate to near zero while preserving the original performance across diverse benchmarks, without retraining the backdoored model or introducing inference-time overhead. Our code is publicly available at this https URL.

58. 【2610.09940】Juno: Taming Predictive Latents for Vision-Language-Action Models

链接:https://arxiv.org/abs/2610.09940

作者:Yuchen Zhu,Chenyi Xu,Yulin Zhang,Gang Xu,Wentao Zhu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Joint-embedding predictive architectures, Joint-embedding predictive, predict masked, offering a natural, natural source

备注: Project Page: [this https URL](https://juno-policy.github.io/)

点击查看摘要

Abstract:Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from $60.9\%$ to $68.5\%$ over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches $72.7\%$; on a real robot, it retains $70\%$--$75\%$ success under background, height, and object shifts where the base policy collapses to $0\%$.

59. 【2610.09928】Do Generative Priors Align with Human Naturalness Perception?

链接:https://arxiv.org/abs/2610.09928

作者:Taiki Fukiage

类目:Computer Vision and Pattern Recognition (cs.CV); Neurons and Cognition (q-bio.NC)

关键词:native priors reflect, regularities governing human, reflect the regularities, regularities governing, open question

备注: 62 pages, 35 figures, including appendices

点击查看摘要

Abstract:Visual generative models are trained to capture the probability distributions of natural images, yet whether their native priors reflect the regularities governing human perception of image naturalness remains an open question. Here, we probe these priors through native prediction errors across 25 open image and video generators. Because raw single-image losses are dominated by scene content and visual complexity, we evaluate directional loss differences using content-preserving, paired relational interventions that selectively disrupt facial configurations or physical illumination consistency while limiting changes in low-level image statistics. Across both domains, these loss differences reproduce human-like selective sensitivities and tolerances, capturing the classic Thatcher effect on faces and shape-dependent responses to illumination inconsistencies. Notably, these loss differences reliably track continuous gradations of human naturalness judgments across individual stimulus pairs (peaking at $r = .84$ on faces and $.64$ on physical scenes) and retain unique human-aligned signals even after controlling for feature distances from frozen vision encoders and standard image quality metrics. We also find that while overall sensitivity to these violations broadly covaries with human alignment across models, the two systematically decouple along denoising schedules, with alignment peaking earlier than sensitivity, revealing that human-like naturalness judgments dissociate from generic violation detection. Together, these findings demonstrate that learning visual distributions yields generative loss landscapes that capture distinct aspects of human naturalness perception.

60. 【2610.09920】Inverting Multi-Vector Visual Document Indices

链接:https://arxiv.org/abs/2610.09920

作者:Zhuchenyang Liu,Yao Zhang,Yu Xiao

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Prevailing multi-vector visual, vector databases run, thousand patch vectors, Prevailing multi-vector, page

备注: 30 pages. Under review

点击查看摘要

Abstract:Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.

61. 【2610.09907】FedSSMCoOp: SSM Encoders for light-weight Federated Prompt Learning for Few-shot Classification

链接:https://arxiv.org/abs/2610.09907

作者:Ankita Das,Ambarish Parthasarathy,Sumohana S. Channappayya,C. Krishna Mohan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:complementary information contained, downstream vision tasks, shown strong performance, shown strong, wide range

备注:

点击查看摘要

Abstract:Vision-Language Models (VLMs) have shown strong performance across a wide range of downstream vision tasks, thanks to the complementary information contained in the respective domains. Despite the performance gains, most of these approaches rely on aligning these domains using the cosine similarity metric, which fails to capture token-level structure and cross-modal interactions prior to the classification stage. This is especially critical in biomedical applications under federated constraints, where data sharing is restricted, labeled data is scarce at each site, and it differs widely across institutions, leading to substantial statistical heterogeneity. To overcome this issue, we propose FedSSMCoOp, a federated few-shot image classification framework that enables multimodal learning while preserving data privacy. With the help of the SSM-based Vision Mamba and Cross Mamba blocks, and by optimizing only the soft-prompt and communication-prompt updates in the federated setting, the framework prioritizes both computation and performance. Importantly, this eliminates the need to use an external Large Language Model (LLM) for feature alignment. The framework is further trained and evaluated on various biomedical image datasets, and its performance is assessed. The proposed framework delivers stable performance relative to the baselines and is, on average, 1.96 times lighter. The corresponding script will be made available soon.

62. 【2610.09875】SANet: Selective Attention Network for Infrared Small Target Detection

链接:https://arxiv.org/abs/2610.09875

作者:Yingmei Zhang,Wangtao Bao,Qin Xiao,Yong Yang,Weiguo Wan,Yitao Luo,Xueting Zou,Lei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Infrared small target, locate dim targets, small target detection, target detection aims, Infrared small

备注:

点击查看摘要

Abstract:Infrared small target detection aims to accurately identify and locate dim targets in complex backgrounds and supports applications such as maritime surveillance and military search and rescue. However, the small size and weak contrast of infrared targets make it difficult to balance detection accuracy and false alarms. This paper proposes a selective attention network (SANet) for infrared small target detection. A dual-path semantic-aware module combines standard and pinwheel-shaped convolutions to preserve local spatial consistency and capture broader contextual information. Spatial and channel attention further refine the features and improve target-background discrimination. To address the limitations of static skip connections in U-Net, a selective attention fusion module adaptively integrates features across scales using spatially varying weights. It selectively enhances salient regions and improves discrimination between true targets and false alarms. Experiments on three public benchmarks, NUAA-SIRST, IRSTD-1K, and NUDT-SIRST, show that SANet achieves competitive performance in intersection over union (IoU), normalized IoU, detection probability, and false alarm rate. Its IoU exceeds that of the second-best method by 1.93, 4.32, and 2.21 percentage points, respectively. These results support the effectiveness of SANet in dim-target perception, discriminative feature representation, and background suppression.

63. 【2610.09873】Bringing BNNs to Fast Event Processing

链接:https://arxiv.org/abs/2610.09873

作者:Paul Longour,Julien Moreau,Franck Davoine

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Binary Neural Networks, Neural Networks, substantially reducing model, resource constrained devices, reducing model size

备注: 19 pages, 3 figures, 8 tables. Accepted by NeVi Workshop at ECCV 2026

点击查看摘要

Abstract:Binary Neural Networks (BNNs) enable efficient deep learning deployment on resource constrained devices with weights and activations compressed to one bit, substantially reducing model size and inference cost. Event cameras offer complementary advantages, including low latency, high dynamic range, and low power consumption, by capturing asynchronous streams of events rather than dense image frames. Despite their shared emphasis on efficiency, the combination of these technologies remains largely unexplored. This work aims at adapting and evaluating modern deep BNN architectures on event data. We also show that cross-modal pretraining from RGB data can improve the classification accuracy of BNNs on neuromorphic datasets. We introduce the Polar-wise Binary Event Volume (PBEV), a binary representation that enables event-camera data to be processed directly by BNNs and represents a step toward fully binarized event-based vision systems. Best evaluated BNN on N-Caltech101 classification benchmarks shows 90.58% accuracy with 7.5x less operations than their full-precision counterparts.

64. 【2610.09863】Global Average Precision for Representation Learning

链接:https://arxiv.org/abs/2610.09863

作者:Bill Psomas,Mohammad Mahdi,Michalis Thomas,Danda Pani Paudel,Giorgos Tolias,Giorgos Kordopatis-Zilos

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Standard information retrieval, information retrieval metrics, Average Precision, Global Average Precision, assess performance

备注:

点击查看摘要

Abstract:Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, $k$NN and zero-shot classification accuracy.

65. 【2610.09860】DeepTopoClustering: Unsupervised Derivation of Surface Process Taxonomy from 4D Point Clouds for Topographic Monitoring

链接:https://arxiv.org/abs/2610.09860

作者:Jiapan Wang,Daan Hulskemper,Mathilde Letard,Roderik Lindenbergh,Katharina Anders

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:point clouds acquired, permanent laser scanning, enable accurate high-frequency, accurate high-frequency monitoring, point clouds

备注: Submitted to ISPRS Journal of Photogrammetry and Remote Sensing

点击查看摘要

Abstract:4D point clouds acquired by permanent laser scanning (PLS) enable accurate high-frequency monitoring of surface change in dynamic topographic environments. However, existing methods remain limited in organizing detected surface activities into meaningful process types. We propose DeepTopoClustering (DTC), an unsupervised framework for deriving a hierarchical process taxonomy from object-based surface activities, so-called 4D objects-by-change (4D-OBCs). We transform each 4D-OBC into a GeoMorphogram, a distributional sequence representing the temporal evolution of topographic change within a spatially bounded surface activity. A convolutional autoencoder learns latent embeddings from GeoMorphograms, which are jointly optimized using a hierarchical deep clustering objective to organize surface activities into a hierarchy. We evaluate the learned hierarchy using expert annotations on two 4D datasets of sandy beach sites and their combination. DTC with GeoMorphograms achieves the highest agreement with expert judgment at the taxonomy level comprising eight major process types ($F_1=0.78$, match accuracy $=0.92$), outperforming dimensionality reduction and conventional flat clustering. The learned taxonomy separates major erosion- and deposition-dominated activities and distinguishes finer subtypes based on change magnitude, duration, compactness, and temporal evolution. DTC thus provides a scalable and interpretable route from 4D change detection to a data-driven, expert-supported surface process taxonomy, advancing automated knowledge derivation for understanding surface dynamics in topographic monitoring.

66. 【2610.09853】DeltaSplat: Iterative Gaussian Refinement for Pose-Free Feed-Forward 3D Gaussian Splatting

链接:https://arxiv.org/abs/2610.09853

作者:Chanung Park,Seunghyeon Song,Joo Chan Lee,Eunbyung Park,Jong Hwan Ko

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:single network pass, Gaussian Splatting, reconstructs a scene, scene from sparse, unposed images

备注: 11 pages, 6 figures

点击查看摘要

Abstract:Pose-free feed-forward 3D Gaussian Splatting (3DGS) reconstructs a scene from sparse, unposed images in a single network pass, removing the need for camera calibration and per-scene optimization. However, camera estimation errors propagate into the predicted Gaussians and compound the geometric and photometric inaccuracies of single-pass prediction. To correct these errors, we introduce DeltaSplat, a lightweight Gaussian refinement module for pose-free feed-forward 3DGS. It iteratively renders the current Gaussians at the input context views and predicts per-Gaussian updates from the resulting residuals. A 2D residual alone, however, underdetermines the 3D correction. DeltaSplat therefore conditions each update on per-pixel Plücker rays and rendered depth as a soft geometric prior. A dual-branch convolutional mixer efficiently encodes these inputs, and per-attribute heads decode the fused features into position, opacity, and color updates. The module adds only ~2.2% parameters to the backbone and remains fully feed-forward at inference. On DL3DV, DeltaSplat reaches 26.64 dB PSNR in the pose-free setting, improving its state-of-the-art backbone by 1.75 dB and surpassing even baselines supplied with ground-truth cameras; consistent gains hold across 6-24 views and all camera regimes.

67. 【2610.09849】Hard, Yet Reducible: Controlled Forward Transfer for Synthetic Degradation Curation

链接:https://arxiv.org/abs/2610.09849

作者:Chunming He,Kailai Zhou,Jiaming Zuo,Hanqi Liu,Fengyang Xiao,Youwei Pang,Xiaofeng Liu,Weisi Lin,Xiaoqi Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dense prediction requires, Selecting synthetic degradations, Selecting synthetic, dense prediction, prediction requires

备注: 17 pages, 4 figures, 9 tables

点击查看摘要

Abstract:Selecting synthetic degradations for dense prediction requires an estimate of their training utility, the generalization gain they bring under a finite training budget. Clean and degraded twins share content and labels, suggesting a score based on how much short training reduces the excess error caused by degradation. However, this gap can also shrink when clean performance deteriorates. Measuring the improvement on degraded images alone avoids that confound, but it still credits progress that the same amount of clean training would have produced. We propose the \textbf{controlled Reducible Degradation Gap} (cRDG) for regions defined by degradation type and severity. From a common checkpoint, cRDG runs two budget-matched probes that differ only in one augmentation slot, which holds either a synthetic degradation or a clean augmentation. The score is the gain on held-out degraded images relative to the clean-control probe. Clean harm is a separate feasibility constraint. cRDG reveals a correctable severity band in which training on the degradation yields high controlled gain under the available budget, and the band moves with the predictor, the starting checkpoint, and the training budget. \textbf{Curation of Reducible Bands} (\method) uses cRDG to select synthetic data without changing the predictor. On semantic segmentation and salient object detection, \method{} improves representative predictors under matched synthetic-data budgets and training schedules, extends to existing data-generation pipelines, and preserves clean performance. Code and supporting materials will be publicly released.

68. 【2610.09844】For Those Who Believe in Faithfulness: Optimizing the Area Under Insertion and Deletion Curves for Ranking Relative Feature Importance

链接:https://arxiv.org/abs/2610.09844

作者:Bjørn Leth Møller,Bulat Ibragimov,Christian Igel

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:machine learning models, explainable artificial intelligence, socially relevant tasks, relevant tasks requires, tasks requires effective

备注:

点击查看摘要

Abstract:The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect the relative importance of input features for a particular prediction. The quality of such maps is often assessed by measuring faithfulness based on the area under insertion and deletion curves, which measures changes in the model output as features are added and removed. In this study, we derive an objective function from this notion of faithfulness and a way to approximate its gradient. We establish the connection between insertion curves and top-$k$ feature selection, which leads to a loss function measuring the quality of attributions. Randomization of the loss allows us to efficiently approximate its gradient. To show the effectiveness of the general approach, we combine the loss function with the neural explanation mask framework. The resulting method, termed Ra-NEM, can be used with any differentiable model without affecting the model's performance. Experiments demonstrate that Ra-NEM provides accurate attributions robustly and efficiently. Compared to other algorithms, the attributions have not only higher faithfulness but also perform well in terms of other XAI metrics. The high inference speed of Ra-NEM makes the method suitable for online applications. The code is available online: this https URL

69. 【2610.09841】ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

链接:https://arxiv.org/abs/2610.09841

作者:Arshia Hemmat,Amirhossein Vahidi,Amitis Shidani,Mohammad Vali Sanian,Hesam Asadollahzadeh,Aryan Yazdan Parast,Mohammad Lotfollahi

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:models fail predictably, scenes lose count, diffusion models fail, loses compositional structure, multi-object scenes lose

备注: Accepted at NeurIPS 2026. 26 pages, 4 figures

点击查看摘要

Abstract:Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA (Orthogonal Residual Compositional Alignment), aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones (DiT-B/2, DiT-L/2, U-ViT-L), ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.

70. 【2610.09830】MOTIF: Person-of-Interest Deepfake Detection Beyond 3DMM Coefficients

链接:https://arxiv.org/abs/2610.09830

作者:Giovanni Affatato,Sara Mandelli,Paolo Bestagini,Stefano Tubaro

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Video deepfakes targeting, specific individual, abundantly recorded, deepfakes targeting, targeting a specific

备注: 6 pages. Accepted at the 2026 IEEE International Workshop on Information Forensics and Security (WIFS)

点击查看摘要

Abstract:Video deepfakes targeting a specific individual, the Person-of-Interest (POI), are the most harmful ones, and, since a public figure is abundantly recorded, a detector can be built from genuine footage of that individual. Such detectors commonly describe a subject through a 3D Morphable Model (3DMM) and adopt its coefficients as a whole, so which part of that description carries the signal has never been measured. We dissect it, holding the encoder, the training corpus and the enrollment protocol fixed and varying only what the encoder observes. The groups of coefficients prove largely redundant, since the shape block alone recovers almost all the accuracy of the full vector, and their temporal evolution contributes a real but bounded amount. We further show that the dense surface the same fit returns, which these detectors discard, carries identity information that the coefficients do not, and that it helps precisely where they are weakest. We assemble the best configuration into MOTIF, a visual-only detector trained on real videos only, with no manipulated video and no POI-specific data. It improves on both state-of-the-art POI detectors in every dataset and manipulation of our benchmark and at two quality levels. Our experimental code will be released at this https URL.

71. 【2610.09823】UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

链接:https://arxiv.org/abs/2610.09823

作者:Deyuan Liu,Yihao Hu,Jingxuan Zhang,Xingying Li,Jun Xie,Jiacheng Liu,Jungang Li,Yu Huang,Xuanyi Liu,Yue Ding,Zecheng Wang,Lei Zhao,Mingda Wang,Zhenglin Cheng,Peng Sun,Tao Lin

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:reproduce long strings, requires image generators, generators to reproduce, reproduce long, Dense visual text

备注:

点击查看摘要

Abstract:Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each human-reviewed prompt supplies exact strings for four to twelve text regions, paired with structured references for their content, placement, and visual attributes. We use the Q-Judger vision-language model to assess each image against the complete reference, reporting text fidelity, text clarity, spatial quality, and scene quality. Across 24 model configurations, these dimensions reveal different strengths: Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings. Performance also varies with workload; Qwen-Image-2512's English composite falls from 86.50 at L1 to 42.86 at L3. Ten participants took part in human evaluation of the automatic scores. Repository: this https URL.

72. 【2610.09821】Efficient 3D Gaussian Head Avatars for Edge Devices

链接:https://arxiv.org/abs/2610.09821

作者:Umar Farooq,Jean-Yves Guillemaut,Adrian Hilton,Marco Volino

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains computationally expensive, representation remains computationally, Gaussian representation remains, avatars provide high-quality, Gaussian head

备注:

点击查看摘要

Abstract:Generative 3D Gaussian head avatars provide high-quality, efficient rendering, but synthesising the Gaussian representation remains computationally expensive, limiting deployment on resource-constrained and edge devices. We introduce an efficient generator architecture for unconditional 3D Gaussian head synthesis, based on a parameter-efficient synthesis block and depth-wise separable convolutions while retaining style-based conditioning. Our architecture reduces generator complexity without requiring model compression or quantisation. Compared with the baseline model, our approach reduces FLOPs by 94%, parameter count by 70%, and model size by 81%, while maintaining competitive generation quality. We further demonstrate practical CPU inference and browser-based execution on mobile devices using ONNX Runtime, enabling 3D Gaussian avatar synthesis without dedicated GPU hardware or application-specific software. In addition to conventional image-quality metrics, we evaluate multi-view consistency, training cost, and deployment performance. Code, trained models, and evaluation tools will be released publicly.

73. 【2610.09807】Concentration, Not Uncertainty: Why Targeted Synthetic Data Doesn't Help Camouflaged Object Detection

链接:https://arxiv.org/abs/2610.09807

作者:Akshat Dobhal,Sanjay Singh

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Camouflaged object detection, requires pixel-accurate masks, object detection requires, detection requires pixel-accurate, Camouflaged object

备注: 29 pages, 3 figures, 12 tables

点击查看摘要

Abstract:Camouflaged object detection requires pixel-accurate masks, but obtaining such annotations is slow and costly, making synthetic training images an attractive alternative. Under a fixed generation budget, however, it remains unclear which real-image regions to target for synthetic data generation. We study an uncertainty-guided generation strategy that clusters the unlabelled real images, identifies clusters on which the model is least certain, allocates synthetic generation toward those clusters, and iteratively retrains the model. Across 103 training runs, uncertainty-based targeting does not outperform random allocation. Five independent controls further show that this null result is not an artifact: targeted training sets are measurably different from random sets, but the difference is explained by concentrating the generation budget rather than by where uncertainty is concentrated, as every concentration rule we test reproduces the effect and, on boundary accuracy, so does aiming at the clusters the model was most certain about. Separately, we find substantial data contamination in CHAMELEON, with 50 of its 76 images duplicated from training data despite the standard overlap check reporting zero overlap. Together, these results show that, under a fixed synthetic-data budget, budget concentration, not uncertainty-based targeting, accounts for the observed training-set effects.

74. 【2610.09802】DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

链接:https://arxiv.org/abs/2610.09802

作者:Adam Pardyl,Siddhartha Gairola,Sukrut Rao,Adam Wróbel,Bartosz Zieliński,Bernt Schiele,Dawid Rymarczyk

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Concept-based vision models, Concept-based vision, vision models represent, intermediate layer, layer of human-inspectable

备注: Under review

点击查看摘要

Abstract:Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quantized attributes), a method that learns spatially grounded, discrete concept representations from a powerful frozen vision-only self-supervised backbone. It requires no class labels and no language supervision. Each image patch is assigned to exactly one concept from a learnable prototype dictionary, and only a sparse subset of concepts may activate per image. To capture how each concept varies across images (e.g., the type of a "wheel"), we learn continuous residuals alongside the concepts and then quantize them into discrete attributes. A spatial decoder reconstructs the backbone's representation from the concepts and attributes alone, so successful reconstruction means that the discrete representation preserves the backbone's information. We evaluate DisParQ across seven datasets, from general recognition (ImageNet, PartImageNet, Places) to fine-grained benchmarks (CUB, Cars, Dogs, Flowers). We show that DisParQ closely matches its frozen DINOv2 teacher on ImageNet linear probing (83.2% top-1), achieves higher concept consistency than language-aligned models, remains competitive on fine-grained recognition, and enables cross-category part-based retrieval.

75. 【2610.09800】Beyond Masks and Trajectories: Flow-Guided Latent Action Injection for Stable Surgical Video Generation

链接:https://arxiv.org/abs/2610.09800

作者:Tsz-Yui Qin,Siyu Zhou,Chi-Keung Tang,Yuxiang Nie,Shu Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:motion remains challenging, holds substantial potential, plausible motion remains, generation holds substantial, surgical videos

备注:

点击查看摘要

Abstract:Surgical video generation holds substantial potential for surgical education, simulation, and data augmentation, yet generating surgical videos with realistic and clinically plausible motion remains challenging. Most existing methods rely on auxiliary conditions, such as masks, trajectories, depth, or reference videos, to achieve visually plausible synthesis. Yet, these auxiliary conditions typically require additional manual annotation or specialized acquisition, making it difficult to scale such methods beyond small, curated datasets. This motivates the need for a reference-free architecture capable of generating high-quality surgical video without requiring auxiliary visual conditions at inference time. We propose FLAIR, a Flow-guided LatentAction Injection framework for Reference-free surgical video generation. FLAIR learns action priors from optical flow of real surgical videos, dynamically predicts corresponding latent action representation from an input prompt, and injects it into a frozen base model to generate surgical videos with improved action consistency. We further construct SurgActionClip-30K, the first large-scale surgical vision dataset comprising action-centric segmented clips and structured caption labels, addressing the persistent lack of fine-grained, action-centric surgical datasets. Lastly, we introduce SurgMetrics, the first surgical domain-specific evaluation metrics for quantifying the quality of generated surgical videos, addressing the persistent absence of clinically grounded evaluation standards in this domain. Extensive experiments demonstrate that FLAIR enables generating high-quality surgical videos using text-only inference without auxiliary conditions, and validation in SurgMetrics demonstrates its strength in alignment with human perception compared to traditional metrics.

76. 【2610.09791】Counterfactual Route Optimization for Gaussian Head Avatar Modeling

链接:https://arxiv.org/abs/2610.09791

作者:Shikun Zhang,Yong Li,Yiqun Wang,Qiuhong Ke,Cunjian Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires jointly optimizing, jointly optimizing multiple, modeling requires jointly, optimizing multiple objectives, Head avatar modeling

备注: 15 pages, 7 figures, 4 tables

点击查看摘要

Abstract:Head avatar modeling requires jointly optimizing multiple objectives with different dominant effects on geometry, appearance, and cross-view consistency. However, their relative effectiveness varies across training states, while existing pipelines typically rely on fixed loss weights or handcrafted stage-wise schedules. A central challenge is therefore to identify which optimization direction is more beneficial at each training state. We propose a counterfactual route optimization framework for Gaussian head avatar modeling, which characterizes state-dependent optimization preference from the realized effects of alternative updates rather than predefined heuristic weighting. Starting from the same training state, we perform short-horizon route-restricted lookahead over geometry, appearance, and joint update routes and evaluate their outcomes under a unified utility. The resulting counterfactual evidence is factorized into a geometry--appearance preference and a residual joint advantage, separately capturing the relative preference between individual update directions and the additional benefit of coordinated optimization. We further amortize this offline evidence into a lightweight controller that directly estimates the current optimization preference and applies bounded modulation to the training objectives during full avatar optimization. Experiments on the NeRSemble dataset validate the effectiveness of the proposed design, consistently outperforming existing methods while preserving clearer local facial structures and finer details.

77. 【2610.09785】UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map

链接:https://arxiv.org/abs/2610.09785

作者:Keke Yang,Erqi Wang,Sainan Guan,Hongliang Ren

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:enable autonomous ultrasound, autonomous ultrasound scanning, enable autonomous, scanning by predicting, predicting the outcomes

备注:

点击查看摘要

Abstract:World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action--observation relationship typically relies on synchronized video--pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound's cross-sectional sampling geometry. We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action--video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29\% and 38\%, respectively, compared with visual servoing. Project Page: this https URL.

78. 【2610.09784】CIRSeg: Coarse-to-Fine Intensity-Robust Liver Segmentation with Source-Free Continual Test-Time Adaptation

链接:https://arxiv.org/abs/2610.09784

作者:Ruoshi Xu,Mingqi Gao,Shengda Luo,Jingkun Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:quantitative hepatic assessment, longitudinal disease monitoring, Reliable liver segmentation, treatment planning, Reliable liver

备注: Accepted at the CARE 2026 Workshop at MICCAI 2026

点击查看摘要

Abstract:Reliable liver segmentation in contrast-enhanced MRI is essential for quantitative hepatic assessment, treatment planning, and longitudinal disease monitoring. However, limited annotated data and scanner- or vendor-dependent intensity variations can cause overfitting and poor generalization to unseen acquisition domains. Moreover, simultaneously achieving robust global localization and precise boundary delineation remains challenging, while predictions may contain isolated false-positive regions outside the main liver component. To address these challenges, we propose CIRSeg, a coarse-to-fine, intensity-robust liver segmentation framework based on nnU-Netv2. CIRSeg combines 3D CutMix with stochastic intensity transfer using either Nyul augmentation or histogram matching to improve robustness to heterogeneous MRI intensities. Its cascaded architecture decouples low-resolution anatomical localization from full-resolution boundary refinement. At inference, source-free test-time adaptation based on confidence-filtered predictions and probability-prior regularization further improves robustness to out-of-distribution inputs. As a final deterministic post-processing step, largest connected component filtering removes isolated false-positive regions. On the CARE 2026 test set, CIRSeg achieves Dice scores of 97.13\% and 97.93\% on the in-domain and unseen-domain subsets, with corresponding HD95 values of 20.18 mm and 11.30 mm, respectively. These results demonstrate consistently accurate segmentation across both in-domain and unseen acquisition settings. The code is available at this https URL

79. 【2610.09780】Relational Abstractions for Spatial Reasoning with Diffusion Models

链接:https://arxiv.org/abs/2610.09780

作者:Ana Ezquerro,Ozan Özdenizci

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reliably satisfy structured, satisfy structured spatial, Diffusion models excel, Diffusion models, remain limited

备注:

点击查看摘要

Abstract:Diffusion models excel at image synthesis, but they remain limited in their ability to reliably satisfy structured spatial reasoning constraints. In conditional data distribution modeling tasks with implicit logical structure, such as puzzles defined by visible clues paired with consistent solutions, state-of-the-art generative models tend to approximate pixel-space distributions without learning the underlying logical rules required for inference. To address this limitation, we present a novel framework for spatial reasoning with diffusion models that leverages unsupervised object discovery and abstractions of object relations. We show that the relational knowledge derived from object-centric representations enriches diffusion models with structural primitives, allowing them to effectively guide the generative representation space during both training and inference, and enabling conditional image generation that satisfies reasoning constraints. Additionally, we introduce a large-scale generative spatial reasoning benchmark with four datasets inspired by human-solvable puzzles. Our results show that relational abstractions significantly improve reasoning capabilities of diffusion models on a variety of complex reasoning tasks, while enabling robust generalization in out-of-distribution settings.

80. 【2610.09761】PARC-Loc: Text-to-Point-Cloud Localization with Partial Assignment and Relational Consistency

链接:https://arxiv.org/abs/2610.09761

作者:Shengkai Ma,Zhenyu Hou,Weihua Cao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:estimates a position, surrounding objects, localization estimates, submap, PARC

备注:

点击查看摘要

Abstract:Text-to-point-cloud localization estimates a position in a city-scale 3D map from descriptions of surrounding objects. Existing coarse-to-fine methods retrieve submaps using aggregate learned compatibility and then localize within a selected submap. However, repetitive or similar urban objects can inflate the embedding similarity between the query and multiple submaps, even when the instance layout within a submap violates the query description. Meanwhile, query-relevant instances often span submap boundaries, leaving the retrieved submap with incomplete contextual evidence. We term these failure modes layout-inconsistent aliasing and boundary evidence incompleteness, respectively. To address them, we propose PARC-Loc, a coarse-to-fine localization framework built on Partial Assignment with Relational Consistency (PARC). PARC jointly models hint-object compatibility and pairwise spatial relations, allowing unmatched elements while favoring assignments consistent with the queried layout. At the coarse stage, its candidate-level assessment complements neural similarity for layout-consistent submap selection. At the fine stage, the context is expanded with query-relevant instances from adjacent submaps, while PARC yields object-level matching weights that guide cross-modal attention. Extensive experiments on KITTI360Pose and CityLoc show that PARC-Loc outperforms conventional coarse-to-fine baselines. On KITTI360Pose, our method improves Top-1 localization recall at 5 m from 0.50 to 0.67, achieving a 34% relative gain over the strongest baseline.

81. 【2610.09755】Diffusion-Generated Image Watermarking: A Two-Axis Taxonomy and Three Protocol-Bounded Case Studies

链接:https://arxiv.org/abs/2610.09755

作者:Sung Ju Lee,Nam Ik Cho

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires balancing provenance, balancing provenance signals, images requires balancing, computational cost, requires balancing

备注: 17 pages, 2 figures. Accepted to the non-archival track of the ECCV 2026 LifeGenIP Workshop. English translation with partial reorganization of our article in Journal of Broadcast Engineering 31(4), 687-699 (2026)

点击查看摘要

Abstract:Watermarking diffusion-generated images requires balancing provenance signals with image quality, robustness, and computational cost. This work organizes methods along two axes: insertion mechanism and primary signal-bearing representation, and formalizes a representative $z_T$-Fourier pipeline for verification and identification. We then use the taxonomy to structure three protocol-bounded case studies. The first examines associations among frequency integrity, detection, quality, and cropping behavior. The second revisits persistence under seed-linked and seed-independent editing and formulates a scoped Semantic Imprinting Hypothesis without claiming a localized carrier or causal mechanism. The third studies single-shot VAE-latent phase modulation, including its efficiency, regeneration robustness, and robustness--quality operating points. Finally, we separate four content-level attack families from model/pipeline adaptation, propose corresponding evaluation protocols and testable conjectures for parameter-tuning threats, and identify additional temporal extensions for video. These analyses do not establish a universal ranking; instead, they provide a framework for matched, protocol-aware comparisons of watermarking systems for diffusion-generated images.

82. 【2610.09746】Flow-of-Thought: A Framework for Visual Reasoning

链接:https://arxiv.org/abs/2610.09746

作者:Mariia Baidachna,Nicolas Pugeault

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, progress Large Language, mind eye, essential aspect, mimicking mental imagery

备注: NeurIPS WiML 2026 version: OpenReview version: [this https URL](https://openreview.net/forum?id=QBcqVOacYO)

点击查看摘要

Abstract:Mental imagery, ``seeing with the mind's eye'' is an essential aspect of human cognition. Despite rapid progress Large Language Models (LLMs) and Vision Transformers (ViTs) still underperform on tasks requiring spatial understanding. To address this, we introduce Flow-of-Thought (FoT), a framework that integrates the generation of visual sketches as intermediate reasoning steps, mimicking mental imagery in humans. We train coordinate-aware trajectory flow fields on $SO(2)$ group orbits and cumulative shortest paths, then freeze the learned dynamics; same vs. different decisions compare competing generative hypotheses using foreground-weighted reconstruction energy. On locked tests FoT reaches 100.0% accuracy on Tetris and 99.0% on colored shapes. Under frozen transfer, the orbit-trained 2D flow improves over its endpoint-only control on BLINK Multi-view (72.2% vs. 63.9% on 133 public validation pairs), supporting continuous visual traces as an effective and interpretable representation for spatial reasoning in some out-of-distribution settings.

83. 【2610.09742】SoccerNet-FoulRet: Retrieving Semantically Similar Soccer Foul Videos

链接:https://arxiv.org/abs/2610.09742

作者:Jacobus Arthur,Ahmad Sait,Batool Hani,Merey Ramazanova,Jan Held,Marc Van Droogenbroeck,Bernard Ghanem,Anthony Cioppa,Silvio Giancola

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:similar past cases, professional soccer remain, soccer remain inconsistent, decisions in professional, professional soccer

备注: ACCV 2026

点击查看摘要

Abstract:Refereeing decisions in professional soccer remain inconsistent because referees cannot easily compare a contentious foul against similar past cases. We cast this as a retrieval problem and introduce SoccerNet-FoulRet, the first benchmark for semantic foul retrieval. Given a query foul, the task is to retrieve past fouls judged to be relevant precedents, regardless of camera angle, teams, or appearance. This differs from prior video-to-video retrieval, which matches clips by visual similarity or a shared event. Here, relevance is defined by refereeing interpretation. We build the benchmark from the SoccerNet-MVFoul dataset and evaluate retrieval ability of zero-shot video and vision-language embedders together with a task-specific fine-tuned baseline on 693 human-verified queries and category-relevance labels. Semantic foul retrieval remains challenging. The strongest zero-shot model achieves under 5% HitRate@10 on human-verified precedents, while category-supervised fine-tuning improves category relevance but transfers only modestly to precedent retrieval. We release SoccerNet-FoulRet to establish semantic foul retrieval as an open problem: this https URL.

84. 【2610.09726】Beyond Group Splits: Specimen-Level Cross-Validation and Visual Attribution for Remaining-Shelf-Life Regression in Climacteric Fruit

链接:https://arxiv.org/abs/2610.09726

作者:Rovhona Mudau,Jean Frederic Isingizwe Nturambirwe,Clement Nthambazale Nyirenda

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Estimating remaining shelf, remaining shelf life, provide affordable decision, affect reported performance, Estimating remaining

备注: 7 pages, 1 figure, 4 tables. Accepted for physical presentation at the 10th IEEE Conference on Information Communications Technology and Society (ICTAS 2026), 14-16 October 2026, Durban, South Africa

点击查看摘要

Abstract:Estimating remaining shelf life (RSL) from images could provide affordable decision support for perishable produce, but evaluation protocols can substantially affect reported performance when repeated images are available from the same biological specimen. We use the Hass Avocado Ripening dataset, comprising 8,834 image-RSL pairs from 426 fruits across three storage regimes, to evaluate a frozen ImageNet-pretrained visual backbone with a lightweight regression head. Our contributions are threefold: we quantify the effect of observation-level versus specimen-disjoint evaluation, compare lightweight and heavier visual backbones under specimen-disjoint cross-validation, and examine their spatial attributions using Grad-CAM. Across ten observation-level random splits, the model achieves a mean RMSE of 2.37 days with a standard deviation of 0.03 days, whereas specimen-disjoint 5-fold cross-validation yields a mean RMSE of 3.12 days with a standard deviation of 0.11 days. The corresponding mean coefficient of determination is 0.553. A matched per-specimen comparison confirms higher error under specimen-disjoint evaluation, with a probability value below 0.001 across 426 specimens, showing that observation-level partitioning gives a substantially more optimistic estimate for this dataset and model configuration. Under specimen-disjoint evaluation, MobileNetV3-Small (0.93 million parameters) achieves accuracy comparable to ResNet-18 while providing substantially higher throughput, and Grad-CAM reveals differences in spatial attribution between the lightweight backbones. These results support specimen-disjoint evaluation and attribution analysis when assessing lightweight vision models for longitudinal shelf-life prediction.

85. 【2610.09723】MeshCarve: Artisan Mesh Generation with Flow Matching in Compact Latent Spaces

链接:https://arxiv.org/abs/2610.09723

作者:Xiyu Wang,Ruocheng Wu,Yufei Wang,Zhihao Li,Lanqing Guo,Bihan Wen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:makes inference slow, largely predict face, face tokens autoregressively, predict face tokens, Prior artisan mesh

备注: 18 pages, 6 figures, 9 tables

点击查看摘要

Abstract:Prior artisan mesh generation works largely predict face tokens autoregressively, which makes inference slow. Recent methods instead flow match continuous latents built by Variational AutoEncoders (VAEs), but reconstruction quality drops significantly when geometry and topology are jointly encoded, and further when the latent space is compressed. We present MeshCarve, a flow matching method that generates entirely in compact latent spaces, generating vertex positions and edge connections separately and sidestepping the difficulty of a joint compact latent. To shorten the token sequence, we propose a hierarchical sparse transformer backbone, instantiated as VertexVAE and EdgeVAE. Instead of encoding fields over the surface voxels, both VAEs anchor on discrete vertices in their latent spaces, which drastically reduces the token sequence length, and our spatial-aware compression shortens it further without costing reconstruction. VertexVAE directly encodes vertex occupancy. For connectivity, we propose vertex-link encoding, which turns arbitrary connectivity between vertices into fixed-length continuous per-vertex embeddings and recovers complex artistic topology faithfully. MeshCarve combines these VAEs with an anchor generator and flow matches on the shortened token sequences. It shows advantages over state-of-the-art autoregressive and flow matching methods on Objaverse and generalizes to Toys4K. To the best of our knowledge, it is among the first artisan mesh generation methods whose every generative stage runs in a spatially compressed latent, with a token sequence only a fraction of the most compressed previous autoregressive and flow matching works.

86. 【2610.09720】DynStream: Online Streaming 4D Gaussian Reconstruction of Dynamic Worlds from Unposed Video

链接:https://arxiv.org/abs/2610.09720

作者:Dingwei Xian,Xiaoyu Zhou,Yajiao Xiong,Yongtao Wang,Ming-Hsuan Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing methods struggle, achieve simultaneously, Existing feed-forward Gaussian, feed-forward Gaussian methods, streaming videos requires

备注:

点击查看摘要

Abstract:Online reconstruction of dynamic 4D scenes from long, unposed streaming videos requires both continuous processing and photorealistic rendering, which existing methods struggle to achieve simultaneously. Existing feed-forward Gaussian methods are restricted to offline processing, whereas online point-cloud approaches struggle to maintain dense geometry and high-fidelity rendering. We present DynStream, a framework for streaming 4D Gaussian reconstruction from long, unposed videos. Given a continuous video stream, DynStream reconstructs the scene within local temporal windows and incrementally aligns and fuses these local reconstructions into a globally consistent scene, enabling online 4D reconstruction without per-scene optimization. By jointly enforcing cross-window geometric consistency and modeling time-varying scene content, DynStream supports efficient reconstruction and photorealistic rendering over extended video streams. Experiments demonstrate that DynStream enables high-fidelity online dynamic reconstruction and rendering from long video streams, achieving state-of-the-art performance across diverse dynamic indoor and outdoor scenes.

87. 【2610.09718】YUBI-STAG: Contact and Semantic-Rich Alignment for VLAs via Automated Video-Language Grounding

链接:https://arxiv.org/abs/2610.09718

作者:Masatoshi Tateno,Takehiko Ohkawa,Yueh-Hua Wu,Hanlong Li,Tatsuya Matsushima,Yoichi Sato,Kei Ota

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:requires fine-grained alignment, acquire broad manipulation, broad manipulation capabilities, models acquire broad, language requires fine-grained

备注: Project page: [this https URL](https://yubi-stag.airoa.io/)

点击查看摘要

Abstract:Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions are executed, including which gripper acts, which object is contacted, and how it is grasped and moved. We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language. Combining contact-object segmentation with vision-language models, YUBI-STAG annotates object identities, attributes and states, per-gripper actions, bimanual coordination, and spatially grounded interactions. To address YUBI-STAG's reliance on localized sequences and multi-stage VLM inference, we distill it into YUBI-VLM. YUBI-VLM directly recovers action structure and annotations from raw, unsegmented video in few inference calls and operates from wrist views alone. We evaluate both frameworks on YUBI-STAG-Bench across temporal, semantic, and spatial grounding tasks. YUBI-VLM retains much of YUBI-STAG's annotation accuracy with fewer inference calls and shorter runtime while generalizing to unseen manipulations. Finally, post-training VLA policies on these annotations aligns them with fine-grained language and contact-aware structure. Bimanual experiments demonstrate improved performance and instruction following, including control over object identity, acting gripper, target location, and spatial relations absent from original labels.

88. 【2610.09706】Enhancing Multi-Region Stylization with Interior-Guided Boundary Repair

链接:https://arxiv.org/abs/2610.09706

作者:Hong-Son Nguyen,Thi-Ngoc-Hanh Le

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Region-based neural style, transfer enables fine-grained, enables fine-grained artistic, fine-grained artistic control, Region-based neural

备注: 10 pages, 7 figures

点击查看摘要

Abstract:Region-based neural style transfer enables fine-grained artistic control by allowing independent stylization of semantic image regions. However, compositing these regions often leads to boundary artifacts, degrading visual quality. We propose Interior-Guided Boundary Repair (IGBR), a lightweight and model-agnostic method that improves boundary handling in multi-region stylization. IGBR repairs boundary pixels using interior-guided propagation and applies inward, distance-based blending restricted to object-background boundaries, preventing inter-object style leakage. The method is derived from a region-wise constrained formulation with a closed-form solution and can be seamlessly integrated into existing stylization pipelines without retraining. To evaluate efficiency of our IGBR, we introduce quantitative metrics that measure boundary consistency, gradient artifacts, inter-object leakage, and interior preservation without requiring annotated stylized images. Our experiments and evaluations demonstrate that the proposed IGBR consistently produces plausible boundaries, outperforming prior blending techniques in boundary consistency, gradient stability, and interior preservation. The code is available at this https URL.

89. 【2610.09702】Latent Watermarks under Generative Editing: A Benchmark and Analysis of Detection Survival

链接:https://arxiv.org/abs/2610.09702

作者:Sung Ju Lee,Nam Ik Cho

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Ordinary prompt-based editing, Ordinary prompt-based, fail without explicitly, explicitly targeting, Ordinary

备注:

点击查看摘要

Abstract:Ordinary prompt-based editing can cause latent watermark detection to fail without explicitly targeting the watermark. We benchmark eight watermark methods against five editors across four generative backbones, four editing strengths, and five semantic categories, with edit-validity and threshold checks. Separating editing from seven subsequent distortions reveals that editing alone primarily distinguishes Tree-Ring, while added distortions expose a broader spectrum of detection survival. Sequential edits reveal a second hidden difference: score separation can decline while detection rates remain near their ceiling. Across methods, standardized clean score separation ($d'$) organizes composite-survival tiers, whereas spatial overlap adds little to predicting edit-only survival beyond clean detectability. Embedding-strength interventions in two methods link higher clean separation to higher post-edit separation. In HSTR, the margin contrast is positive, while the angular layout contrast at matched clean separation remains unresolved. Together, outcome decomposition and continuous separation expose differences hidden by aggregate TPR. Method tiers are stable under threshold recalibration at the main operating points and alternative composite weights. Clean $d'$ is thus a useful empirical diagnostic within this benchmark, with mixed transfer to unseen methods. Code and supporting artifacts are planned for a separate release.

90. 【2610.09700】What Makes Synthetic Hard Negatives Work in Vision-Language Pretraining?

链接:https://arxiv.org/abs/2610.09700

作者:Nikos Giakoumoglou,Paschalis Giakoumoglou,Andreas Floros,Kleanthis Marios Papadopoulos,Tania Stathaki

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:unimodal self-supervised learning, Synthetic hard negatives, hard negatives generated, vision-language pretraining, self-supervised learning

备注: ACCV 2026

点击查看摘要

Abstract:Synthetic hard negatives generated in the representation space have proven effective for unimodal self-supervised learning, but transferring this idea to vision-language pretraining is not straightforward. We analyze six representation-space synthesis strategies and identify two failure modes in their transfer to vision-language pretraining: cross-modal constructions that produce overly easy negatives or pull them toward the query, and intra-modal constructions that incorporate the matched positive. We also observe logit-scale saturation when training with synthetic hard negatives and a learnable temperature, and find that fixing the temperature improves downstream performance. Using this geometric analysis we propose SNAP, which generates intra-modal hard negatives that never involve the positive from either modality, avoiding both failure modes entirely. SNAP is model-agnostic, requires no external generative models, and adds less than 10% training time overhead. Evaluated on top of CLIP and FLIP across multiple architectures and datasets, SNAP delivers consistent improvements on zero-shot retrieval, zero-shot classification, and linear probe evaluation.

91. 【2610.09695】Do Better Visual Representations Always Lead to Better End-to-End Autonomous Driving?

链接:https://arxiv.org/abs/2610.09695

作者:Zihao Zhang,Haochen Tian,Tianyu Li,Changhui Jing,Jingliang He,Naisheng Ye,Ziyuan Pu,Zhenjie Yang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Visual foundation models, foundation models, increasingly integrated, remains unclear, improve driving performance

备注:

点击查看摘要

Abstract:Visual foundation models (VFMs) are increasingly integrated into end-to-end autonomous driving for their powerful representations, yet it remains unclear when these representations improve driving performance. To investigate this question, we introduce ViRA, a planner-agnostic visual representation alignment framework that keeps the planner architecture and inference cost unchanged. Our study reveals three findings: (1) VFM-guided visual representations consistently improve driving performance across diverse end-to-end planners, with gains extending to zero-shot closed-loop evaluation. (2) The choice of VFM target matters for planning performance, and alignment to a different VFM can further benefit planners with pre-trained VFM encoders. (3) Auxiliary perception supervision reduces sensitivity to VFM target selection, narrowing the EPDMS spread across five targets from 2.7 to 0.5 points and potentially compensating for less effective VFM targets. Guided by these findings, we develop ViRA-Diffusion, a diffusion-based planner trained without auxiliary perception supervision, which achieves 92.3 EPDMS on NAVSIM v2 navtest, outperforming recent methods in our comparison by at least 1.9 points. The results motivate jointly considering target selection and planner supervision when integrating VFMs into end-to-end autonomous driving. The results and demo are available at this https URL.

92. 【2610.09677】A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods

链接:https://arxiv.org/abs/2610.09677

作者:Marco Riedenauer,Daniel Kienzle,Pratik Mayekar,Rainer Lienhart

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Self-supervised anomaly detection, anomaly detection, promising paradigm, paradigm for medical, easier to obtain

备注: 7 pages, 3 figures

点击查看摘要

Abstract:Self-supervised anomaly detection is a promising paradigm for medical ultrasound, as normal images are often easier to obtain than exhaustive annotations of all possible pathologies. However, most existing evaluations are limited to a single anatomy or task, making it unclear whether models learn a robust notion of normal ultrasound appearance or only a source-specific representation. We introduce the SADUSI benchmark, a multi-source ultrasound dataset designed to train and evaluate anomaly detection methods across a broad range of anatomical regions, views, and acquisition protocols. The goal of SADUSI is to provide a diverse normal ultrasound distribution and a benchmark for visible structural anomalies that can be assessed from single images. We evaluate representative self-supervised anomaly detection methods and find that current approaches struggle in this setting. In particular, reconstruction-based diffusion methods such as AnoDDPM and DeCo-Diff achieve pixel-level AUROC values of 0.56-0.72 and maximum F1 scores of 0.10-0.26, indicating limited separation of pathology from normal image regions. Feature-based PatchCore variants perform better, reaching pixel-level AUROC values of 0.76-0.83, but remain limited with maximum F1 scores of 0.14-0.40. These findings suggest that broad multi-source ultrasound anomaly detection remains an open challenge and that SADUSI can serve as a resource for developing methods that generalize beyond anatomy-specific settings.

93. 【2610.09670】Quasi-Binarized Autoencoders: An Architecture-Independent Information Bottleneck for Medical Image Anomaly Detection

链接:https://arxiv.org/abs/2610.09670

作者:Shouhei Hanaoka,Takahiro Nakao,Atsushi Takamatsu,Takeharu Yoshikawa,Osamu Abe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Unsupervised anomaly detection, medical image analysis, Unsupervised anomaly, anomaly detection, open problem

备注:

点击查看摘要

Abstract:Unsupervised anomaly detection, which learns only from normal images, is a central task in medical image analysis and remains an open problem. Reconstruction-based methods pass an image through an encoder-decoder network trained on normal data and detect anomalies from the residual between the image and its reconstruction. This works only if the information passed from the encoder to the decoder is limited; otherwise the network learns an identity mapping and reconstructs anomalies too. This limit is usually imposed through architectural choices, tuned per dataset, that cannot be stated in bits. We introduce the quasi-binarizing (QB) layer, which squashes each latent element into [0, 1] and adds Laplace noise of scale 1/epsilon. Each element is then epsilon-locally differentially private, and the mutual information between an image and its reconstruction is bounded by a quantity that depends only on epsilon and the number of QB elements, whatever the encoder and decoder. Placing a QB layer on every encoder-decoder path, including all skip connections, we build QBAE, a seven-level attention U-Net with 32,768 QB elements. On the seven datasets of the MedIAnomaly benchmark, QBAE with one architecture and one configuration reaches a mean image-level AUROC of 0.828, the highest among methods that do not adapt to each dataset, and the best reported results on BraTS2021 (AUROC 0.911, pixel-level AP 0.838). The noise is kept at test time, so that every reconstruction satisfies the bound. Without input corruption, the bottleneck alone prevents identity collapse (mean AUROC 0.805 vs. 0.590). Code is available at this https URL.

94. 【2610.09614】Identity-Duplication Auditing in National-Scale Neuroimaging Repositories

链接:https://arxiv.org/abs/2610.09614

作者:Jiheng Li,Michael E. Kim,Trent M. Schwartz,Yuhan Cui,Gaurav Rudravaram,Derek B. Archer,Timothy J. Hohman,Lori L. Beason-Held,Victoria L. Morgan,Dario J. Englot,Angela L. Jefferson, for theAlzheimer's Disease Neuroimaging Initiative, for theBIOCARD Study team, for theHealth,Aging Brain Study:Health Disparities(HABS-HD)Study Team,Lianrui Zuo,Guray Erus,Christos Davatzikos,Bennett A. Landman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:National-scale magnetic resonance, magnetic resonance imaging, National-scale magnetic, increasingly integrate data, repositories increasingly integrate

备注:

点击查看摘要

Abstract:National-scale magnetic resonance imaging (MRI) repositories increasingly integrate data from different studies and institutions. However, subject identifiers that are valid only within individual datasets are no longer guaranteed to remain globally unique after aggregation, making it possible for the same subject to be assigned multiple identifiers, which we define as identity duplication. Such duplication can create leakage between training and test data and inflate apparent performance in downstream biomedical studies. Existing methods do not provide an end-to-end, image-based workflow for auditing this problem at repository scale. In this work, we present HAPPEN, a human-in-the-loop pipeline for auditing identity duplication in T1-weighted brain MRI repositories. It combines SHA-256 fingerprinting for exact-duplicate detection with supervised contrastive retrieval of non-identical scans that may originate from the same person. Retrieved pairs are reviewed as candidates in a locally hosted interface rather than automatically classified as duplicates. We deployed the workflow in a 95,129-scan aggregated repository and assessed end-to-end recovery using 54 genetic-reference pairs. Transferability was assessed by locally deploying the same workflow on 22,386 scans at an independent institution without model retraining or image transfer. Deployment in the study repository identified 1,316 exact-duplicate scan groups and 1,275 reviewer-supported near-duplicate subject groups. Of these groups, 56% and 82%, respectively, crossed dataset boundaries. All 54 genetic-reference pairs were recovered. The external team independently completed the full workflow using a locally selected operating threshold and review standard.

95. 【2610.09598】STORK: Spatio-Temporal Observation of uterine contRactions via neural networKs

链接:https://arxiv.org/abs/2610.09598

作者:Melissa Schween,Tristan Gottwald,Jordina Aviles Verdera,Lisa Story,Mary Rutherford,Jana Hutter

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:typically identified manually, Contractile Activity Detection, Uterine Contractile Activity, manually and discarded, limiting insights

备注: Accepted at the PIPPI Workshop at MICCAI 2026 and will appear in the workshop proceedings (Springer)

点击查看摘要

Abstract:Uterine contractions in fetal MRI are typically identified manually and discarded, limiting insights into contraction dynamics. We formalize Uterine Contractile Activity Detection (UCAD) as a weakly-supervised learning problem and introduce STORK, a multi-instance learning model trained on dynamic MRI series using only coarse, series-level labels. STORK factorizes 3D spatio-temporal convolutions into parallel branches across temporal hyperplanes to capture coherent tissue motion without the cost of full 4D convolutions. Per-frame embeddings, combining intensity and Demons-estimated displacement fields, are aggregated by a linear mean-pooling head. This ensures that frame-level contraction scores can be recovered post-hoc without frame-level training supervision. Evaluated on around 700 multi-vendor dynamic fetal MRI series, STORK achieves a series-level AUROC of 95.0% and AUPRC of 94.6%, substantially outperforming 3D ResNet and ConvNeXt baselines. Grad-CAM analysis suggests that the model draws on predictive features extending beyond the placenta into the uterine tissue, offering an automated tool for richer phenotyping of uterine behavior.

96. 【2610.09579】Gradient-Based Trajectory Optimisation over Continuous Poses for Sparse-View Cone-Beam CT

链接:https://arxiv.org/abs/2610.09579

作者:Linda-Sophie Schneider,Simon Wittl,Gabriel Herl,Andreas Maier

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cone-beam computed tomography, sparse-view scans acquire, information sparse-view scans, computed tomography, determines which information

备注: Submitted to TPAMI

点击查看摘要

Abstract:Trajectory optimisation for cone-beam computed tomography (CT) determines which information sparse-view scans acquire. Fixed candidate pools prevent off-grid refinement and require new object-specific precomputation for each acquisition manifold. We make every source pose an individual continuous variable and move all poses jointly by gradient ascent on the scanner's kinematic manifold. The objective combines soft-Tuy plane coverage, continuous View Covariance Loss, and an analytic attenuation-aware ray-bundle penalty. The same optimiser handles circular, limited C-arm, two-axis, and freesphere parametrisations. On a Defrise flange, continuous selection recovers laminar defects invisible to a circular orbit, matches discrete swap search on the free sphere at the sparser budget, and leads at the denser one, with the same objective evaluated in every arm. A moderate elevation band already recovers most of the free-sphere gain at the defects, so the same optimiser transfers to bounded scanner envelopes. Photon noise preserves the ordering on the flange and compresses it on a dense fuel nozzle. Sparseprescan planning benefits from matching prescan and planned acquisition manifolds. Selection takes seconds rather than minutes without an object-specific reconstruction basis. Prescan-planned poses were executed on a robot CT bench and reconstructed in a common frame, demonstrating feasibility but no consistent metric gain over uniform band sampling. Continuous pose optimisation incorporates attenuation and scanner constraints directly into sparse-view acquisition design.

97. 【2610.09550】Visual Evidence Under Cross-Examination: Evaluating and Controlling Decision-Level Evidence Use in Vision-Language Models

链接:https://arxiv.org/abs/2610.09550

作者:Huiyao Zhang,Jin Bai,Zilong Su,Rui Guo,Chaofan Qin,Jinze Lv,Wenhui Yu,Hongfei Wang,Ye Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision-language models increasingly, models increasingly reason, reason through crops, Vision-language models, increasingly reason

备注:

点击查看摘要

Abstract:Vision-language models increasingly reason through crops, regions, and tool-produced observations. Yet an observation can influence the answer without benefiting the candidate it supports. We study candidate-bound visual contribution: valid evidence should help, invalidating its supporting relation should remove its additional effect, and valid rebinding should redirect that effect to the newly supported candidate. We introduce CROSS-Bench, a benchmark of 28,000 decision problems, with matched invalidation and rebinding tests on a dedicated evaluation subset. Our RIVET interface preserves evidence identity and uncertainty, composes a candidate-conditioned response, and separately controls its strength. Shared-evidence experiments show that task accuracy and evidence ownership can diverge. Under matched capacity and training, RIVET increases normalized effect transfer from 0.512 to 0.651 where clean evidence has a positive effect. The advantage persists on common evaluation examples and across repeated decision-layer fits. With evidence predicted from raw inputs, RIVET improves CROSS-Bench accuracy by an average of 5.70 pp across four frozen backbones, relative to the same models without auxiliary evidence. These results separate the utility of visual evidence from the candidate-specific destination of its effect.

98. 【2610.09535】WAPR: A Foundation Model for Wide-Angle Refinement in Unseen Object Pose Estimation

链接:https://arxiv.org/abs/2610.09535

作者:Yulin Wang,Mengting Hu,Hongli Li,Jianghao Zhou,Chen Luo

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:Real-world applications require, Real-world applications, applications require, scalable to unseen, Real-world

备注: Accepted to ECCV 2026. 19 pages, 4 figures. Yulin Wang and Mengting Hu contributed equally. Corresponding author: Chen Luo

点击查看摘要

Abstract:Real-world applications require 6D pose estimation to be accurate, fast, and scalable to unseen objects. This paper introduces WAPR, a zero-shot wide-angle pose refinement model that refines candidate poses with rotational deviations up to 90 degrees. With as few as 12 candidate poses per detected object instance, WAPR supports fast inference within 1 s per frame and reaches a pose-estimation throughput of up to 25 detected object instances per second. To support wide-angle training for rotationally symmetric objects, WAPR uses rotational symmetry priors to canonicalize symmetry-equivalent pose targets before loss computation. We further construct SA6D, a large-scale 6D training dataset with such priors. SA6D obtains KASAL-assisted rotational symmetry priors for 944 GSO scans and expands them through geometry and texture augmentation into about 50K augmented object instances and about 2M rendered RGB-D images. In addition, an angle-balanced loss stabilizes learning across different angular ranges by reducing the influence of uninformative large-error cases. Experiments on seven BOP core datasets show that WAPR achieves state-of-the-art performance in unseen-object 6D pose localization and detection under both fast and unconstrained inference settings. Project page: this https URL.

99. 【2610.09534】KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization

链接:https://arxiv.org/abs/2610.09534

作者:Mengxin Zhang,Yulin Wang,Chen Luo,Yongzhe Li,Yijun Zhou

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:supporting symmetry-aware evaluation, symmetry-aware evaluation, supporting symmetry-aware, Rotational, improving pose

备注: CVPR 2026. 10 pages, 4 figures. Mengxin Zhang and Yulin Wang contributed equally. Corresponding authors: Chen Luo and Yijun Zhou

点击查看摘要

Abstract:Rotational symmetry is an important prior in 6D pose estimation, improving pose accuracy and supporting symmetry-aware evaluation. However, current symmetry annotations for 3D objects remain largely manual or semi-automatic, often requiring predefined types or orders, which limits scalability. This work introduces a fully automatic, reference-free framework for symmetry-type classification, rotational-order identification, and full-axis localization across all eight canonical 3D rotational symmetry types. The method localizes a dominant high-order axis, infers its rotational order through self-consistency analysis, and reconstructs the complete symmetry structure under a hierarchy-guided formulation. A texture-aware extension further models appearance-induced reductions in rotational order while preserving axis orientations. Experiments on idealized and real-world datasets demonstrate strong accuracy and generalization, achieving 94.75% accuracy on 438 symmetric objects in GSO. Training FoundationPose with these priors improves accuracy by up to 0.9% across five BOP datasets, showing that automatically estimated rotational priors improve downstream 6D pose estimation. Code is available at this https URL.

100. 【2610.09518】ActiveLang: Active Open-Vocabulary 3D Mapping with Semantic-Uncertainty-Guided Exploration

链接:https://arxiv.org/abs/2610.09518

作者:Liyan Chen,Hairong Yin,Huangying Zhan,Yi Xu,Raymond A. Yeh,Philippos Mordohai

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:increasingly assist humans, robots increasingly assist, increasingly assist, assist humans, humans with diverse

备注:

点击查看摘要

Abstract:As robots increasingly assist humans with diverse tasks, they need both geometric and semantic understanding of their surroundings. Moreover, robots often operate in unfamiliar environments and take on new tasks without knowing the relevant concepts ahead of time. This motivates language-annotated 3D maps that support open-vocabulary scene understanding and human-robot interaction. We introduce ActiveLang, an autonomous system for active open-vocabulary 3D mapping with semantic-uncertainty-guided exploration. ActiveLang performs online language-feature adaptation on a compact dual-Gaussian representation to jointly reconstruct scene geometry, appearance, and open-vocabulary semantics with modest memory overhead. Its planner efficiently selects informative viewpoints, enabling effective mapping with fewer observations and lower computational cost. Experiments on Replica and ScanNet++ demonstrate substantial improvements in 2D and 3D open-vocabulary segmentation over both online and offline baselines, highlighting that actively exploring scenes builds language-annotated 3D maps more efficiently.

101. 【2610.09517】It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

链接:https://arxiv.org/abs/2610.09517

作者:Samuel Tetteh,Cody Fleming

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:vision-language models increasingly, models increasingly provide, Frozen vision-language models, increasingly provide safety, reinforcement learning

备注:

点击查看摘要

Abstract:Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.

102. 【2610.09514】STRIKE: Learning Visual State Transitions for Physical World Modeling

链接:https://arxiv.org/abs/2610.09514

作者:Wenbin Teng,Tianshuo Xu,Depu Meng,Yuelei Li,Quentin Herau,Yihan Hu,Yajie Zhao,Wei Zhan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generating coherent motion, modeling requires predicting, coherent motion, requires predicting, predicting how interactions

备注:

点击查看摘要

Abstract:Physical world modeling requires predicting how interactions change a scene, not merely generating coherent motion. We propose STRIKE, a framework that separates visual state transition learning from dense video generation. We construct event-aligned supervision by extracting observed states from training videos and pairing them with transition descriptions and temporal offsets. An image-based transition model learns to predict the next scene configuration from the current image, a local transition specification, and elapsed time. At inference, a pretrained vision-language planner predicts time transition specifications, and recursive application of the learned transition model produces a sequence of future visual states. A separately trained dynamic model then generates the complete rollout conditioned on these states and their temporal locations. Experiments on Physics-IQ Verified, PhyGenBench, Pisa-Experiments, and RoboTwin2.0 show improvements of STRIKE over the corresponding video-backbone baselines in benchmark measures of physical consistency and manipulation-video fidelity. These results support learned visual state transitions as an effective intermediate representation for physical world modeling.

103. 【2610.09513】OmniCam: Omni-Camera Trajectory Generation via Geometry-Grounded Pose Token Learning

链接:https://arxiv.org/abs/2610.09513

作者:Zhenyang Liu,Chenjie Cao,Yisu Zhang,Xuhui Zuo,Xiangyang Xue,Yanwei Fu,Tengfei Wang,Chunchao Guo

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:trajectories control viewpoint, Camera trajectories control, control viewpoint, scene reconstruction, trajectories control

备注:

点击查看摘要

Abstract:Camera trajectories control viewpoint changes in video generation, scene reconstruction, and robotic perception. Generating them from language requires both scene geometry and target-aware framing. We introduce OmniCam, an autoregressive model that generates camera pose sequences from a single panorama and textual trajectory descriptions. Its geometry-grounded pose token learning combines three components: a panoramic point-cloud encoder for omnidirectional geometric context; hybrid absolute-rotation and relative-translation tokenization with temporally consistent quaternion signs; and separate geometric and semantic conditioning streams with an explicit 3D target anchor. We also construct OmniCaT, containing 267,700 trajectories across four camera behaviors. On the reported OmniCaT evaluation, OmniCam reduces trajectory errors by 28--47% and collision rate by 65.8% relative to GenDoP retrained on OmniCaT. Against the best baseline for each metric, the ATE and collision reductions are 43.0% and 62.3%, respectively. Component ablations support the use of geometric and target-aware conditioning, while downstream experiments examine camera-controlled video generation and robotic active perception.

104. 【2610.09511】LiG-DETR: Local-in-Global Reassembly in Latent Space for Aerial Object Detection

链接:https://arxiv.org/abs/2610.09511

作者:Yupeng Zhang,Fangzhuo Gao,Juntao Cheng,Ziyi Zhao,Liang Wan,Ruize Han

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:faces substantial scale, object detection faces, density variations, scale and density, Aerial object detection

备注:

点击查看摘要

Abstract:Aerial object detection faces substantial scale and density variations. Small objects are easily degraded by downsampling and feature compression, while medium and large objects require sufficient global context. Existing methods mainly follow two paradigms: image slicing provides clearer local evidence but relies on independent crop-level prediction and post-processing, whereas feature- and query-level optimization preserves unified inference but operates on already compressed full-image representations, limiting recovery of fine-grained information. This raises a key question: can aerial detection directly acquire high-fidelity local evidence before feature degradation and integrate it into a unified end-to-end framework? To this end, we propose LiG-DETR, an Efficient Global-Local Reassembly framework that reformulates image slicing as high-fidelity local feature acquisition. A shared encoder extracts global and locally magnified features, which are projected into the detector feature space. The projected local features are reassembled according to their original spatial locations to form a globally aligned local feature level, and a single DETR decoder jointly decodes global and local features. To reduce redundant computation, Context-Preserved Selective Reassembly focuses high-resolution encoding on informative regions while preserving a dense feature layout, and Density-Aware Adaptive Query Allocation adapts the decoder query budget using encoder proposal scores. Experiments show substantial gains on small and medium objects while retaining strong large-object performance, with favorable accuracy--efficiency trade-offs and improved cross-domain generalization. The code will be released.

105. 【2610.09509】ERRA: Learning Transportable Latent Actions through Temporal Effect Representation and Relational Alignment

链接:https://arxiv.org/abs/2610.09509

作者:Tianxingjian Ding,Mubarak Shah,Yu Tian

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)

关键词:action-like codes inferred, supervise robot policies, actions supervise robot, action-like codes, codes inferred

备注: Preprint. Code and project page coming soon

点击查看摘要

Abstract:Latent actions supervise robot policies with action-like codes inferred from visual transitions, and their usefulness hinges on two questions: what a code keeps from a transition, and whether it still means the same thing when reused in a different initial state. The first is a tension in time: an endpoint difference discards how motion unfolds, while the full sequence admits nuisance variation. The second is left open by reconstruction, which only ever observes a latent together with the state it came from. We argue that both questions can be answered in the same place. TERRA (Temporal Effect Representation and Relational Alignment) describes a transition by a compact temporal effect, its net feature change together with a low-order within-window dynamics component, and learns a continuous latent from this effect. The same effect space then serves as the reference for reuse: Effect-Anchored Transport (EAT) decodes a latent in other initial states and anchors the resulting effect to the one observed at its source, so that the latent is shaped by what it does across contexts rather than only by the transition it came from. With frozen linear readers, TERRA predicts actions more accurately than UniVLA and a LAPA-style baseline, degrades more slowly under visual distractors, and keeps transported transitions faithful to the donor action as the recipient context moves farther away; a same-budget control shows that these gains come largely from EAT. At matched pretraining scale, the complete system reaches 93.4% average success on LIBERO, compared with 91.8% for UniVLA.

106. 【2610.09502】RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection

链接:https://arxiv.org/abs/2610.09502

作者:Yupeng Zhang,Ziyi Zhao,Juntao Cheng,Sheng Wang,Ningnan Guo,Ruize Han,Liang Wan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recognizes categories unseen, Open-vocabulary detection, achieving strong generalization, efficiency remains challenging, recognizes categories

备注:

点击查看摘要

Abstract:Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.

107. 【2610.09498】SpatialUQ: Post-Hoc Uncertainty Quantification from Spatial Consistency in Black-Box Vision Models

链接:https://arxiv.org/abs/2610.09498

作者:Md Kawsher Mahbub,Milon Biswas,Mirza Niaz Morshed,Wei Yu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:frozen black boxes, Clinical vision models, Clinical vision, access to internals, inference time

备注:

点击查看摘要

Abstract:Clinical vision models are often deployed as frozen black boxes with no access to internals, retraining, or ground truth at inference time. We introduce \textbf{SpatialUQ}, a post-hoc uncertainty method using only output probabilities. It measures the Jensen-Shannon divergence between the global prediction and the mean of five fixed spatial crops in six deterministic forward passes. The premise is simple, trustworthy predictions are spatially consistent. On NIH ChestX-ray14 (DenseNet-121, $N{=}25{,}596$), our Multicrop Uncertainty Score (MUS) reaches $0.784$ failure-detection AUC versus $0.664$ for MC-Dropout ($p{}10^{-6}$) at one-fifth the compute, with native calibration ($\text{SCE}{=}0.049$ vs.\ $0.127$ for $\ell_1$), the best-calibrated among methods above 0.78 AUC. A supervised fusion of MUS with entropy, confidence, and $\ell_1$ reaches $0.832$, outperforming a five-member ensemble ($0.813$). MUS scales with model quality, reaching $0.899$ with BiomedCLIP ($\rho = 0.846$), while this relationship remains meaningful in-distribution ($\rho = 0.523$) but breaks down under severe distribution shift (VinBigData, $\rho = 0.027$). MUS is well-suited to diffuse findings but is less dependable for small focal lesions such as nodules. Code and experimental materials are publicly available at this https URL.

108. 【2610.09492】Gaussian Material Fields for Volumetric Multi-Energy CT Decomposition

链接:https://arxiv.org/abs/2610.09492

作者:Jian Lin,Jiancheng Fang,Hongming Shan,Shaoyu Wang,Yang Chen,Qiegen Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:computed tomography requires, organizes multiple three-dimensional, common spatial domain, multi-energy computed tomography, Gaussian material fields

备注: 13 pages, 12 figures

点击查看摘要

Abstract:Volumetric material decomposition in multi-energy computed tomography requires a representation that organizes multiple three-dimensional material fields in a common spatial domain while retaining differences in composition and local structure. We observe that spatial primitives can be shared across materials without tying their coefficients, but their local capacity must respond to material-specific reconstruction needs. We introduce Gaussian material fields, which represent multiple material distributions with shared anisotropic 3D Gaussian primitives and independent nonnegative material coefficients. The shared geometry defines a continuous spatial basis, while the coefficients determine each primitive's contribution to the individual material fields. To reconstruct this representation from multi-energy projections, a differentiable spectral forward model combines Gaussian material path integrals with a calibrated basis matrix, enabling joint optimization of spatial geometry and material composition. Material-aware adaptive density control retains material-specific refinement evidence before aggregation and adjusts local representation capacity to accommodate both spatially extensive components and sparse details. Experiments use synthesized multi-energy projections generated from pseudo-reference material maps constructed by conventional methods from publicly available CT data. Across 15 cases, our approach improves average PSNR by 4.03 dB and SSIM by 4.96% over the strongest baseline, while reducing NRMSE by 33.45%. Material-wise comparisons and component ablations support improved recovery of localized structures, while runtime and memory measurements show favorable computational scaling. These results establish Gaussian material fields as an explicit, adaptive representation for volumetric multi-material reconstruction.

109. 【2610.09478】InstanceBench: Diagnosing Referential Reasoning and Target Identity in Referring Expression Segmentation

链接:https://arxiv.org/abs/2610.09478

作者:Yuchen Li,Shaoyang Zhou,Yiran Wang,Ruiyi Deng,Haoyu Wang,Ziru Wei,Zhen Zhao,Luping Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Referring Expression Segmentation, links natural-language descriptions, Referring Expression, Expression Segmentation, pixel-level object masks

备注:

点击查看摘要

Abstract:Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insight into instance-level referential reasoning: it does not systematically distinguish referential logics, test target preservation across valid grounding paths, or separate target-selection from mask-generation errors. We introduce InstanceBench, an instance-centered diagnostic benchmark comprising 6,194 images, 9,264 target instances, and 25,077 human-verified expressions. Each target-centric expression set (TCES) fixes the image and target mask while pairing a minimal expression with a same-target variant that uses another valid cue or grounding path. A compact referential-logic taxonomy spans direct target evidence, same-class selection, relational and compositional grounding, and exclusion, while logic-critical construction suppresses simpler shortcuts. Identity-aware metrics measure target retention and set-level success while separating selection from mask-generation errors. Across 22 native-mask RES checkpoints from 18 model families, the strongest checkpoint reaches 67.1% mIoU but only 59.6% All@0.7. Controlled interventions confirm language sensitivity, while failure decomposition identifies target selection rather than mask decoding as the main bottleneck. On a controlled training subset, matched supervision improves identity-aware performance, showing that the diagnosed capability responds to targeted supervision. Collectively, InstanceBench supports a measure-diagnose-improve cycle: measuring target consistency across grounding paths, localizing failure sources, and evaluating targeted interventions.

110. 【2610.09459】An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Fragment Adjacency Prediction

链接:https://arxiv.org/abs/2610.09459

作者:Guillaume Brouillette,Alain Goupil,Pierre-Olivier Parisé,Fadel Touré(Université du Québec à Trois-Rivières, Trois-Rivières, Canada)

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:addresses the prediction, network scores pairs, fragments based, rotation-equivariant Siamese convolutional, Siamese convolutional neural

备注: 14 pages, 4 figures, 6 tables

点击查看摘要

Abstract:This paper addresses the prediction of adjacency between pairs of 2D fragments based on their contours. We improved the two-stage architecture proposed in Beaulac's thesis, in which a rotation-equivariant Siamese convolutional neural network scores pairs of local image windows along the two contours of two fragments. The scores are gathered in an adjacency matrix in which a ResNet detects the partial anti-diagonal band that reveals the adjacency of two fragments. In the current work, we keep the pipeline and replace the local score by a comparison of tangent-angle profiles of contour windows, making it, by construction, invariant to fragment rotation and agnostic to the selected contour-starting point. These adaptations may be either a training-free likelihood ratio or a small one-dimensional convolutional model trained on corresponding points. We also replaced the final classifier by a band U-Net that segments the band and classifies the pair, so that the shared arc is obtained along with the decision. In the synthetic data set of the original thesis, the tangent descriptor performs as well as or better than the image-window approach in all tested configurations. The proposed pipeline reaches an accuracy of 98%, vs 93% to 95% for the original approach once its evaluation is corrected. We tested our pipeline, with models trained only on synthetic data, on the PairingNet benchmark, and obtained an AUC of 0.93. Furthermore, under the PairingNet pair-searching protocol conditions, our learned descriptor obtains a Recall@10 of 0.82 on the real set against 0.56 from the best model of the original paper.

111. 【2610.09455】RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

链接:https://arxiv.org/abs/2610.09455

作者:Seungjun Moon,Subin Jeon,Sangwoo Kim,Hanbyul Joo,Jinwoo Shin

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:robot policy training, approaches that leverage, increasingly prevalent, leverage human video, Recently

备注: 34 pages, 15 figures

点击查看摘要

Abstract:Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at this https URL.

112. 【2610.09450】Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning

链接:https://arxiv.org/abs/2610.09450

作者:Hanqiu Li Cai(SperidLabs),Chema Garabito(SperidLabs)

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:fine-grained detail matters, lossy VAE, diffusion models avoid, detail matters, avoid the lossy

备注: 19 pages, 13 figures, 6 tables. Project lead: Hanqiu Li Cai. Code and models: [this https URL](https://github.com/speridlabs)

点击查看摘要

Abstract:Pixel-space diffusion models avoid the lossy VAE of latent models, which suggests an advantage on downstream tasks where fine-grained detail matters. We test this claim along both routes to a pixel-space backbone. We pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, from scratch through a $256\to512\to1024$ curriculum, after first ablating the prediction target and representation alignment at $256^2$ to decide what to scale. We also convert a pretrained latent model, FLUX.2 Klein base 4B, to pixel space. We fine-tune both families for monocular depth estimation and for image restoration/super-resolution. We find no significant improvement from using a pixel-space generative prior. Fine-tuned for depth with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein and the converted pixel FLUX.2 Klein falls behind it, and on $4\times$ DIV2K restoration neither pixel model beats a latent FLUX.2 Klein fine-tune, the converted one trailing it slightly. We document the recipes, the failure modes and the remaining confounds behind this negative result. Nevertheless, Iris-3B shows that pixel-space pretraining with the pixel-transformer (PiT) head of PixelDiT scales to 3B parameters and to text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at $1024^2$. We release its weights and training code in the hope that they help pave the way for further work on pixel-space generation.

113. 【2610.09449】From Global Alignment to Local Grounding: Zero-Shot Chinese Character Recognition with Radical Verification

链接:https://arxiv.org/abs/2610.09449

作者:Yu-Heng Shih,Bing-Chen Wu,Tsz-To Wong,Ting-En Yen,Hong-Han Shuai,Bin-Hua Hsieh,Chien-An Chen,Yi-Ren Yeh,Ching-Chun Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Chinese character recognition, Zero-shot Chinese character, compositional structure shared, Ideographic Description Sequence, Chinese character

备注:

点击查看摘要

Abstract:Zero-shot Chinese character recognition (ZS-CCR) aims to recognize characters whose categories are never observed during training, and typically relies on the compositional structure shared between seen and unseen characters. Recent CLIP-style methods represent this structure with the Ideographic Description Sequence (IDS) and align it with glyph images in a shared embedding space. However, they rely on a single global image--IDS similarity that discards the spatial layout of radicals and, being learned only implicitly from seen classes, generalizes poorly to unseen ones; moreover, global matching often retrieves the correct character within the top candidates yet fails to rank it first when characters differ only in subtle local radicals. To address these issues, we propose a global-to-local two-stage framework. In the first stage, STG-CLIP augments the IDS with explicit tree-position and radical-level geometric priors, yielding a spatial-aware prototype that provides a consistent spatial description across seen and unseen categories for high-recall global retrieval. In the second stage, the Radical Verification Module (RVM) uses the radical instances of each retrieved candidate as queries to verify whether the corresponding radicals can be matched to spatially compatible regions in the input glyph. A margin-based gating rule activates the RVM only when the leading global candidates receive similar similarity scores. Experiments on the ICDAR2013 benchmark demonstrate that our method achieves state-of-the-art performance under the character-level zero-shot setting, obtaining 83.06% top-1 accuracy with 2,755 seen classes. Ablation studies further show that the explicit geometric priors and radical-level verification provide complementary improvements.

114. 【2610.09444】LighTROcc: Lightweight 4D Occupancy Forecasting via Instance-Centric 3D Gaussians

链接:https://arxiv.org/abs/2610.09444

作者:Hwanhee Jung,SeungHyeon Kim,Inkyu Koo,Qixing Huang,Sang Ho Yoon,Sangpil Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:existing approaches rely, cost grows rapidly, autonomous driving, prediction horizon, essential for autonomous

备注:

点击查看摘要

Abstract:Forecasting future 3D occupancy from surround-view cameras is essential for autonomous driving, yet existing approaches rely on dense voxel or bird's-eye-view representations whose cost grows rapidly with spatial resolution and prediction horizon. Because these representations do not explicitly maintain object identities, they also struggle to preserve instance consistency over time. We present LighTROcc, a lightweight instance-centric framework that represents movable objects with a compact set of learned queries and predicts present and future occupancy in a single forward pass. LighTROcc localizes each query through attention-guided forward lifting, combining image-space cross-attention, query-specific depth, and camera geometry to estimate its 3D center. Each instance is modeled as a mixture of anisotropic 3D Gaussians and propagated across future steps using predicted displacements, producing continuous, temporally consistent occupancy forecasts. Experiments on nuScenes and supplemented nuScenes-Occupancy show that LighTROcc outperforms the evaluated dense and instance-wise baselines in instance-level forecasting accuracy while maintaining strong voxel-level occupancy quality. Across different model configurations, LighTROcc achieves a favorable balance between forecasting accuracy and computational efficiency, demonstrating the potential of compact instance-centric modeling for camera-based 4D occupancy forecasting.

115. 【2610.09441】IRA: Tumor Immune Representation Adaptation for Zero-Shot Cross-Cancer MSI and TMB Prediction

链接:https://arxiv.org/abs/2610.09441

作者:Dasari Naga Raju

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:clinically relevant biomarkers, tumor mutational burden, Microsatellite instability-high, distinct cancer types, high tumor mutational

备注:

点击查看摘要

Abstract:Microsatellite instability-high (MSI-H) and high tumor mutational burden (TMB-H) are clinically relevant biomarkers, yet their histopathological prediction remains challenging when models are transferred across morphologically distinct cancer types. Immune-associated spatial patterns can persist across cancers despite these morphological differences, but foundation-model-based predictors trained on a single cancer do not explicitly use this information, limiting cross-cancer generalization. To address this limitation, we propose TIRA (Tumor Immune Representation Adaptation), a target-free framework that refines frozen foundation-model representations using spatial immune topology, without requiring target-domain data during model development or test-time adaptation. TIRA uses a topology-supervised biology representation to condition tile-level attention while pooling only morphological features for joint MSI and TMB prediction. We train TIRA on TCGA-COAD+READ and evaluate it zero-shot on CPTAC-COAD, TCGA-STAD, TCGA-UCEC, and CPTAC-UCEC, covering cross-site, cross-cancer, and combined cross-cancer-site distribution shifts under UNI2, CONCH, and Virchow2. With UNI2, TIRA improved zero-shot AUROC on TCGA-STAD from 0.633 to 0.766 for MSI and from 0.651 to 0.772 for TMB. Source-derived spatial immune topology improved the cross-cancer robustness of frozen pathology foundation-model representations.

116. 【2610.09440】Mixture of Layers: Dynamic Layer Routing for Visual Reasoning

链接:https://arxiv.org/abs/2610.09440

作者:Jeonghwan Kim,Sofia Stoica,Jiwan Chung,Ansel Blume,Hyeonjeong Ha,Zhenhailong Wang,Xin Luna Dong,Heng Ji

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Pre-trained vision encoders, Large Language Models, Multimodal Large Language, Pre-trained vision, vision encoder layers

备注: NeurIPS 2026

点击查看摘要

Abstract:Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders' receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at this https URL.

117. 【2610.09439】InscriptionOCR: A Dataset and Method for Understanding Inscriptions

链接:https://arxiv.org/abs/2610.09439

作者:Jaidev Sanjay Khalane,Akbar Ali,V. N. Prabhakar,Shanmuganathan Raman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:fundamental problem, affects the reliable, interpretation of historical, historical documents, Emperor Ashoka

备注:

点击查看摘要

Abstract:Ancient script image restoration is a fundamental problem in computer vision, as it directly affects the reliable analysis and interpretation of historical documents and inscriptions. Ashokan Brahmi is an ancient script extensively used during the reign of Emperor Ashoka in the 3rd century BC, primarily for inscriptions in Prakrit. These inscriptions, including major and minor rock and pillar edicts, constitute a valuable yet largely unexplored source of data for computational analysis. The degraded nature of inscription imagery and the lack of standardized digital resources pose significant challenges for automated processing. We present an end-to-end AI-based framework for understanding ancient inscriptions that encompasses image enhancement, optical character recognition (OCR), transliteration, and neural machine translation (NMT). The proposed pipeline processes low-quality images captured directly from stone inscriptions, performs image restoration and Brahmi script character recognition, maps the recognized characters to the Roman script, and finally translates the resulting Prakrit text into English. We also introduce two new datasets: (i) InscriptionOCR Dataset: the largest publicly usable digital OCR dataset for Brahmi script to date, consisting of over 200,000 character images across about 600 classes, and (ii) a bilingual Prakrit-English parallel corpus comprising over 2,000 sentence pairs for NMT. We believe that the proposed framework and datasets will facilitate future research in ancient script analysis, low-resource OCR, and digital epigraphy.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.09439 [cs.CV]

(or
arXiv:2610.09439v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2610.09439

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
118. 【2610.09438】Controllable Crowd Generation through World-Model Planning

链接:https://arxiv.org/abs/2610.09438

作者:JunGyu Lee,Jisu Shin,Seunghyun Shin,Hae-Gon Jeon

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous driving, Crowd, robot navigation, plays a central, central role

备注: 28 pages, 6 figures. Project page: [this https URL](https://jungyu0413.github.io/Ctrl-CWM)

点击查看摘要

Abstract:Crowd simulation plays a central role in robot navigation, autonomous driving, and urban planning. For these applications, realistic simulation requires crowds to adapt their behavior to environmental changes and user objectives. However, existing methods that rely on predefined control settings have limited flexibility in accommodating new user-specified objectives. To address this limitation, we propose Ctrl-CWM, a multi-agent Controllable Crowd World Model that integrates crowd generation and run-time control. Our key idea is to adapt the world-model principle of planning using imagined futures to crowd simulation. To this end, Ctrl-CWM consists of an encoder that learns a representation of human motion dynamics, an actor that proposes pedestrian displacements, a critic that evaluates imagined crowd trajectories, and a planner that selects actions. We first learn human motion dynamics through trajectory prediction on real-world pedestrian videos and then freeze the encoder to preserve them. Using this representation, the actor generates imagined crowd trajectories through repeated state updates, and the planner combines the critic's scores with user costs to select actions. Repeated planning advances the simulated crowd, while additional user costs introduce new control objectives without retraining. We extensively evaluate crowd generation under varied agent arrival conditions and run-time control across avoidance and attraction scenarios. Ctrl-CWM outperforms the state-of-the-art method on most crowd realism and collision metrics, and adapts crowd behaviors to user-specified objectives introduced during simulation. The project page is available at this https URL

119. 【2610.09430】SkillCycle: Co-Evolving Agent Policies and Skill Banks

链接:https://arxiv.org/abs/2610.09430

作者:Ling Li,Qiuyu Shen,Zheng Jiang,Qinwei Ma,Yuxuan Liu,Zhidong Deng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:newly encountered decisions, Internalizing external skills, Internalizing external, encountered decisions, insufficient for newly

备注:

点击查看摘要

Abstract:Internalizing external skills changes a language agent's capabilities and, with them, the value of its remaining guidance: rules can become redundant, misleading, or insufficient for newly encountered decisions. This creates a coupled problem of learning from skills and adapting the skills that supervise further learning. We introduce SkillCycle, a framework for co-evolving agent policies and skill banks through a feedback loop between skill internalization and rule revision. Our central contribution is to give distillation feedback a second role: token-level contextual differences help locate rules for inspection, while interaction outcomes guide edits to their content and applicability. SkillCycle alternates between two phases: policy learning with a fixed skill bank and router, and rule revision with a frozen policy. Candidate edits undergo rule-level and whole-bank environment comparisons before they guide the next learning cycle. On WebShop, SkillCycle with a 3B model achieves a success rate of 74.74% and a score of 88.37 without inference-time skill inputs, representing relative improvements of 0.73% and 3.96% over the state-of-the-art (SOTA) model, respectively. In Cycle 3 ablations on ALFWorld and WebShop, SkillCycle's no-skill success rates improve by 10.18% and 18.11% relative to a static skill bank, and by 2.41% and 2.50% relative to a single bank update, respectively. These results show that continually revising skill guidance as the agent's capabilities change helps transform external skills into policy capabilities that require no skill inputs at inference. We will release code, configurations, skill banks, and evaluation protocols.

120. 【2610.09427】Event-Aligned Visual Action Reasoning for World Action Models

链接:https://arxiv.org/abs/2610.09427

作者:Xiaomeng Yang,Yushu Wu,Yi Gao,Yuhao Lei,Xuan Zhang,Pu Zhao,Yanzhi Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:World-Action Models, utilize future visual, intermediate reasoning process, utilize future, process to guide

备注: Project Page: [this https URL](https://xiaomeng-yang.github.io/Event-aligned-WAM/)

点击查看摘要

Abstract:World-Action Models (WAMs) utilize future visual prediction as an intermediate reasoning process to guide action generation. However, existing WAMs typically structure visual imagination according to predefined temporal intervals, without explicitly accounting for the different roles of task-critical interactions and connecting transitions. We argue that effective visual foresight should align directly with task-relevant interactions and their corresponding reasoning demands. To this end, we introduce an event-aligned visual action reasoning framework that organizes visual-action prediction around interaction events. Through event-aligned visual-action supervision, WAM learns to generate event-aligned visual context in each imagined rollout, placing greater emphasis on critical state changes that inform action generation. This shapes the visual reasoning granularity according to the underlying interaction dynamics, with detailed reasoning around task-critical events and coarser progression through connecting transitions. Furthermore, we introduce an execution validity head that identifies the valid portion of each predicted action sequence, avoiding redundant actions during chunked inference. Experiments demonstrate a 10.26 percentage point improvement in DOMINO success rate over baseline and competitive performance on RoboTwin 2.0. It also transfers from DOMINO Level 1 to Levels 2 and 3 without target-level adaptation.

121. 【2610.09418】Spatial Latent Reasoning for Embodied Reference Understanding

链接:https://arxiv.org/abs/2610.09418

作者:Ling Li,Jianhui Zhong,Wei Liu,Zheng Jiang aand Yuxuan Liu,Jingyu Li,Zhidong Deng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires connecting hand, connecting hand geometry, Pointing-gesture visual grounding, grounding requires connecting, Pointing-gesture visual

备注:

点击查看摘要

Abstract:Pointing-gesture visual grounding requires connecting hand geometry with the visual identity and extent of a referred object. A central challenge for continuous latent reasoning is how to organize these complementary cues into useful intermediate supervision. We propose Spatial Latent Reasoning (SLR), a framework that structures this supervision around an ordered sequence of geometric and visual states. A spatial ray state is supervised by fingertip position and pointing direction, followed by four states aligned with target-region features. To construct the visual targets, we introduce parity pooling, which applies polyphase grouping to average region tokens on four interleaved spatial supports. All states are generated recurrently during training and inference; auxiliary annotations are required only during training. On EgoPoint-Ground, the framework improves mIoU over same-backbone supervised fine-tuning by 2.8, 17.5, and 21.1 percentage points on Qwen3.5-4B, Qwen2.5-VL-7B, and Qwen3-VL-8B, respectively, with improvements on both hard subsets. On YouRefIt, it achieves 77.6% precision at IoU 0.5, a numerical margin of 5.2 percentage points over the reported state of the art under differing evaluation protocols. Ablations support joint geometric and visual supervision on the standard and similar-object sets, and favor parity over three alternative pooling operators on the standard set. These results support task-structured supervision for continuous pointing grounding. We will release the code and supporting materials.

122. 【2610.09417】rACT: temporal revelation Airborne Camera Trap

链接:https://arxiv.org/abs/2610.09417

作者:Oliver Bimber,Rakesh John Amala Arokia Nathan,Mohamed Youssef,Vinayak Lal Bhatnagar,Ralf Berger,Klaus Hackländer

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Effective remote monitoring, Airborne Camera Trap, Effective remote, revelation Airborne Camera, dynamic vegetation

备注:

点击查看摘要

Abstract:Effective remote monitoring and surveillance using drones are frequently impeded by severe environmental and thermal clutter, dynamic vegetation, target camouflage, and system latency. Drawing inspiration from the hunting strategies of birds of prey that hover and stabilize their vision to isolate subtle ground motion, we introduce trACT (temporal revelation Airborne Camera Trap), a lightweight, real-time aerial robotics framework designed for autonomous consumer drones. The system integrates Temporal Max Pooling (TMP), a low-level signal processing method that transforms imperceptible movement across a rolling integration window into robust value and time encodings, with self-supervised motion anomaly detection to isolate target motion from background environmental motion caused by wind gusts and drone drift. To overcome mechanical and processing delays, trACT combines motion prediction with automated gimbal-stabilized optical zoom verification and equitable multi-target verification balancing. Extensive real-world field experiments in densely forested wildlife habitats and surveillance scenarios demonstrate that trACT successfully bridges the gap between wide-area aerial monitoring and precise, autonomous target verification under challenging operational conditions.

123. 【2610.09411】Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data

链接:https://arxiv.org/abs/2610.09411

作者:Dominik Schnaus,Thomas Dagès,Daniel Cremers,Xi Wang,Phillip Isola

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:requires large amounts, Multimodal representations enable, Platonic Representation Hypothesis, Multimodal representations, classification and retrieval

备注: Project: [this https URL](https://dominik-schnaus.github.io/unpaired-rosetta/) , Code: [this https URL](https://github.com/dominik-schnaus/unpaired-rosetta)

点击查看摘要

Abstract:Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently align independently trained representations without pairs, and standard geometric alignment metrics accurately predict when this is possible. Nevertheless, we can naturally benefit from paired examples. In the very few-pair regime, our method substantially outperforms existing ones, while staying competitive with pair-based methods with more added examples. Finally, we demonstrate that the resulting alignments can enable text-to-image generation without paired examples. These results show that independently trained models often share enough geometry to establish cross-modal correspondence with little or no paired data.

124. 【2610.09408】ok: Audio-Visual LLM for Multi-Segment Temporal Grounding

链接:https://arxiv.org/abs/2610.09408

作者:Eunji Shin,Dahyun Choi,Seungyeon Jo,Yejin Hong,Jiyoung Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Audio-visual multi-segment grounding, predicting multiple segments, multi-segment grounding, untrimmed videos, remains challenging

备注: ACCV'2026

点击查看摘要

Abstract:Audio-visual multi-segment grounding (AV-MSG) in untrimmed videos, reasoning over audio-visual evidence and predicting multiple segments for a query, is a fundamental problem but remains challenging. Visual-only models overlook complementary acoustic cues, while audio-visual models often fail to calibrate the number of events - a phenomenon we refer to as count miscalibration. We present TiTok, an audio-visual large language model (AV-LLM) that localizes an arbitrary number of temporal event segments for each query. For precise boundary prediction, we introduce the Time Token Interleaving (TTI) method, which explicitly injects special time tokens into the audio-visual stream to align input-side temporal perception with output-side temporal prediction. We further propose decoupled, multi-segment-oriented rewards for reinforcement learning, consisting of global, local, count, precision, and format rewards, optimized with Group reward-Decoupled Normalization Policy Optimization (GDPO). To assess the performance on AV-MSG, we establish a new UnAV-100-based evaluation protocol, and propose the CountF1 metric for quantifying count miscalibration that overlap metrics fail to capture. TiTok reaches 65.7 mIoU and 0.58 CountF1, achieving state-of-the-art performance. Our code is available at this link.

125. 【2610.09397】One Frame, Full Heartbeat: ECG-Free Cardiac Cine MRI Synthesis via Phase-Conditioned Flow Matching

链接:https://arxiv.org/abs/2610.09397

作者:Shiyi Wang,Ruochen Sun,Peirong Liu,Xiang Li,Fangxu Xing

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Cine cardiovascular magnetic, cardiovascular magnetic resonance, Cine cardiovascular, full cardiac cycle, magnetic resonance

备注:

点击查看摘要

Abstract:Cine cardiovascular magnetic resonance (CMR) analysis relies on multi-frame sequences capturing the full cardiac cycle. However, standard multi-frame acquisition depends heavily on electrocardiogram (ECG) gating and repeated breath-holds, posing challenges in uncooperative populations, resource-limited settings, and temporally corrupted datasets. Existing methods that synthesize full cardiac sequences either rely on explicit ECG signals to parameterize myocardium function, or employ deformable registration without physiological constraints, failing to faithfully reproduce clinically relevant dynamic metrics such as ejection fraction (EF) and ventricular contraction magnitude. We present PhaseFlow, a unified generative framework that overcomes both limitations. PhaseFlow estimates a non-linear cardiac phase signal directly from the input sequence via a segmentation-derived left-ventricular (LV) area curve, capturing the asymmetric dynamics of systole and diastole without any ECG dependency. At inference, this phase signal is provided by a pathology-specific template, informing phase-specific frame generation. A rectified flow model conditioned on the phase and slice position synthesizes the full cardiac motion trajectory in the latent space, decoded into a diffeomorphic displacement field that warps end-diastole pixel intensities directly, eliminating the reconstruction blur often accompanying the variational autoencoder. On the ACDC benchmark, PhaseFlow achieves superior physiological fidelity and image realism, with best LV volume curve $R^2$, structural similarity (SSIM) and generative quality (FID) among all baselines. Ablation studies confirm that each proposed component contributes measurably to the overall performance.

126. 【2610.09392】Quantifying Volumetric Risk: Class-Aware Asymmetric Weighted Conformal Prediction for 3D Medical Image Segmentation

链接:https://arxiv.org/abs/2610.09392

作者:Shadi Alijani,Fereshteh Aghaee Meibodi,Homayoun Najjaran

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reliable volumetric segmentation, MedSAM remain deterministic, Reliable volumetric, Weighted Conformal Prediction, conformal prediction

备注:

点击查看摘要

Abstract:Reliable volumetric segmentation is critical for clinical diagnostics, yet foundation models such as MedSAM remain deterministic and lack calibrated uncertainty under distribution shift. Existing conformal prediction methods offer statistical guarantees but are frequently applied in 2D and assume symmetric error distributions, so they do not capture the class-specific biases that arise in 3D multi-class segmentation. We propose Class-Aware Asymmetric Weighted Conformal Prediction (CA-WCP), which combines latent-space density-ratio weighting for covariate shift with directional quantiles for the lower and upper volume bounds, and scales each bound by a class-specific asymmetry factor derived from validation-set false-positive and false-negative rates. We prove that CA-WCP retains the weighted-exchangeability marginal coverage guarantee for every class, and we evaluate it on 3D brain tumor segmentation (BraTS 2020) and on a synthetic multi-organ CT benchmark constructed under covariate shift. On both benchmarks the 95\% Clopper--Pearson interval for the observed coverage of CA-WCP contains the nominal 90\% level for every semantic class, while interval width is reduced by 8--14\% relative to symmetric weighted conformal prediction. We further encode the calibrated intervals into structured prompts for a multimodal large language model to produce uncertainty-conditioned radiology reports, linking distribution-shift-aware uncertainty quantification to interpretable clinical communication.

127. 【2610.09383】PPCAR-Net: Projection-Refined Parametric 3D Coronary Artery Reconstruction from Sparse X-ray Angiographic Views

链接:https://arxiv.org/abs/2610.09383

作者:Yu Ren,Hwee Kuan Lee,Tat-Jen Cham,Jonathan Yap,Khung Keong Yeo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reconstruction commonly relies, directly provide centrelines, vascular-graph extraction, coronary reconstruction commonly, commonly relies

备注: 22 pages, 5 figures, 10 tables, including references and appendix. Code and pretrained models: [this https URL](https://github.com/G2304138H/PPCAR-Net) . Project page with video results: [this https URL](https://G2304138H.github.io/PPCAR-Net/)

点击查看摘要

Abstract:Sparse-view 3D coronary reconstruction commonly relies on cross-view correspondence and triangulation, which are vulnerable to vessel overlap and foreshortening, or on volumetric prediction followed by vascular-graph extraction, which does not directly provide centrelines and radii. We introduce PPCAR-Net, a projection-refined parametric coronary artery reconstruction network that directly predicts a branch-structured centreline-and-radius representation without explicit point matching, triangulation, or an intermediate volume. Given a variable number of segmented views, a coarse predictor combines frozen VGGT features with learned branch queries to estimate branch presence, B-spline centreline trajectories, and dense radius profiles. Projection-guided geometry and radius refiners then sample local evidence from the input views and apply residual corrections learned with 3D supervision. We evaluate representation fidelity and sparse-view reconstruction quantitatively and qualitatively. On simulated angiographic masks generated from CT-derived coronary anatomy, PPCAR-Net produces better connected artery reconstructions and achieves strong centreline accuracy, particularly for RCA, while maintaining competitive volumetric overlap. Coarse-to-fine inference takes 121 ms, enabling real-time reconstruction.

128. 【2610.09382】ScribbleEdit: A Benchmark for Scribble-Only Image Editing

链接:https://arxiv.org/abs/2610.09382

作者:Jie Ren,Hao Kang,Kai Guo,Yiding Yang,Bo Liu,Liming Jiang,Qing Yan,Zichuan Liu,Yizhi Song,Yue Xing,Hui Liu,Xin Lu

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:Scribble-based interaction, interactive editing tools, image editing, image editing models, image editing intents

备注:

点击查看摘要

Abstract:Scribble-based interaction provides a lightweight and intuitive way for users to specify image editing intents in interactive editing tools. However, current image editing models based on VLMs or LLMs struggle to understand and execute edits based solely on scribble inputs. To systematically study this problem, we construct a new benchmark, ScribbleEdit, that evaluates the ability of image editing models to perform image editing conditioned on scribbles. This task requires both a deep understanding of the intention of the scribble and an accurate interpretation of its spatial information. In ScribbleEdit, we design an automated data construction pipeline and introduce a dedicated evaluation protocol that explicitly measures intention alignment. Our analysis reveals that existing VLM/LLM-based editing models fail to accurately capture scribble intentions. To guide future progress on scribble-only image editing, we propose a simple yet effective soft-token baseline, which enhances the model's understanding of scribble semantics and outperforms standard image editing models on our benchmark. Our evaluation and baseline together provide a concrete foundation for assessing and improving the scribble-driven image editing.

129. 【2610.09376】Unified Multi-plane Autoregressive Diffusion for 3D Multi-contrast MRI Synthesis

链接:https://arxiv.org/abs/2610.09376

作者:Yejee Shin,Geonhui Son,Jinglu Wang,Minwoo Jung,Yan Lu,Dosik Hwang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:magnetic resonance imaging, Acquiring a complete, resonance imaging, multi-contrast imaging, uncomfortable for patients

备注: Accepted to ECCV 2026

点击查看摘要

Abstract:Acquiring a complete set of magnetic resonance imaging (MRI) contrasts is time-intensive and uncomfortable for patients, despite the diagnostic value of multi-contrast imaging. This motivates synthesizing missing contrasts from those already acquired, which is an inherently 3D problem requiring anatomical coherence across axial, sagittal, and coronal planes. However, fully 3D generative models are often impracti- cal under computational resources that scale cubically with volume size. We propose a unified Multi-Plane Autoregressive Diffusion (MPAD), a latent diffusion framework that achieves full-volume 3D synthesis using efficient plane-wise 2D operations while preserving volumetric coherence. A 3D autoencoder first compresses MRI scans into an isotropic 3D la- tent representation. A 2D diffusion model is then trained to reconstruct masked latent slices of the target contrast, conditioned on both source- contrast slices and unmasked target-contrast slices. During inference, we introduce plane-wise autoregressive synthesis with inter-plane priors. Slices are generated autoregressively in random order within one plane orientation to maintain intra-plane continuity, then propagated as con- ditioning priors to orthogonal plane orientations to enforce inter-plane consistency. Compared to 3D latent diffusion baselines, MPAD reduces training and inference FLOPs by 7x and 3x, respectively, while also lowering inference time and peak memory consumption. Experiments on multiple datasets demonstrate that MPAD achieves superior perfor- mance, generating high-fidelity 3D volumes and supporting one-to-many translation within a single unified model.

130. 【2610.09363】Closing the Loop on Contrail Avoidance with Satellite Verification

链接:https://arxiv.org/abs/2610.09363

作者:Spandan Ghose Chowdhury

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:thin ice clouds, ice clouds, clouds that aircraft, aircraft leave, model

备注: Accepted into Tackling Climate Change with Machine Learning: workshop at NeurIPS 2026

点击查看摘要

Abstract:Contrails are the thin ice clouds that aircraft leave behind. They cause a large share of aviation's warming, and rerouting the few flights that produce them could avoid much of it. However, an avoided contrail only counts if a satellite can confirm that it never formed, and this check is hard: contrails are one to two pixels wide, cover only 0.18% of pixels, and look very similar to natural cirrus. We build a small diffusion model (8.4M parameters, trained on one GPU) that detects them, and we run a controlled study to find out which components matter. The model reaches 0.476 PR-AUC, compared with 0.414 for a DeepLabV3+ baseline and 0.119 for an adapted MedSegDiff. Doubling the input resolution of the CNN brings it to parity (0.499, p=0.07). Three lessons apply beyond contrails. First, check the input resolution before designing a new architecture. Second, simple flips and rotations more than double accuracy and matter more than any architectural choice we measured. Third, pretraining the model on contrail shapes is harmful: the model learns that thin strokes appear everywhere and paints them onto empty scenes. Precision collapses to 1% while recall-based metrics still rate the degraded model as excellent, and no threshold or guidance heuristic repairs this failure.

131. 【2610.09355】Multimodal LLMs Can Learn to Read Brain Signals: A Vision--Language Model for Unified Multi-Task EEG Decoding

链接:https://arxiv.org/abs/2610.09355

作者:Parastoo Azizeddin,Omid Sharafi,Maryam M. Shanechi

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:recording conditions remains, EEG, challenge in electroencephalography, recording conditions, key challenge

备注:

点击查看摘要

Abstract:Learning EEG representations that generalize across cognitive tasks, subjects, and recording conditions remains a key challenge in electroencephalography (EEG) decoding. Recent advances in foundation models have improved EEG decoding performance, yet a fundamental open question remains: how to effectively interface neural signals with these models to enable multi-task learning across datasets. To investigate this question, we introduce BraVista, a visual-language framework that encodes multichannel EEG signals as structured images and enables multi-task learning through instruction-conditioned vision-language models (VLMs). Our approach relies on continued post-training of a general-domain VLM, leveraging its visual and linguistic priors to adapt to neural signals without a separate large-scale EEG-specific pretraining stage. We evaluate BraVista on four datasets spanning sleep staging, emotion recognition, cognitive workload classification, and abnormal EEG detection, showing strong performance across these tasks. Further analyses show that the choice of EEG-to-image representation is critical to performance. Moreover, through controlled perturbations of the EEG signal, we observe a gradual performance degradation under increasing noise, suggesting that the model relies on EEG-relevant information rather than superficial visual patterns. Together, these findings establish structured visual representations as an effective and scalable interface between neural signals and general-domain foundation models for unified multi-task EEG decoding.

132. 【2610.09344】Do Image Editors Follow Depth-Dependent Blur and Aperture Response? A Rendered-Ground-Truth Pilot Audit

链接:https://arxiv.org/abs/2610.09344

作者:Zhihan Chen,Yuhuan Zhao,Yijie Zhu,Xinyu Yao,Mengcong Ren,Yuchen Sun,Yunqing Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:General image editors, General image, asked to make, make a photo, rarely checked

备注:

点击查看摘要

Abstract:General image editors are asked to make a photo look as if it were taken at f/1.4, yet it is rarely checked whether the blur they add follows thin-lens optics. A physical aperture edit spreads blur across depth in thin-lens proportions and changes the blur when the aperture changes; prior evaluations check blur monotonicity, sharpness-trend correlation, effective-aperture error, or vision-language judgments, and none we found reports the two properties separately at known depths. In this pilot audit of two editors (Gemini~3.1 Flash Image and GPT-image-2.5) we compare against a rendered oracle: Blender Cycles scenes with true thin-lens depth of field, one blur-width estimator applied identically to oracle and editor outputs, and preregistered depth and aperture indices. In 24 texture scenes rendered in one three-panel geometry, accepted and measurable panels show f/1.4-to-f/2.8 width ratios, $\bar\sigma_{1.4}/\bar\sigma_{2.8}$, of 0.99--1.18 against 1.98--2.13 for the oracle, and ratios of pooled median Gaussian-equivalent near/far blur widths of about 1.18--1.25 (Gemini) and 0.97--1.05 (GPT-image) against 1.69--1.82. Preregistered black-box interventions show that qualitative wording changes blur strength by roughly 2--10 times, whereas a request for 2 versus 6 px changes it 1.1--1.3 times and no tested wording of the f-number meets the registered ``followed'' criterion. The depth compression appears in the original and reversed centre-focus layouts; with the focus on the near panel, the available-panel depth index reaches the registered threshold, and the aperture response stays attenuated in every layout tested. Scalar metrics adapted from published ones give oracle-like scores to synthetic editors whose proportions are compressed.

133. 【2610.09343】Skipper: Region-Adaptive Tile Pruning for 3D Gaussian Splatting

链接:https://arxiv.org/abs/2610.09343

作者:Jingxing Li,Yongjae Lee,Deliang Fan,Abhay Kumar Yadav,Cheng Peng,Rama Chellappa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scene-wide contribution cutoff, Gaussian Splatting rasterizers, Gaussian Splatting, tile enumeration, support truncation

备注:

点击查看摘要

Abstract:Tiled 3D Gaussian Splatting rasterizers often use one scene-wide contribution cutoff for tile enumeration, although content differs in its sensitivity to support truncation. TileSkipper selects a static per-Gaussian cutoff policy for a frozen checkpoint. Calibration renders measure candidate pair savings and an isolated-removal distortion proxy that accounts for front transmittance and background color. The method allocates cutoffs across 64 Gaussian groups and accepts policies only after complete renders on disjoint selection views. The exported policy uses one byte per Gaussian, with no parameter updates, additional kernel, or per-frame policy inference. Across 13 scenes from Mip-NeRF 360, Tanks Temples, and Deep Blending, a fixed-policy AccuTile sweep gives dataset-macro speedups of $1.088\times$ at standard resolution and $1.238\times$ at 3840 pixels wide, with $-0.007/-0.023$ dB mean PSNR change. Six integrations with existing opacity-aware bounds yield $1.009\times$--$1.121\times$ compiler-only speedups. For four ports from $3\sigma$ rasterizers, we separately attribute the prior exact-bound transition and our incremental gain. Matched-quality ablations show modest gains over scene-global calibration and parity with per-Gaussian control; the standalone comparison with AdaGScale is regime-dependent.

134. 【2610.09330】CRT-HMAR: Causal Requirement Tracing-Guided Hierarchical Multi-Agent Regulation for Open-Task-Aware Infrared-Visible Image Fusion

链接:https://arxiv.org/abs/2610.09330

作者:Zengyi Yang,Shuai Yuan,Zhong-Cheng Wu,Juan Cheng,Huafeng Li,Yu Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Infrared and visible, integrates complementary multimodal, single fused image, fusion integrates complementary, IR-VIS image fusion

备注: 17 pages, 11 figures

点击查看摘要

Abstract:Infrared and visible (IR-VIS) image fusion integrates complementary multimodal information into a single fused image to support downstream vision tasks. However, existing methods are typically tailored to seen tasks within a fixed task set and struggle to generalize to unseen tasks, which restricts their applicability in real-world open-task scenarios. To address this issue, this paper proposes CRT-HMAR, a Causal Requirement Tracing-Guided Hierarchical Multi-Agent Regulation Framework for open-task-aware IR-VIS image fusion. CRT-HMAR introduces a Causal Requirement Tracing Task Localization mechanism, which actively intervenes in key image information and observes task-network response variations to map task-specific semantic preferences into image-level causal requirement maps. Based on these maps, a requirement analysis agent aggregates task-specific requirement knowledge to adaptively guide requirement-customized image fusion. Moreover, CRT-HMAR incorporates History-Analysis Multi-Objective Balancing and Task-Level-Correction Conflict Mitigation mechanisms, jointly constructing a hierarchical regulation chain of "requirement interpretation - task balancing - conflict mitigation". Through multiple collaborative agents, CRT-HMAR dynamically regulates key processes including open-task requirement modeling, multi-task balanced optimization, and gradient conflict mitigation. Extensive experiments on open-task scenarios involving five downstream tasks demonstrate that CRT-HMAR significantly improves generalization to unseen tasks while maintaining the performance and balance of seen tasks. Overall, CRT-HMAR shifts IR-VIS image fusion from task-oriented modeling toward requirement-oriented modeling, promoting its extension from closed-task settings to real-world open-task scenarios.

135. 【2610.09329】GRC-Net: Global Representation Consistency Network for Unsupervised Multimodal Anomaly Detection

链接:https://arxiv.org/abs/2610.09329

作者:Seyoung Jeong,Jong Pil Yun,Sang Jun Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Automated quality inspection, http URL image-based, motivating multimodal approaches, multimodal approaches incorporating, URL image-based methods

备注: 5 pages, 3 figures, Under Review

点击查看摘要

Abstract:Automated quality inspection is essential for ensuring product reliability in this http URL image-based methods effectively capture appearance-related defects, these methods are limited in detecting structural and geometric anomalies, motivating multimodal approaches incorporating 3D information. However, existing methods mainly rely on local patch-level representations, which often lead to unstable reconstruction errors even in normal regions. To address this limitation, we propose GRC-Net, which integrates a global-attention MLP to enforce global representation consistency across patch embeddings with a stable reconstruction module to improve reconstruction stability. The proposed method captures holistic contextual information through a global token and suppresses reconstruction noise by minimizing discrepancies between original and predicted embeddings. Experiments on MVTec 3D-AD and Eyecandies demonstrate that GRC-Net consistently outperforms existing methods at both image and pixel levels. Qualitative results further demonstrate reduced reconstruction errors in normal regions and more distinct reconstruction differences between normal and anomalous regions.

136. 【2610.09328】Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation

链接:https://arxiv.org/abs/2610.09328

作者:Baoteng Li,Wenzhuo Wu,Kongming Liang,Zhanyu Ma

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:requested attributes, image generation requires, Visual Jev, Multi-subject image generation, generation requires rewards

备注:

点击查看摘要

Abstract:Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn these visual decisions into generator training signals. Each subject-related question receives a positive label only when the requested condition and the relevant reference identities hold jointly. We construct fixed questions offline, train a Qwen3.5-4B verifier with binary supervision, and directly read Yes probabilities from its language-model head. Their mean supplies a GRPO reward while retaining individual judgments for inspection. Using 200 MICo-150K training tasks and 30 updates, the framework raises a GPT-5.4 composite score from 41.78 to 52.50 on a manually selected 897-task MICo-Bench subset; direct 27B rewards yield 51.84. Each reward is tested in one GRPO run, and offline human evaluation does not establish a statistically significant advantage over direct scoring. The study provides an initial implementation and evaluation of Visual Jev as a reference-bound reward for multi-subject image generation.

137. 【2610.09326】VIS-Ground: Video Interactive Storytelling with Contextual Grounding

链接:https://arxiv.org/abs/2610.09326

作者:Bingxuan Li,Yiwen Song,Xueqing Wu,Yanzhou Pan,Yang Li,Kuang Su,Jingyun Liu,Sebastian Ko,Huan Zhang,Tong Zhang,Nanyun Peng,Tomas Pfister,Yale Song

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:storytelling enables viewers, interactive storytelling enables, video generation, Video, generation

备注: Project Page: [this https URL](https://bx126.github.io/vis-ground.github.io)

点击查看摘要

Abstract:Video interactive storytelling enables viewers to actively steer how a video unfolds. However, once we allow viewers to intervene during generation, a new challenge arises: The viewer's request can have latent dependencies on both the grounding source and the current rendered video state. These dependencies may not be explicitly stated in any individual input, but emerge only when the source, rendered history, and new viewer intent are considered jointly. Existing interactive video generation systems primarily emphasize following viewer instructions, while source-grounded video generation methods focus on aligning generated content with an external narrative or knowledge source. This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context? In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation. To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering. Across three video generation backbones, VIS-Ground consistently achieves the highest overall composite score, reaching an average absolute improvement of 10.3 points over the strongest per-backbone baselines. Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.

138. 【2610.09317】SAREO-FM: Decoupled Semantic Supervision for SAR-EO Foundation Models

链接:https://arxiv.org/abs/2610.09317

作者:Jeonghyeok Do,Munchurl Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Synthetic aperture radar, provides rich appearance, Synthetic aperture, fine-grained semantic cues, imagery provide complementary

备注: Please visit our project page at [this https URL](https://kaist-viclab.github.io/SAREO-FM_site/)

点击查看摘要

Abstract:Synthetic aperture radar (SAR) and electro-optical (EO) imagery provide complementary observations: SAR enables day-and-night, weather-resilient sensing, whereas EO provides rich appearance and fine-grained semantic cues. We introduce SAREO-FM, which avoids forcing a single token stream to serve two distinct roles: modality tokens preserve how each sensor observes the scene through masked reconstruction, while learnable semantic queries capture what the scene contains under guidance from a pretrained vision foundation model (VFM). By jointly encoding these queries with SAR and EO tokens, the queries acquire modality-grounded semantic context, while the modality-token outputs remain the explicit targets of masked reconstruction. This design assigns semantic and reconstruction supervision to separate token streams while preserving their interaction within the shared encoder. Pretrained on the million-scale SAR-1M corpus, SAREO-FM achieves strong unimodal transfer for both SAR-only and EO-only inputs, while delivering substantial gains from joint SAR--EO observations on tasks that benefit from complementary sensing.

139. 【2610.09313】Why VLMs Miss Small Objects, and When Zooming In Is Safe

链接:https://arxiv.org/abs/2610.09313

作者:Junzhe Shi,Yuan Gan,Shida Jiang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:miss small objects, Vision-language models, object, image, miss small

备注:

点击查看摘要

Abstract:Vision-language models (VLMs) often miss small objects in large images. We ask three questions: what limits them, which of these limits better models can remove, and whether the classical way of handling large images, local decomposition, still has a future. We answer them with a theory built on two quantities of the image interface: S, the number of visual tokens across an object's side, and L, the content a call must cover. The limits: a W x H image sent whole within N tokens gives an object of side m at most m*sqrt(N/(WH)) tokens per side. Doubling the token budget N therefore raises the largest S a whole image can reach by only 41%, and seeing an image at the S an object needs costs at least order S^2 tokens whatever the model. If recognition improves gradually with S, any search strategy, zoom agents included, obeys a recall-cost frontier. What better models can change: the S an object needs and how much one call can carry, measured in bits per object found; the coverage cost remains. Decomposition: yes. Assuming only that more tokens per object and less content per call do not hurt on average, splitting an image cannot lower recall if no view zooms out relative to the whole image and views overlap by one object. Neither part of this condition can be dropped; we bound the cost of every such decomposition, and a simple rule approaches the bound as the image grows. We test the theory in about 177,000 requests on 797 images. On controlled images, none of 40 orderings predicted for 8 VLMs is violated. Checked after the fact on drawings, floor plans, natural and synthetic images, 61 of 89 implied orderings hold significantly and 4 fail, all for OpenAI models given more pixels than their default path. On construction drawings the rule never significantly lowered recall relative to the whole image and raised it by up to 0.28. Code and data: this https URL

140. 【2610.09305】Kuration SDK: Addressing the Virtual2Real Gap via Data Curation

链接:https://arxiv.org/abs/2610.09305

作者:Nirmit Desai,Eric Song,Mayank Sengupta,Tejal Bedmutha,Siri Reddy,Sahiti Dharmavaram,Kunal Sawarkar

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:visual similarity-based metrics, action-conditioned world models, action-conditioned world model, world models, action-conditioned world

备注:

点击查看摘要

Abstract:Benchmarks for measuring the quality of action-conditioned world models are still evolving and shifting away from visual similarity-based metrics to action-semantic and physically-grounded metrics. However, for domain and task-agnostic action-conditioned world model training, existing benchmarks provide a limited signal. By training and evaluating diffusion world models on CounterStrike gameplay data, we confirm that qualitative playability does not correspond with metrics such as FVD, LPIPS, and JEDi. We term this the Virtual2Real gap. We posit that, in lieu of reliable benchmarks, curating raw gameplay data and measuring a variety of diagnostic properties provides a more robust signal to bridge the gap, before the training even begins. We present several curation strategies and a general-purpose kit for physical AI data curation called Kuration SDK, which is being open-sourced with this paper. The SDK was instrumental in uncovering the root cause of the virtual2real gap in a specific case: why two world models trained on identical gameplay map, action and state distribution, behaved very differently when played in spite of having very similar LPIPS and FVD scores. Thus, Kuration SDK has the potential to uncover the root causes of Virtual2Real gap in specific datasets and accelerate development of sample-efficient training datasets.

141. 【2610.09304】mg2face: Expressive Facial Animation with High-Density Surface EMG

链接:https://arxiv.org/abs/2610.09304

作者:Ganidhu Abey,Wendy Greening,Ashika Kamboj,Leonhard Helminger,Abhijeet Ghosh,Karel Petranek,Sergio Orts Escolano,Dinesh K. Pai

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Machine Learning (cs.LG)

关键词:movements convey subtle, human social communication, social communication, convey subtle, subtle and important

备注: 12 pages plus supplmentary material

点击查看摘要

Abstract:Facial movements convey subtle and important information that is critical for human social communication. Optical methods for face capture are difficult or impossible to use when the face is occluded by head-mounted devices (HMDs), such as VR headsets. Even with a clear line of sight, such methods raise privacy concerns and require head-mounted capture rigs that offset cameras and lighting from the face. We show that high-density surface electromyography (HD-sEMG) provides a viable non-optical alternative that addresses these challenges. We measured 64 EMG channels, using two textile EMG grids, with 32 from the forehead (typically occluded by an HMD) and 32 from the side of the face. EMG data were digitized at 2048 Hz and filtered. Facial movements were simultaneously recorded and used to estimate 478 3D facial landmarks using MediaPipe's Face Landmarker. A major challenge in such multimodal recordings is synchronizing EMG and video data, which have different sampling frequencies and independent clocks. We developed a novel synchronization method using analog audio bursts that is capable of sub-millisecond synchronization. We also developed a staged fitting method that fits a recent high-resolution parametric head model (GNM), with 253 identity blendshapes and 383 expression blendshapes, to the MediaPipe landmarks as participants performed different facial expressions. We trained a deep neural network comprising per-grid spatial encoders followed by a dilated temporal convolutional network (TCN) to predict blendshape parameters from HD-sEMG signals at 100 Hz. Once trained, the network can predict expression blendshapes solely from HD-sEMG recordings. The output can be rendered using standard real-time blendshape animation methods. We demonstrate the methods using recordings from 25 participants, and direct expression transfer to a variety of human faces and non-human characters.

Comments:
12 pages plus supplmentary material

Subjects:

Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET); Machine Learning (cs.LG)

Cite as:
arXiv:2610.09304 [cs.GR]

(or
arXiv:2610.09304v1 [cs.GR] for this version)

https://doi.org/10.48550/arXiv.2610.09304

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
142. 【2610.09285】LeCuration: A Tiny World Model as a Data Curation Multi-Tool

链接:https://arxiv.org/abs/2610.09285

作者:Mayank Sengupta,Nirmit Desai,Eric Song,Kunal Sawarkar

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:governing object behavior, laws governing object, physical laws governing, closed physical worlds, object behavior

备注: 9 pages

点击查看摘要

Abstract:Many applications of physical AI run within finite or closed physical worlds with a limited set of physical laws governing object behavior. Examples include robots working in a warehouse and agents moving around in a video game. In order to better organize, filter, and curate data for physical AI applications, we propose a new approach centered on the unique settings and physical laws of individual datasets. We train LeCuration, a small world model intended to serve as a data curation tool for a separate, larger downstream model. To build this model, we choose LeWorldModel (LeWM)as our latent encoder and predictor, adding a diffusion transformer (DiT) decoder to add visuals to autoregressive gameplay rollout. We find that the embeddings of this model can be used as an anomaly detection signal and as a content-based clustering heuristic, and that auto-regressively predicting the game state with this model allows us to qualitatively check for action-state consistency. This paper presents a qualitative, proof-of-concept case study on CS:GO gameplay data; we do not yet report quantitative curation metrics or downstream training results, which we identify as the key next step.

143. 【2610.09275】EM-SNN: Efficiently Modulated Spiking Neural Network for Remote Sensing Image Dehazing

链接:https://arxiv.org/abs/2610.09275

作者:Jie Shao,Jiaqi Ma,Wenwen Min,Beihang Song,Ning Chen,Youfa Liu,Jun Wan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:artificial neural networks, dehazing remains limited, neural networks, spiking neural networks, Spiking Neural Network

备注:

点击查看摘要

Abstract:Although spiking neural networks (SNNs) provide an energy-efficient alternative to artificial neural networks (ANNs), their application to remote sensing image dehazing remains limited. A key challenge arises from the coupling between haze-induced high-frequency attenuation and discrete spike thresholding. This interaction suppresses weak responses and fundamentally limits the recovery of edges, textures, and fine details in spiking dehazing models. To address this challenge, we propose the Efficiently Modulated Spiking Neural Network (EM-SNN), a dedicated spiking framework tailored to remote sensing image dehazing. EM-SNN integrates a statistics-driven Threshold-Modulated Leaky Integrate-and-Fire (TM-LIF) neuron to adaptively compensate for haze-induced contrast compression, together with a Spike Sobel Modulation (SSM) module that enhances structural cues and reduces depth-wise attenuation during spiking feature propagation. By jointly modulating activation scales and structural representations, EM-SNN improves dehazing performance while preserving the inherent event-driven sparsity of SNNs. Experiments on HRSD, RICE, RRSHID, and SateHaze1K demonstrate that EM-SNN achieves competitive dehazing performance while consuming only one quarter of the energy of the strong ANN baseline SFRDP-Net.

144. 【2610.09274】Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers

链接:https://arxiv.org/abs/2610.09274

作者:Weitian Wang,Shubham Rai,Cecilia De La Parra,Akash Kumar

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Geometry Grounded Transformer, Visual Geometry Grounded, Grounded Transformer, Visual Geometry, Geometry Grounded

备注: Accepted to IJCNN26

点击查看摘要

Abstract:The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63$\times$ and the whole backbone by 1.77-2.35$\times$ with negligible loss (1%) for large scenes. With a small performance loss ( 5%), calibrated BC attention further achieves a 2.26-2.87$\times$ latency improvement on the global attention layers and a 1.90-2.55$\times$ improvement on the backbone.

145. 【2610.09253】PhyDiCT: Plug-and-Play CT Reconstruction from Sparse X-Rays via Differentiable Rendering and Strong Priors

链接:https://arxiv.org/abs/2610.09253

作者:Weicheng Dai,Shantanu Ghosh,Kayhan Batmanghelich

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Computed Tomography, highly ill-posed inverse, ill-posed inverse problem, inverse problem due, volumetric information

备注: Accepted at MICCAI 2026; to appear in LNCS 16888

点击查看摘要

Abstract:Reconstructing 3D Computed Tomography (CT) images from a few X-ray projections is a highly ill-posed inverse problem due to the loss of volumetric information. We propose PhyDiCT, a training-free framework that integrates a differentiable Physics-based forward model, grounded in the Beer-Lambert law, with a text-conditioned Diffusion as a strong prior to reconstruct 3D lung CT images. We refer to our approach as training-free since the prior model is used without fine-tuning, and our goal is to steer the denoising procedure to generate samples consistent with X-ray observations. We guide the diffusion generation using Split Gibbs sampling to jointly optimize for projection fidelity (reward) and consistency with prior knowledge. Also, we introduce a test-time refinement step that enhances image realism and anatomical coherence. We extensively evaluate our method on publicly available 3D CT datasets using both perceptual and semantic metrics, demonstrating that it surpasses existing plug-and-play diffusion and fully trained reconstruction approaches. Our findings highlight that combining a strong generative prior with the underlying physics of image formation substantially improves reconstruction quality, e.g., 7.5\% improvement on SSIM compared to full training methods. Code will be released at this https URL.

146. 【2610.09252】Adaptive Visual Token Reduction for Accelerated Image Understanding

链接:https://arxiv.org/abs/2610.09252

作者:Seyoung Jeong,Jong Pil Yun,Sang Jun Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Vision-Language Models, Vision-Language Models achieve, Models achieve strong, information-rich images requires, requires substantial computation

备注: 5 pages, 2 figures. Under review

点击查看摘要

Abstract:Large Vision-Language Models achieve strong VQA performance, but processing high-resolution, information-rich images requires substantial computation, motivating visual token reduction. However, existing methods often prune individual tokens or rely on fixed-size cropping, limiting their ability to preserve spatially structured information such as horizontally or vertically elongated text. To address this limitation, we propose ReFIT, an instruction-guided visual token reduction framework for efficient LVLM inference. ReFIT consists of Relevance-Guided Window Reshaping (RWR) and Instruction-Guided Token Refinement (ITR), where RWR captures instruction-relevant regions by adapting to their spatial characteristics, while ITR further removes unnecessary visual tokens. Experiments on four VQA benchmarks demonstrate that ReFIT improves answer accuracy while reducing computational cost, and qualitative results demonstrate its effectiveness in localizing relevant regions and removing unnecessary visual information.

147. 【2610.09242】Pooling Representation Autoencoders for Efficient Diffusion

链接:https://arxiv.org/abs/2610.09242

作者:Ramón Calvo-González,Youssef Saied,François Fleuret

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Representation Autoencoders, pre-trained visual fea, generative modeling expensive, grids make generative, make generative modeling

备注:

点击查看摘要

Abstract:Representation Autoencoders (RAEs) generate images from pre-trained visual fea- tures, but their dense token grids make generative modeling expensive. Motivated by local feature correlations, we introduce PoolDINO, a learned affine pooling operator that merges neighboring tokens. Training the pooling operator jointly with the RGB decoder preserves the standard two-stage RAE procedure without a separate feature auto-encoder. On ImageNet-256, 4x token compression retains comparable generation quality under internal guidance, while 16x compression trades some quality for greater efficiency. At a fixed budget of 100 sampling steps, latent-sampling throughput increases by 3.7x and 9.0x, respectively, relative to the unpooled baseline. Classification and dense prediction evaluations show that comparable guided generation quality can coexist with weaker performance on other tasks.

148. 【2610.09223】PVSync: A Unified Lip-Sync Expert for Timing and Articulation

链接:https://arxiv.org/abs/2610.09223

作者:Kevin Stephen,Varun Menon,Timo Mertens,Nikita Drobyshev

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Lip movements, spoken sounds, movements can match, match the timing, matching the spoken

备注:

点击查看摘要

Abstract:Lip movements can match the timing of speech without matching the spoken sounds. We introduce PVSync, a unified model for audio-visual offset estimation and phoneme-level articulation scoring. PVSync combines window-level contrastive learning for synchronisation with a phoneme-level articulation objective that aligns audio and video embeddings of the same viseme class across clips. Visemes group phonemes with similar visible articulation. Viseme labels are derived automatically from forced-aligned transcripts, without manual annotations. On offset-corrected videos from 13 talking-head video generation models, PVSync matches human rankings of lip-sync quality more closely than LSE-C, achieving a Spearman correlation of 0.83 versus 0.34. On an automatically constructed benchmark from held-out speech, PVSync distinguishes viseme-matched from mismatched audio-visual pairs with an ROC AUC of 0.91. PVSync also outperforms SyncNet and MTD-VocaLiST in temporal offset recovery on held-out in-the-wild clips. Code and benchmark data will be released upon acceptance.

149. 【2610.09221】Consistent Distribution Matching for Data-Free Diffusion Distillation

链接:https://arxiv.org/abs/2610.09221

作者:Yuxiang Fu,Qi Yan,Zike Wu,Yongxing Zhang,Purang Abolmaesumi,Lele Wang,Renjie Liao

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:expensive numerical integration, slow inference due, computationally expensive numerical, numerical integration, suffer from slow

备注:

点击查看摘要

Abstract:Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher's dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, costly teacher rollouts, or auxiliary proxy networks, which complicate model training and scaling. In this work, we propose Consistent Distribution Matching, a simulation-free and data-free distillation method for accelerating diffusion and flow models while preserving strong generative capacity. Our key insight is to unify sample generation and score estimation with one student network. Thus, our framework uses only two models, a frozen teacher and a trainable student, and optimizes one objective. We prove that minimizing our objective indicates Wasserstein convergence of the student flow-map pushforwards to the teacher marginals. On ImageNet 256$\times$256, our method attains an FID of 2.04 with a single function evaluation (1-NFE) and a 4-NFE FID of 1.37 within 40 epochs of training, surpassing the state-of-the-art distillation baselines without data. Our code code and model are available at this https URL.

150. 【2610.09217】A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

链接:https://arxiv.org/abs/2610.09217

作者:Wenqi Li,Mindi Ruan,Chuanbo Hu,Shuo Wang,Xin Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Autism spectrum disorder, child social behavior, Autism spectrum, spectrum disorder, early identification

备注:

点击查看摘要

Abstract:Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~0, across eight pipeline configurations on one backbone, 16--37\% of clips change their predicted label between repeated runs, and the cause lies in the serving stack. We keep the VLM frozen and move the decision out of the model. Under our grounded perception constraints the VLM writes an event table of timestamped, glossary-labeled events that names the eliciting press and logs counter-evidence; a text-only stage age-calibrates the confidence of each row; a deterministic weight-of-evidence scorer sums it into an evidence total and stratifies it into a risk category, so every decision decomposes into named per-feature contributions and can be re-scored from the saved table. On 43 caregiver-recorded, protocol-free home free-play clips of preschool children, the pipeline reaches AUC $0.851 \pm 0.012$, 86.0\% accuracy, and F$_1$ 71.8 over three runs. It labels 74.4\% of clips correctly in every run (60.5\% for the zero-shot baseline) and flags no typically developing clip in every run (9 of 31 at zero-shot). An ablation on the same backbone attributes the gain to the grounded perception constraints read through the deterministic scorer.

151. 【2610.09205】Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency

链接:https://arxiv.org/abs/2610.09205

作者:M. Moein Esfahani,Sepehr Salem,Mohammed Alser,Vince Calhoun

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Modern image editing, image editing models, Modern image, editing models, models can satisfy

备注: Accepted in NeurIPS 2026 the 1st Workshop on Physical World AI

点击查看摘要

Abstract:Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited image violates scene physics, naming the violation type, localizing the affected region, and explaining the failure in language. We introduce a counterfactual benchmark whose controlled synthetic component uses Mitsuba~3 to generate 5,500 images from 500 scene families. Each family contains one clean image and ten matched violations involving shadows, reflection, support, surface response, and occlusion. The renderer pipeline provides category labels, affected-region masks and boxes, scene metadata, and explanation targets. We use LLaVA-1.5-7B, Qwen2.5-VL-7B, and InternVL3.5-8B as diagnostic baselines rather than proposed methods. On a 1,650-image synthetic test set, the adapted baselines reach 64.0--67.8\% category macro-F1 on standard held-out scenes. For LLaVA-1.5-7B, category macro-F1 falls from 64.0\% on the standard split to 40.8\% under intervention shift. This gap shows that high in-distribution accuracy partly reflects cues tied to rendering and counterfactual construction.

152. 【2610.09200】StyleFields: Multi-Scale AdaIN-Modulated Implicit SDFs for Coarse-to-Fine 3D Shape Reconstruction and Editing

链接:https://arxiv.org/abs/2610.09200

作者:Ehsan Garaaghaji,Nicolas Talabot,Pascal Fua,Doruk Oner

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:geometric style mixing, Adaptive Instance Normalization, enables controllable geometric, controllable geometric style, multi-level Adaptive Instance

备注: 39 pages, 20 figures, 3 tables. Includes supplementary material

点击查看摘要

Abstract:We introduce StyleFields, a DeepSDF-based architecture for high-fidelity 3D reconstruction that enables controllable geometric style mixing: the coarse structure of one object can be combined with the fine-scale details of another. The core idea is depth-aware modulation: instead of a single global code, we inject latents via multi-level Adaptive Instance Normalization at several decoder depths, and supervise matching auxiliary heads with a coarse-to-fine schedule while gradually growing network depth. This aligns early layers with global shape and later layers with high-frequency detail, achieving content-style decoupling without part labels or adversarial training. StyleFields delivers faithful reconstructions, convincing cross-instance hybrids, and consistent gains in ablations over injection depth and supervision granularity. We further demonstrate a practical application in automotive aerodynamics: a learned surrogate drag predictor serves as a differentiable objective to optimize reconstructed cars, allowing targeted edits of global form or surface details by freezing the complementary latent stream. StyleFields offers a simple, effective recipe for controllable implicit reconstruction and downstream performance-driven design.

153. 【2610.09195】StableGrasp: Reconstructing Physically Stable Human Hand Grasps from Single Images

链接:https://arxiv.org/abs/2610.09195

作者:Han Jiang,Etienne Vouga,Qixing Huang,Georgios Pavlakos

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR); Robotics (cs.RO)

关键词:single RGB image, problem requires estimating, single RGB, constrained hand pose, visually constrained hand

备注:

点击查看摘要

Abstract:Reconstructing a physically stable human grasp from a single RGB image is challenging because physically modeling grasps is itself difficult, and the problem requires estimating not only a visually constrained hand pose but also a control target that stabilizes the grasp. Existing methods either model only visual hand geometry without considering physics, or rely on less plausible physical modeling, which limits the physical validity of the resulting grasps. In this paper, we present StableGrasp, a differentiable simulation-based optimization framework that explicitly separates the visual hand pose from the control target that determines the grasping forces. Our method jointly optimizes hand geometry and control by minimizing the kinetic energy of the grasp in a differentiable simulator, while regularizing the hand geometry to preserve visual consistency and geometric plausibility. The reconstructed grasps are substantially more stable under rigorous physical simulation, while remaining visually consistent with the input images and geometrically plausible. Experiments show that our approach produces far more stable grasps than alternative hand-control strategies, benefiting visual-only grasp reconstruction pipelines by turning their outputs into physically stable grasps.

154. 【2610.09191】R-CNN-Based Chess Position Recognition

链接:https://arxiv.org/abs/2610.09191

作者:Paras Govind,Ognjen Arandjelović

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Performing chess game, Performing chess, chess game position, three-dimensional board requires, board requires predicting

备注:

点击查看摘要

Abstract:Performing chess game position recognition solely from a single image of a three-dimensional board requires predicting the position and orientation of the board relative to the camera, the occupancy of squares and the piece type, which includes its colour. We propose an R-CNN-based framework with independent components for piece recognition and board geometry estimation, whose predictions are combined to reconstruct the position. For piece recognition, we adapt Faster R-CNN using a class-weighted objective and a deeper classification head. The detector operates directly on the input image, retaining alternative piece hypotheses that are subsequently refined using constraints on piece counts and square occupancy. For board detection, we introduce an octagonal arrangement of eight labelled boundary keypoints, predicted using the keypoint head of Mask R-CNN. These provide redundant correspondences for homography estimation and encode board orientation. The estimated homography maps representative points from the piece boxes to an 8x8 grid. On a synthetic dataset, the modifications to piece detection increase mean average precision from 61.59% to 90.14%. Of the predicted board keypoints, 97.11% are within 1% of the image diagonal of their labelled targets. Using ground-truth piece boxes with the predicted homographies gives correct square assignments for every test position. The complete framework recovers 76.61% of test positions exactly and 96.49% with at most one incorrect square.

155. 【2610.09185】One Frame, Full Heartbeat: ECG-Free 4D Cardiac Cine MRI Synthesis via Radial-Decomposed Flow Matching

链接:https://arxiv.org/abs/2610.09185

作者:Shiyi Wang,Ruochen Sun,Xiang Li,Peirong Liu,Fangxu Xing

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cardiovascular magnetic resonance, acquisition requires electrocardiogram, repeated breath holds, standard acquisition requires, Cine cardiovascular magnetic

备注:

点击查看摘要

Abstract:Cine cardiovascular magnetic resonance (CMR) captures the cardiac cycle as a four-dimensional (4D) sequence, but standard acquisition requires electrocardiogram (ECG) gating and repeated breath holds. Visual realism alone does not establish accurate patient-specific ejection fraction (EF) or ventricular volumes. We present PhaseFlow3D, a generative framework that synthesizes a complete 4D cine sequence from a single end-diastolic (ED) three-dimensional (3D) volume without ECG. To capture asymmetric systolic and diastolic dynamics, it represents the cardiac cycle as a piecewise linear phase anchored at ED and end-systolic (ES) time points. At inference, a population-level canonical template supplies this phase without patient-specific temporal information. A phase-conditioned rectified flow model generates a cardiac motion trajectory in latent space. Radial Contraction Decomposition converts each latent state into a 3D displacement field, combining a physics-informed radial component for centripetal myocardial contraction with an image-conditioned residual for rotation and out-of-plane motion. Each frame is generated by directly warping the ED volume, bypassing variational autoencoder decoding. On the combined ACDC and MMs benchmark, PhaseFlow3D achieves the lowest EF mean absolute error, the only positive left-ventricular volume-curve $R^2$, and the best distributional quality among compared methods. Ablations confirm each component's contribution. Downstream evaluations demonstrate the utility of the synthesized sequences and displacement fields for segmentation, pathology classification, label propagation, and myocardial strain analysis.

156. 【2610.09173】RDGSplat: Render-Dedicated Geometry for Novel View Synthesis

链接:https://arxiv.org/abs/2610.09173

作者:Zhijie Zheng,Xinhao Xiang,Jiawei Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:carrying a Gaussian, Gaussian head, models enable efficient, foundation models enable, enable efficient

备注:

点击查看摘要

Abstract:3D foundation models enable efficient novel view synthesis by carrying a Gaussian head on the representation they already use for reconstruction. However, the views they render fall short of the geometry they recover, because that geometry is estimated under a metric objective and never scored on how it renders. Recent methods alleviate this by updating the backbone weights, but they thereby discard the metric predictions the model was built for and must be repeated for every new backbone. To this end, we propose RDGSplat, a framework that decodes a second geometry dedicated to rendering from a frozen 3D foundation model, leaving its metric predictions intact. In particular, we devise Render-Dedicated Geometry Decoding, which duplicates the pretrained decoders and optimizes the duplicates under photometric supervision alone. Then, a Target-Pose Conditioned Adapter is introduced to reformulate the representation those decoders read, conditioned on the target camera pose rather than the target image. Extensive experiments show that RDGSplat improves novel view synthesis across three feed-forward backbones on four benchmarks, with every pretrained weight frozen. On RE10K, it raises WM2.0 from 20.918 to 24.266\,dB while training 205.5\,M added parameters against a frozen 1.4\,B backbone, and the depth and pose the same model predicts are unchanged.

157. 【2610.09125】Depth-to-RGB: Repurposing a Frozen Depth Estimator for Geometry-Guided Compositing

链接:https://arxiv.org/abs/2610.09125

作者:Sanghyun Jo,Chae Yeon Lim,Donghwan Lee,Sihyun Kim,Soo Ye Kim,Kyungsu Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reference-based object compositing, Reference-based object, background image, object compositing inserts, inserts or replaces

备注:

点击查看摘要

Abstract:Reference-based object compositing inserts or replaces an object using a background image, a reference image, and a 2D compositing mask. These inputs guide appearance and placement but leave the completed scene's geometry implicit, which can distort object structure or alter the surroundings. Our Depth-to-RGB (D2R) framework predicts composite depth for a scene not yet observed in the RGB inputs. It learns reference-conditioned corrections to a frozen depth estimator using encoder features of paired completed scenes as targets. The unchanged decoder maps the corrected representation to the intended scene's depth, which a separately trained renderer holds fixed during RGB synthesis. Under matched architecture and training, encoder-feature supervision reduces OOD Stage-1 AbsRel by 31.4% relative to decoded-depth supervision. We also introduce AnyInsertion++ with paired in-distribution and category-disjoint splits to evaluate generalization beyond compositing training categories. The complete D2R system leads 12 open-source and 3 closed-source baselines in estimator-derived geometry and photometric quality on both paired splits. On category-disjoint data, D2R reduces AbsRel by 43.7% and improves PSNR by 2.4 dB over the matched RGB baseline. Across three unpaired benchmarks, D2R leads both identity metrics and reduces mean CLIP reference cosine distance by 55% relative to the strongest baseline. Project page: this https URL

158. 【2610.09118】opoCurve: Geometry-Aware Topology Reasoning via Bézier Curves in Autonomous Driving

链接:https://arxiv.org/abs/2610.09118

作者:Mihai Bogdan Deaconu,Laura Dioşan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reasoning jointly detects, jointly detects, structural connectivity, Topology reasoning jointly, traffic elements

备注: Accepted at NeurIPS 2026 (main track). 15 pages, 4 figures

点击查看摘要

Abstract:Topology reasoning jointly detects 3D lanes and traffic elements from multi-view images and infers their structural connectivity. Current methods model lanes as discrete polylines, lacking smoothness, analytical tangent directions, and global spatial support for attention, while providing sparse topology supervision. We propose TopoCurve, a geometry-driven architecture for 3D topology reasoning grounded in a structured parametric lane representation. Lanes are modeled as endpoint-fixed cubic Bézier curves, enabling continuous geometry with exact endpoints and analytically defined directionality. We exploit this shared curve geometry across the entire pipeline. Endpoint distance and tangent alignment are encoded with multi-scale Fourier features and injected into the topology head. Sampled curve points serve as geometry-aligned references for deformable cross-attention spanning the full lane. Parallel curve-anchored attention branches provide diverse predictions for one-to-many topology supervision. These components form a tightly coupled cascade where representation enables geometric reasoning, guides feature aggregation, and supports denser supervision. TopoCurve achieves 50.6 OLS on the OpenLane-V2 benchmark without any post-processing, establishing a new state-of-the-art among end-to-end camera-only methods, and outperforms all existing approaches on endpoint detection (56.8 vs. 52.6 on DET_p).

159. 【2610.09116】SPLATIFY: Reproduce, Discover, Innovate! From Papers and Ideas to Trainable 3DGS Code

链接:https://arxiv.org/abs/2610.09116

作者:Seemandhar Jain,Keshav Gupta,Manmohan Chandraker

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, research demands significant, demands significant effort, research demands, rapid growth

备注:

点击查看摘要

Abstract:The rapid growth of 3D Gaussian Splatting (3DGS) research demands significant effort to reimplement papers before building on them. We introduce SPLATIFY, a multi-agent framework that converts 3DGS papers into trainable gsplat-based implementations, where generic paper-to-code methods and frontier models fail. SPLATIFY achieves this through five innovations: (1) A context-free grammar for gsplat over a modular method template with extension points for losses, densification, rendering, and optimization, constraining synthesis so generated code satisfies gsplat's architectural invariants by construction. (2) Architectural elements for faithful reproduction: fork-aware citation recovery retrieving component-level code at function-level granularity, Graph-of-Thought synthesis in topological dependency order, RAG-guided in-context example selection from over 20 verified implementations, and visual feedback combining PSNR-guided regeneration, Gaussian-level structural checks, and VLM-driven patching. (3) Knowledge-driven compositional improvement that autonomously finds weaknesses and composes complementary regularizers, losses, and densification strategies to improve upon original results. (4) Interdisciplinary method discovery where agents retrieve physical priors from outside the 3DGS literature and compose them with rendering knowledge to produce methods for previously unaddressed scene types. (5) SPLATIFY-Bench, an evaluation framework across 30 diverse 3DGS papers. On papers without public code, SPLATIFY matches expert implementations while reducing development time from weeks to minutes, and through compositional discovery further improves PSNR by up to 2.4 dB. We additionally demonstrate novel methods for volumetric nebula rendering and other scientific domains, synthesized entirely by SPLATIFY.

160. 【2610.09085】Anaximander: Interactively Running Geospatial Deep Learning Models on Any Compute Backend

链接:https://arxiv.org/abs/2610.09085

作者:Satej S. Soman,Akram Zaytar,Girmaw A. Tadesse,Gilles Q. Hacheme,Muhammad S. Danish,Inbal Becker-Reshef,Rahul Dodhia,Juan Lavista Ferres

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Applying deep learning, Applying deep, geographic information systems, deep learning models, remote sensing practitioners

备注: Accepted as a poster at the TerraBytes II workshop, ECCV 2026

点击查看摘要

Abstract:Applying deep learning models to satellite imagery from within geographic information systems (GIS) remains high-friction for remote sensing practitioners. Models arrive in incompatible formats and target different compute environments, from local workstations to serverless cloud services. As a result, every evaluation demands custom deployment, tiling, and georeferencing code before a single prediction reaches the analyst's map. This friction discourages systematic comparison in a domain where model choice directly affects operational outcomes such as field delineation, crop monitoring, and disaster response. We present Anaximander, an open-source system that unifies model source and compute location choice behind one interactive interface. The system's backend is an inference server that loads models from multiple commonly-used sources and serves them on any accessible compute backend. The server provides session management and model caching, and streams results back per tile. The backend is paired with a QGIS plugin that drives tiling, result reassembly, georeferencing, and real-time per-tile status visualization. An additional user-interface path injects layer legends as prompts into vision-language models. We demonstrate the system in a code-free side-by-side comparison of three heterogeneous models on an agricultural field delineation task: gpt-image-1 via a cloud API, Segment Anything Model 3 (SAM3) on a remote GPU, and DelineateAnything on a local CPU. The inference backend and protocol are open-source and available at this https URL.

161. 【2610.09079】Large-scale Repository Engineering via Agent-Native Reusable Code Primitives

链接:https://arxiv.org/abs/2610.09079

作者:Haibo Jin,Peng Kuang,Xucheng Yu,Jerry Wang,Dehao Wu,Haohan Wang

类目:oftware Engineering (cs.SE); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large language models, language models equipped, building complete repositories, complete repositories remains, repositories remains difficult

备注: 44 pages

点击查看摘要

Abstract:Large language models equipped with development environments have moved code generation toward repository-scale construction, yet building complete repositories remains difficult because interacting modules, interfaces, configurations, tests, and dependencies must work together. We introduce Code Primitives, agent-native reusable executable components with interface contracts, dependency closures, validation tests, and provenance. Each primitive uses a resident LLM to assess relevance and adapt its implementation, interfaces, and dependencies to the target repository, and we organize 1,424 validated primitives in CodeFace, a searchable library for repository construction. We introduce LEGO (Large-scale repository Engineering via aGent-native reusable cOde primitives), which activates task-relevant primitives, integrates their adapted implementations with task-specific code while resolving cross-component constraints, and revises the result against executed tests. To measure construction end to end, we build LEGO-REPO, a benchmark of 522 executable reconstruction tasks spanning seven software domains, 22 capability tracks, and five difficulty levels, scored against native test suites between an empty-package floor and original-source ceiling. The strongest of 13 evaluated backbones reaches a delivery score of 0.318 and scores zero on 41.0% of tasks; LEGO improves all 13 by 0.1474 on average and raises GPT-5.6-terra from 0.3180 to 0.5134 (+61.4%). In controlled comparisons, adapted primitives outperform retrieved code supplied as context or vendored unchanged. The effect persists against independent repository agents, across three external benchmarks, and with a disjointly re-mined CodeFace; GPT-OSS-20B for adaptation and diagnosis retains 95.1% of the homogeneous score at 24.0% lower cost.

162. 【2610.09077】DISRQAD: Diffusion Image Super-Resolution Quality Assessment Dataset and Benchmark

链接:https://arxiv.org/abs/2610.09077

作者:Nikita Kukuzei,Artem Borisov,Evgeney Bogatyrev,Khaled Abud,Egor Chistov,Dmitriy Vatolin

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:create visually plausible, visually plausible detail, Diffusion-based image super-resolution, image super-resolution, create visually

备注:

点击查看摘要

Abstract:Diffusion-based image super-resolution (SR) can create visually plausible detail that is not supported by the low-resolution input. We introduce DISRQAD, a subjective-quality dataset and diagnostic benchmark for this setting. It contains mean opinion scores (MOS) for 14,000 SR outputs from ten diffusion and four non-diffusion methods, spanning four low-resolution degradation conditions and x2/x4 upscaling. We evaluate 51 standard full-reference and no-reference metric configurations and 11 adapted variants. Agreement with MOS is substantially weaker on diffusion outputs: the strongest standard no-reference baseline reaches 0.431 SRCC on diffusion SR versus 0.813 on non-diffusion SR. As a case study in benchmark use, a pruned and distilled Q-ReAlign-mini student reaches 0.496 SRCC on diffusion SR. DISRQAD measures perceived output quality, not faithfulness to the input; it enables analysis of metric behavior across generator families and input conditions. Our findings reveal a substantial gap in the assessment of diffusion-based SR and provide a basis for developing quality models sensitive to diffusion-specific artifacts.

163. 【2610.09071】OverLay++: Dense-Overlap Layout-to-Image Generation Dataset

链接:https://arxiv.org/abs/2610.09071

作者:Shivansh Aggarwal,Shresth Grover,Divyansh Srivastava,Haiyang Xu,Bingnan Li,Xiang Zhang,Ethan J. Armand,Chuan Li,Jianwen Xie,Zhuowen Tu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:made substantial progress, object-level control, made substantial, substantial progress, progress in spatial

备注: Accepted at NeurIPS 2026, Evaluations Datasets Track. Project website: [this https URL](https://mlpc-ucsd.github.io/OverLayPP) . Dataset: [this https URL](https://huggingface.co/datasets/mlpcucsd/OverLayPP)

点击查看摘要

Abstract:Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.

164. 【2610.09068】GUARD: Geometric Uncertainty-Aware Point Cloud Denoising and Segmentation for Robotic Hard Disk Drive Disassembly

链接:https://arxiv.org/abs/2610.09068

作者:Zuoxu Wang,Xiao Liang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:disassembly requires part-level, distinguish genuine component, requires part-level representations, robotic disassembly requires, genuine component geometry

备注:

点击查看摘要

Abstract:Reliable robotic disassembly requires part-level representations that distinguish genuine component geometry from scanning and reconstruction artifacts. In point clouds of hard disk drives (HDDs), structured ghost artifacts can resemble valid components locally while remaining inconsistent with the overall geometry, allowing erroneous measurements to receive plausible semantic labels. This creates an engineering information problem: semantic prediction confidence alone does not establish whether the underlying geometry is reliable. We propose \textbf{GUARD}, a geometric uncertainty-aware framework that performs point filtering and segmentation within a single forward pass by modeling the reliability of learned geometric representations. GUARD combines a multi-scale geometric transformer with a multi-bandwidth random Fourier feature Gaussian Process to estimate per-point geometric uncertainty, complemented by predictive entropy to suppress unreliable measurements while preserving informative structures. Evaluation on 2,745 real HDD point clouds shows that GUARD improves PointNet++ segmentation mean intersection over union from 0.7739 to 0.8318. Additional experiments on ShapeNetPart and ScanNet examine robustness across corruption types, point-cloud domains, and segmentation backbones. On manually annotated ScanNet samples, geometric uncertainty achieves a corrupted-point detection F1 score of 0.7931, compared with 0.2212 for predictive entropy. The results demonstrate the value of distinguishing geometric reliability from semantic confidence and reveal a tradeoff between artifact suppression and preservation of informative structures. GUARD contributes a reliability-aware approach to interpreting imperfect 3D measurements for component identification and subsequent robotic handling. Project website: this https URL.

165. 【2610.09032】Shape-Bayes: Bayesian Inference of Structured Shapes under Visual Ambiguity

链接:https://arxiv.org/abs/2610.09032

作者:Mani Kumar Tellamekala,Tosh Brown,Michel Valstar

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Perceiving structured shapes, Perceiving structured, inherently ambiguous task, real-world conditions, Perceiving

备注:

点击查看摘要

Abstract:Perceiving structured shapes, such as human faces, from pixels is an inherently ambiguous task in real-world conditions. Yet, shape inference is largely posed as a deterministic regression task predicting fixed spatial coordinates. We find that deterministic regression is brittle when visual evidence is ambiguous or incomplete; under severe occlusions deterministic models exhibit structural collapse, predicting incoherent shapes or reverting to generic averages. To address this, we introduce Shape-Bayes, a probabilistic framework that couples uncertainty-aware visual perception with Bayesian shape reasoning. Rather than forcing point estimates, Shape-Bayes dynamically weights visual evidence against geometric priors to infer a structurally valid shape posterior. Demonstrated on human face shape regression, a rigorous testbed featuring complex non-rigid deformations and strict anatomical constraints, Shape-Bayes comprises: (1) a base model predicting noisy landmarks alongside distilled aleatoric uncertainties; (2) a lightweight Transformer encoding these observations into an adaptive prior over a PCA shape manifold; and (3) a differentiable Bayesian solver computing closed-form posteriors by balancing the noisy predictions against this prior. By guaranteeing complete structural integrity, Shape-Bayes achieves an absolute improvement of up to ~34% IDR over state-of-the-art deterministic models. Simultaneously, it yields highly calibrated uncertainty bounds and reduces relative error by up to 12.5%, establishing a new state-of-the-art for robust 2D face shape regression under severe occlusion. The project page is at this https URL.

166. 【2610.09031】Beyond Explanation: Debugging Medical Imaging Models via Concept Intervention

链接:https://arxiv.org/abs/2610.09031

作者:Samrajya Thapa,Daniel J. Quest,Timothy L. Kline,Carrie L. Langstraat,Emanuel C. Trabuco,Wei Le

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Medical imaging models, Medical imaging, black boxes, limiting interpretability, systematic debugging

备注: Accepted at the 5th Workshop on Applications of Medical AI (AMAI), MICCAI 2026

点击查看摘要

Abstract:Medical imaging models often operate as black boxes, limiting interpretability and systematic debugging. We introduce an easy-to-use, plug-and-play framework for concept-based interpretation and model refinement. By aligning a single-modality encoder to BioMedCLIP, we construct a Concept Bottleneck Model (CBM) that enables concept-level interventions. These interventions allow us to isolate causal versus spuriously correlated concepts, validate insights with domain experts, and generate counterfactual samples for targeted fine-tuning. We evaluate our framework on a Mayo Clinic ultrasound dataset and the CheXpert 5x200 chest X-ray dataset. Results demonstrate that concept intervention enables reliable model diagnosis while maintaining, and occasionally improving predictive performance via guided fine-tuning. Our findings highlight the practical value of this framework for controlled, interpretable refinement of clinical deep learning models.

167. 【2610.09027】Visual Memory Attacks Can Persist Through The KV Cache

链接:https://arxiv.org/abs/2610.09027

作者:David Dobre,Leo Schwinn,Gauthier Gidel,Spandana Gella,Perouz Taslakian,Pierre-André Noël

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Modern language model, systems operate autonomously, Visual Memory Injection, Modern language, increasingly long contexts

备注:

点击查看摘要

Abstract:Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tokens, allowing adversarial influence to outlive direct access to its this http URL consider the Visual Memory Injection (VMI; Schlarmann and Hein, 2026) attack setting, in which an adversarial image that stays in the context plants a hidden backdoor: the model behaves normally until a chosen trigger elicits an attacker-chosen response. We first demonstrate persistence in this setting with optimized soft prompts, which remain effective after we mask the prompt from attention. We then introduce Persistent Visual Memory Injection (P-VMI), which optimizes images to preserve this adversarial behaviour after they are masked from attention. These attacks persist over conversations substantially longer than those used during optimization. On Qwen3-VL-8B-Instruct, P-VMI achieves up to approximately $90\%$ target success in its strongest configuration and remains effective under a stricter removal setting that exposes the image only on the first turn. A cache-swap ablation localizes the persistent influence to the KV cache. Finally, we show that these attacks can be trained to survive compaction that retains the KV cache of a summary generated by the same model, demonstrating that adversarial behaviour can persist in cached state without continued access to its source.

168. 【2610.09015】Personalize at Test Time: Learning User Preferences for Image Generation

链接:https://arxiv.org/abs/2610.09015

作者:Jiamu Bai,Jiaming Hu,Yanhong Wu,Zellux Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:preferences remains challenging, generate high-quality images, user preferences remains, remains challenging, user preferences

备注:

点击查看摘要

Abstract:Diffusion models can generate high-quality images, yet aligning their outputs with individual user preferences remains challenging. A key bottleneck is accurately modeling diverse user preferences from limited feedback. Existing approaches often rely on labor-intensive manual preference annotations or vision-language models (VLM) to extract preference information from user interaction histories, introducing substantial annotation or computational costs that limit scalability. We propose an approach that learns personalized reward models directly from users' historical image preference pairs. First, we use an autoencoder to compress hundreds of visual attributes into 50 attribute-anchored preference dimensions and train an evaluator to score images along these dimensions. We then represent each user's preferences as a linear combination of the shared dimension scores, estimating the user-specific weights by maximizing the likelihood of their observed pairwise preferences under the Bradley-Terry model. This formulation reduces per-user adaptation to optimizing a low-dimensional weight vector, simplifying optimization and enabling data-efficient personalization from sparse feedback. The learned personalized rewards guide image generation at inference time while keeping the diffusion model frozen. Experiments on real-user preference data show that our approach achieves approximately 77% held-out pairwise preference prediction accuracy and improves the alignment of generated images with individual user preferences.

169. 【2610.08983】Zero-Shot Brain MRI Inpainting with 2.5D Unconditional Flow Priors

链接:https://arxiv.org/abs/2610.08983

作者:Arnela Hadzic,Franz Thaler,Simon Johannes Joham,Martin Urschler

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:downstream brain analysis, brain analysis applications, automated downstream brain, brain MRI volumes, Generative inpainting

备注:

点击查看摘要

Abstract:Generative inpainting of brain MRI volumes is essential for synthesizing healthy tissue in pathological regions, improving the accuracy and reliability of automated downstream brain analysis applications such as image registration, brain extraction, and segmentation. However, standard 3D approaches are computationally prohibitive, while efficient 2D slice-wise methods suffer from severe inter-slice discontinuities. Furthermore, traditional models rely on conditional training, requiring task-specific learning of masked inputs. We propose a zero-shot brain MRI inpainting framework utilizing 2.5D unconditional flow priors to capture spatial context along the superior-inferior axis without the overhead of full 3D convolutions. During training, our flow matching model learns the joint distribution of adjacent axial slice triplets, modeling the manifold of healthy brain anatomy while explicitly excluding pathological regions from the loss function. At inference, the model processes the input triplets autoregressively along the depth axis. We employ the Restora-Flow solver to constrain the unconditional prior using the input mask, achieving accurate zero-shot inpainting. Evaluations show our 2.5D strategy resolves the structural discontinuities of 2D baselines, synthesizing plausible healthy tissue while maintaining volumetric consistency across the axial, sagittal, and coronal planes. As a final step, we generate and average an ensemble of multiple stochastic reconstructions to form the final prediction. Quantitative results benchmarked on the official BraTS 2026 Inpainting Challenge validation set demonstrate the effectiveness of our proposed approach, yielding an SSIM of 0.816 $\pm$ 0.112, MSE of 0.007 $\pm$ 0.005, and PSNR of 22.923 $\pm$ 4.343. Code is available at this https URL.

170. 【2610.08978】S2Tok: Streaming 3D Gaussian Reconstruction with Persistent Spatial Tokens

链接:https://arxiv.org/abs/2610.08978

作者:Fang Li,Jiraphon Yenphraphai,Quentin Herau,Depu Meng,Yihan Hu,Tianshuo Xu,Narendra Ahuja,Wei Zhan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:geometric predictions, sequence of geometric, incorporate new evidence, evidence and remain, remain renderable

备注: Project Page: [this https URL](https://s2tok.github.io/)

点击查看摘要

Abstract:Streaming 3D reconstruction requires more than a sequence of geometric predictions: it requires a persistent scene state that can incorporate new evidence and remain renderable as observations arrive. Latent spatial tokens offer a promising representation for this purpose, but constructing them from an image collection leaves open how to maintain them online, where each observation may both revisit known regions and reveal new content. We introduce S2Tok, a feed-forward framework that maintains a size-adaptive, persistent scene state from uncalibrated image streams. Its central idea is to distinguish updates to the existing representation from selective expansion. A spatially informed transformer integrates each incoming observation with the persistent scene tokens, while a learned admission module selectively expands the representation to limit redundant storage. A hierarchical decoder and Gaussian head convert the evolving state into non-pixel-aligned 3D Gaussians, enabling novel-view rendering without caching previous frames. Experiments across four benchmarks demonstrate competitive streaming rendering quality with compact Gaussian representations. These results support latent spatial tokens as a persistent computational state for online 3D reconstruction, combining learned scene updates with explicit Gaussian rendering.

171. 【2610.08954】RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

链接:https://arxiv.org/abs/2610.08954

作者:Yiyang Huang,Yitian Zhang,Yizhou Wang,Jianglin Lu,Qihua Dong,Hailing Wang,Huimin Zeng,Mingyuan Zhang,Yun Fu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:diverse video-language tasks, Query Comprehension Gap, Video large language, frame selection, Comprehension Gap

备注:

点击查看摘要

Abstract:Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates dominant approaches to long-video frame selection from a task-decomposition perspective, identifying two key challenges: the Query Comprehension Gap in similarity-based methods and the Interpretation--Selection Gap in judgment-based methods. To address them, we propose RACER, a training-free reflective agentic framework that decomposes long-video frame selection into query interpretation driven by a lightweight Vid-LLM and evidence localization supported by an embedding model serving as a retrieval tool. Specifically, the Vid-LLM is responsible solely for reformulating the complex query into sub-queries that make implicit information requirements explicit, mitigating the Query Comprehension Gap. Meanwhile, the retrieval tool leverages these sub-queries to localize relevant evidence, relieving the Vid-LLM of direct frame selection and thus addressing the Interpretation--Selection Gap. Finally, the retrieved frames are fed back to the Vid-LLM for sub-query refinement, forming a reflection loop that iteratively improves query interpretation and frame selection. Experiments across multiple benchmarks show that RACER consistently improves long video understanding. Notably, RACER achieves effective frame selection even with limited-capability components, demonstrating that agentic integration enables these components to enhance more capable Vid-LLMs.

172. 【2610.08944】One-Slide Calibration of Pathology Foundation Models

链接:https://arxiv.org/abs/2610.08944

作者:Ming Ren Hou,Tianyi Huang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:foundation models represent, Scanner variation, models represent, foundation model fixed, Abstract

备注: Accepted at the NeurIPS 2026 Workshop AI at Scale for Clinical Impact (ASCI): Cancer Pathology Foundation Models

点击查看摘要

Abstract:Scanner variation changes how pathology foundation models represent the same tissue. We introduce SlideRuler, which uses regions within a slide as internal controls to estimate and correct acquisition-induced shifts in other regions. A transfer map learned from paired rescans enables calibration from a single scan at inference while keeping the foundation model fixed. Across two encoders and five SCORPION scanners, learned transfer reduces mean target-to-source embedding distance by 16.3-38.5% relative to raw embeddings. Comparisons with unrelated same-scanner controls reveal a positive same-slide contribution across all four evaluation settings, including scanner holdout. A source-anchored variant reduces source-feature displacement by 47.7-83.6% relative to learned transfer while retaining most of its alignment gain. By drawing calibration information from the slide itself, SlideRuler offers a path toward more consistent use of frozen pathology models across imaging systems.

173. 【2610.08941】SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation

链接:https://arxiv.org/abs/2610.08941

作者:Yunheng Liu,Ziqi Cai,Siqi Yang,Yimu Wang,Minggui Teng,Jiaming Tan,Shuchen Weng,Erwin Wu,Kaipeng Zhang,Boxin Shi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:virtual reality experiences, embodied agent training, Language-guided panoramic video, video generation benefits, Language-guided panoramic

备注: Project page: [this https URL](https://alaya-lab.github.io/SPW-Nav)

点击查看摘要

Abstract:Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views. We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion. Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions. Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.

174. 【2610.08936】VCR-Bench: A Modular Open-Source Benchmark for Video Classification Robustness

链接:https://arxiv.org/abs/2610.08936

作者:Maksim Plinskiy,Aleksandr Gushchin,Sergey Lavrushkin,Dmitriy S. Vatolin,Anastasia Antsiferova

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:counterparts are absent, video, adversarial attacks, video classification, video classification temporal

备注: 6 pages,1 figure, accepted at ACM MM 2026

点击查看摘要

Abstract:Robustness of image classification has several benchmarks, but their video counterparts are absent. In video classification temporal dimension introduces additional degrees of freedom for adversarial attacks, defenses, and preprocessing. Temporal sampling, perturbation budgets, and metric aggregation also interact in ways with no direct analogue in the image setting. Therefore, robustness for video classifiers is studied across scattered, incompatible implementations, making reported numbers hard to reproduce and analyze. We introduce VCR-Bench, a modular open-source benchmark framework that standardizes video loading, wrappers for classifiers, adversarial attacks and defenses, perceptual metrics, configuration presets, and result logging. VCR-Bench currently integrates 30 video classification models, 14 adversarial attacks, and 10 defense wrappers under a common evaluation protocol. We evaluate representative video classifiers, attacks, and defenses on Kinetics-400 subset, reporting clean accuracy, attack success rate, perceptual quality, runtime, and memory usage. VCR-Bench is released with documented installation, reproducible run presets, component-extension interfaces, and scripts for reproducing the reported results at this https URL.

175. 【2610.08862】CIRRA: Dual-Level Continual Instruction Reconciliation with Ongoing Execution for Embodied Robot Agents in Interactive Household Tasks

链接:https://arxiv.org/abs/2610.08862

作者:Ci Zhang,Enfu Nan,Arman Akbari,Lin Zhao,Li Wang,Chen Wang,Weiwei Chen,Yanzhi Wang,Geng Yuan

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:continual instruction reconciliation, instruction reconciliation, Household Instruction Reconciliation, executing ongoing tasks, Continual Household Instruction

备注:

点击查看摘要

Abstract:Household robots must accommodate new user instructions while executing ongoing tasks. Existing agents often regenerate or extensively revise the remaining task sequence, introducing plan ambiguity, logical inconsistency, and redundant execution. We formulate continual instruction reconciliation and propose CIRRA (Continual Instruction Reconciliation for Robot Agents), a dual-level framework combining LLM-based semantic reasoning with rule-constrained structural integration. CIRRA first grounds incoming instructions to unique executable skills and resolves underspecified actions and execution locations. It then preserves the ongoing subtask sequence as an execution backbone and generates integration candidates by inserting incoming subtasks into location-matched segments. The semantic reasoner evaluates only modified segments to identify dependencies and conflicts and select the most logically coherent local integration. This structure-preserving process maintains alignment with ongoing execution, mitigates ambiguity and inconsistency, and reuses shared subtasks to reduce redundant execution. We also introduce CHIRP (Continual Household Instruction Reconciliation and Planning), a text-based benchmark of 120 episodes across eight household environments and six categories of everyday activities. On CHIRP, CIRRA achieves 74.2% decision agreement, exceeding the strongest replanning baseline by 30 percentage points; every correct fusion decision yields a correctly placed, conflict-free schedule. On a Unitree G1 humanoid, CIRRA interrupts ongoing skills at the correct moment in every trial and significantly outperforms all baselines on every metric.

176. 【2610.08830】MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models

链接:https://arxiv.org/abs/2610.08830

作者:Pengcheng Zheng,Chaoning Zhang,Jiaxin Yan,Sihan Cao,Jianwei Zhang,Xudong Wang,Jiaquan Zhang,Jewon Lee,Tae-Ho Kim,Yang Yang,Heng Tao Shen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:demonstrated remarkable reasoning, remarkable reasoning capabilities, Large Language Models, Multimodal Large Language, Large Language

备注: 13 pages

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.

177. 【2610.08826】PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking

链接:https://arxiv.org/abs/2610.08826

作者:Qinfeng Zhu,Weiguang Zhao,Yunxi Jiang,Anh Nguyen,Lei Fan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Full-sphere panoramic cameras, planar bounding box, Full-sphere panoramic, mobile robots track, robots track people

备注:

点击查看摘要

Abstract:Full-sphere panoramic cameras let fixed monitoring systems and mobile robots track people in every direction, but a planar bounding box does not fully describe where a person is on the sphere. We introduce PanoPed, a sim-to-real benchmark for pedestrian tracking on the full sphere. PanoPed-S contains 108,000 frames from fixed, quadruped-mounted, and drone-mounted cameras, with synchronized masks, depth, camera poses, and 3D pedestrian states. PanoPed-R adds 28,002 real frames from fixed cameras, 16,247 of them densely annotated. We find that an ERP rectangle cannot uniquely determine the spherical center and angular extent of the visible person, while the detector's visual query still carries information about them. Inspired by the sextant's use of angular measurements to locate objects, we propose Sextant, a plug-and-play angular localization head with only about 0.035M parameters. It reuses a frozen detector, keeps track identities unchanged, and needs no extra image encoder. Sextant gives the best result in our PanoPed-S test comparison, raising the strongest baseline, MOTIP, from 47.30 to 49.49 HOTA, with gains on all eight test sequences. Without fine-tuning on real data, the same synthetic-trained heads improve MOTIP and HAT by 0.96-1.14 HOTA on real video, and both seeds improve every real sequence. HAT+Sextant scores best among the compared systems that add no localization image encoder.

178. 【2610.08825】Autonomous Driving Research Requires a Community-Driven Data Paradigm

链接:https://arxiv.org/abs/2610.08825

作者:Jinsu Yoo,Zanming Huang,Katie Z Luo,Zheda Mai,Qiyuan Wu,Bharath Hariharan,Mark Campbell,Wei-Lun Chao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:made remarkable progress, reshaping urban mobility, advances enabling commercial, enabling commercial deployments, remarkable progress

备注: NeurIPS 2026 Position Paper

点击查看摘要

Abstract:Autonomous driving has made remarkable progress, with recent AI advances enabling commercial deployments that are reshaping urban mobility. Yet the field remains far from its universal social promise: autonomous systems that can operate robustly anywhere, anytime, for anyone. We posit that this gap is not merely a modeling problem, but a problem of the prevailing data paradigm. Current research relies heavily on a few benchmark datasets with limited spatial and scenario coverage, even though the community has collectively produced over 600 autonomous driving datasets across nearly 50 countries. However, this abundance has not translated into broad research impact: most datasets remain significantly underused due to fragmentation, limited visibility, incompatible protocols, and benchmark incentives that concentrate attention on a few dominant datasets. We therefore argue that autonomous driving research requires a collaborative, community-driven data paradigm. Such a paradigm would improve the discovery, reuse, integration, and evaluation of diverse datasets; make underexplored data easier and more rewarding to study; and lower the barrier for new contributors. We outline its key principles, illustrate an early realization, and call for collaboration across academia and industry to transform fragmented datasets into shared community infrastructure for anytime-anywhere autonomy.

179. 【2610.08823】HCPN-GCN: Scaling Hierarchical Prototype Networks with Cone Geometry for Continual Graph Learning

链接:https://arxiv.org/abs/2610.08823

作者:Sammuel R. Silva,Vander L. S. Freitas,Gladston Moreira,Eduardo J. S. Luz,Rodrigo Silva

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)

关键词:preserving knowledge acquired, Continual Graph Learning, Graph Convolutional Networks, aims to incrementally, preserving knowledge

备注:

点击查看摘要

Abstract:Continual Graph Learning (CGL) aims to incrementally learn from graph-structured data while preserving knowledge acquired from previous tasks. A major challenge in this setting is catastrophic forgetting, where learning new tasks degrades performance on previously learned ones. Hierarchical Prototype Networks (HPNs) address this problem through a prototype-based memory mechanism that avoids storing historical data, but their reliance on linear feature extractors limits their ability to exploit graph topology, while point-based prototypes often lead to inefficient prototype growth on structurally diverse graphs. In this work, we propose HCPN-GCN, a graph-aware extension of HPN that replaces the original linear feature extractors with Graph Convolutional Networks (GCNs) and introduces cone-based prototypes with a diversity regularization objective. The proposed design produces richer graph-aware representations while compactly modeling the embedding space, reducing prototype proliferation without sacrificing discriminability. Experimental results on six continual graph learning benchmarks demonstrate that HCPN-GCN consistently improves average classification accuracy over the original HPN and representative continual learning baselines while maintaining near-zero forgetting. Furthermore, our analysis shows that the proposed model learns substantially richer class-level prototype hierarchies using approximately $30\times$ fewer atomic prototypes than the original HPN, providing a more compact and effective memory representation for continual graph learning.

180. 【2610.08813】Pre-training, Reasoning, Benchmarking: X-ray Report Generation on CheXpert Plus Dataset

链接:https://arxiv.org/abs/2610.08813

作者:Xiao Wang,Yuxiang Zhang,Dan Xu,Yuehang Li,Shiao Wang,Bo Jiang,Yaowei Wang,Yonghong Tian,Jin Tang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:medical artificial intelligence, patient waiting periods, critical research direction, alleviate clinicians' diagnostic, clinicians' diagnostic workload

备注:

点击查看摘要

Abstract:X-ray image-based Radiology Report Generation (RRG) constitutes a critical research direction within medical artificial intelligence, with great potential to alleviate clinicians' diagnostic workload and shorten patient waiting periods. Despite substantial advances over recent years, the field faces evident bottlenecks stemming from insufficient standardized benchmarks and inadequate domain adaptation of generic large models. Notably, the newly released CheXpert Plus dataset is provided without accompanying baseline implementations and evaluation results, which impedes standardized training, quantitative evaluation and fair comparison among follow-up algorithms. To mitigate this limitation, we establish a comprehensive benchmark encompassing prevailing X-ray report generation models and Large Language Models on CheXpert Plus. This benchmark delivers a reliable comparative foundation for upcoming methods and enables researchers to rapidly identify state-of-the-art approaches within this domain. Beyond benchmark construction, we rethink X-ray RRG under the paradigm of large models and propose a novel framework termed MambaXray-PRB. Our framework improves report generation performance and enhances model interpretability via multi-stage large-model pre-training and multi-modal Chain-of-Thought reasoning. The pipeline consists of three successive phases: self-supervised auto-regressive modeling, X-ray-report contrastive learning, and post-training optimization for reasoning and report generation. Extensive experiments on IU X-ray, MIMIC-CXR, and CheXpert Plus datasets validate the effectiveness of MambaXray-PRB for radiology report generation. The source code of this paper is available on this https URL

181. 【2610.08800】A Review Of Robotic World Models For Dynamic Environments Based On Factor And Scene Graphs

链接:https://arxiv.org/abs/2610.08800

作者:Marco Giberna,Miguel Fernandez-Cortizas,Jose Luis Sanchez Lopez,Holger Voos

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:successful robotic solutions, powerful foundation, foundation for internal, related literature, scene graphs

备注: 34 pages, 7 tables, 8 figures

点击查看摘要

Abstract:Models based on graphs have emerged in robotics as a powerful foundation for internal world representations, where factor and scene graphs are among the most prominent model types found in the related literature and in successful robotic solutions. Initially, many of these models were assuming static environments as a simplification. Herein, factor graphs mainly provide uncertainty-aware geometric estimations while scene graphs enable a structured semantic abstraction. However, real-world robotic environments are often dynamic, posing severe challenges for purely static world representations. Therefore, this review presents a comprehensive view on how dynamic aspects of real-world environments can be addressed in such graph-based world models. We organize our assessments around three main aspects: (I) suitable representations, (II) pipelines to construct and update the representations, and (III) their exploitation for downstream tasks. We review approaches that are either based on factor or scene graphs, but put special emphasis on novel approaches that combine both types to form hybrid models. We mainly analyze how different types of dynamics can be modeled herein, and categorize common architectural patterns. Finally, emerging trends and open challenges are identified, including uncertainty propagation from learned perception through the representation layers, the observability of dynamic-entity motion and scale under minimal sensing, scalable lifelong maintenance, and the lack of datasets and evaluation protocols that ground world-model quality in downstream task performance under dynamics.

182. 【2603.18166】Efficient Dense Crowd Trajectory Prediction Via Dynamic Clustering

链接:https://arxiv.org/abs/2603.18166

作者:Antonius Bima Murti Wijaya,Paul Henderson,Marwa Mahmoud

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:trajectory prediction plays, safety and management, prediction plays, plays a crucial, crucial role

备注:

点击查看摘要

Abstract:Crowd trajectory prediction plays a crucial role in public safety and management, where it can help prevent disasters such as stampedes. Recent works address the problem by predicting individual trajectories and considering surrounding objects based on manually annotated data. However, these approaches tend to overlook dense crowd scenarios, where the challenges of automation become more pronounced due to the massiveness, noisiness, and inaccuracy of the tracking outputs, resulting in high computational costs. To address these challenges, we propose and extensively evaluate a novel cluster-based approach that groups individuals based on similar attributes over time, enabling faster execution through accurate group summarisation. Our plug-and-play method can be combined with existing trajectory predictors by using our output centroid in place of their pedestrian input. We evaluate our proposed method on several challenging dense crowd scenes. We demonstrated that our approach leads to faster processing and lower memory usage when compared with state-of-the-art methods, while maintaining the accuracy

183. 【2610.09288】FaceKit: a Toolkit for Interpretable Facial Phenotyping, Synthetic Image Generation and Privacy Analysis in Rare Diseases

链接:https://arxiv.org/abs/2610.09288

作者:Hongzhuo Chen,Zhanliang Wang,Florent Pollet,Mian Umair Ahsan,Joshua Bie,Tzung-Chien Hsieh,Peter Krawitz,Cong Liu,Wendy K Chung,Chunhua Weng,Gamze Gürsoy,Kai Wang

类目:Quantitative Methods (q-bio.QM); Computer Vision and Pattern Recognition (cs.CV)

关键词:recognizable craniofacial features, Human Phenotype Ontology, recognizable craniofacial, facial, Phenotype Ontology

备注:

点击查看摘要

Abstract:Many rare genetic diseases are associated with recognizable craniofacial features. However, traditional approaches for describing facial morphology rely largely on qualitative clinical observation and free-text descriptions, which are often subjective, non-standardized, and difficult to reproduce across observers and institutions. Although the Human Phenotype Ontology (HPO) provides controlled terms for describing facial features, these terms are typically categorical rather than quantitative and may vary depending on examiner experience and interpretation. Here, we present FaceKit, a computational framework for quantitative facial phenotyping from frontal facial photographs. FaceKit extracts standardized measurements of facial landmarks and derived 120 morphological features, then reports feature-level z-scores representing deviation from population reference distributions. The reference distributions are built from the FairFace dataset spanning diverse ancestral groups. We evaluated FaceKit on a curated subset of the GestaltMatcher Database covering 50 rare-disease cohorts. In addition to quantitative facial analysis, FaceKit includes synthetic facial image generation to support rare disease model development and data augmentation. We also performed privacy evaluation to assess whether synthetic images reveal identifiable information from real patient photographs and could compromise patient privacy. Across disease case studies, FaceKit-derived quantitative measurements captured known facial features associated with rare genetic disorders and provided objective support for clinical phenotyping. Together, these results establish FaceKit as a useful tool for quantitative phenotyping, and has the potential to improve rare disease diagnosis, support genotype-phenotype studies, and enable more reproducible clinical characterization across diverse patient populations.

184. 【2610.08849】Abdominal Ultrasound Simulation from Semantic Labels using Paired Label-to-Physics-Based Image Translation

链接:https://arxiv.org/abs/2610.08849

作者:Santiago Vitale,Duilio Deangeli,Ignacio Larrabide,José Ignacio Orlando

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Current abdominal ultrasound, Current abdominal, abdominal ultrasound, pathology variability, Stage

备注:

点击查看摘要

Abstract:Purpose: Current abdominal ultrasound (US) simulation methods often require CT-based anatomical references for ray-casting, limiting deformation and pathology variability. We propose a learning-based pipeline trained to predict physics-based images derived from CT scans from semantic labels, enabling controlled simulations without patient-specific CT volumes at inference time. Methods: We introduce a two-stage pipeline that maps anatomical segmentations to realistic US images through a simplified US image. Stage~I synthesizes this image from semantic labels using models trained on CT-based ray-casting outputs. Stage~II refines it into a realistic US scan using anatomically guided unpaired translation. Deformations and pathologies are generated by editing anatomical maps. Results: We evaluated Pix2Pix and the Semantic Diffusion Model (SDM) in Stage~I, followed by segmentation-guided CycleGAN (SG-CycleGAN) refinement in Stage~II. SDM significantly outperformed Pix2Pix in morphological metrics, including MAE (19.85 vs. 21.65), SSIM (0.28 vs. 0.24), and mIoU (0.43 vs. 0.29), whereas Pix2Pix yielded better perceptual point estimates (LPIPS: 0.17 vs. 0.19; FID: 0.32 vs. 0.37; KID: 0.25 vs. 0.48). Conclusion: Training paired generative models with physics-based supervision enables approximation of CT-derived ray-casting outputs at inference time directly from semantic labels. Although the pipeline does not require patient-specific CT volumes at inference time, CT-derived segmentations and ray-casting simulations remain necessary to train Stage~I. Once trained, the framework enables controllable healthy and pathological simulations through semantic-map modification.

Subjects:

Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2610.08849 [eess.IV]

(or
arXiv:2610.08849v1 [eess.IV] for this version)

https://doi.org/10.48550/arXiv.2610.08849

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Duilio Deangeli [view email] [v1]
Fri, 2 Oct 2026 16:26:46 UTC (19,611 KB)

185. 【2610.08848】STRIDE: Spatial-Temporal Representation for Interval-conditioned Disease Evolution in Longitudinal Glioblastoma MRI

链接:https://arxiv.org/abs/2610.08848

作者:Wenhao Guo,Changchang Yin,Pierre Giglio,Weidan Cao,Ping Zhang,Golrokh Mirzaei

类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:primary brain tumor, aggressive primary brain, brain tumor, aggressive primary, primary brain

备注: 37 pages, 11 figures

点击查看摘要

Abstract:Glioblastoma (GBM), an aggressive primary brain tumor, is routinely monitored with longitudinal MRI after treatment. Distinguishing stable disease (SD), pseudoprogression (PsP), and true progression (TP) remains challenging because these states can show overlapping MRI appearances despite different temporal trajectories. Existing longitudinal methods still face challenges in modeling scan-specific spatial variability, variable follow-up intervals, and complementary information from the observed follow-up state and its longitudinal change. We propose STRIDE, a framework for spatial-temporal representation of interval-conditioned disease evolution that takes paired post-treatment MRI scans and their inter-scan interval as input and predicts SD, PsP, or TP. The lesion-prior-guided spatial representation combines SoftGate and an Adaptive-window Hierarchical Transformer (AWHT) to emphasize lesion-related regions while preserving surrounding context. The time-conditioned latent transition uses pair-level context and the actual inter-scan interval to estimate interval-dependent representation changes between visits. The observed--transition fusion integrates the transition-estimated follow-up representation with the directly observed follow-up representation to jointly characterize the follow-up state and its longitudinal change. BraTS2024 is used to develop and evaluate the lesion-prior generator, while longitudinal pretraining on LUMIERE supports transfer before downstream adaptation to Burdenko. On the Burdenko three-class task, STRIDE achieves a macro ROC--AUC of 0.816 and a macro F1-score of 0.796. These results support its potential for more reliable longitudinal post-treatment GBM state assessment.

186. 【2610.08838】Deep Learning for Longitudinal Medical Imaging: A Scoping Review

链接:https://arxiv.org/abs/2610.08838

作者:Francesca Mussa,Divyanshu Tak,Atlas H. Avval,Sarah Brueningk,Ray H. Mak,Hugo J. W. L. Aerts,Andreas M Rauschecker,Benjamin H. Kann

类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:modern medical practice, patient care, cornerstone of modern, practice and patient, Longitudinal medical imaging

备注:

点击查看摘要

Abstract:Longitudinal medical imaging analysis is a cornerstone of modern medical practice and patient care. Deep learning applied to longitudinal imaging offers wide potential to enhance diagnosis and track disease progression by capturing spatial changes over time. With major advances in single-timepoint deep learning for imaging, there has been growing interest in longitudinal image analysis, given its increased clinical relevance, though technical challenges remain. Several recent innovations may lead to a new era of multi-timepoint image evaluation, yet the scientific landscape, recent progress, and areas of need remain under-characterized. To address this gap, we conducted a scoping review of deep learning methodologies applied to longitudinal medical imaging, yielding 102 studies published between 2018 and 2025. Neurological disorders (48%) and ophthalmic conditions (12%) were the most common clinical applications, with MRI serving as the predominant imaging modality (67%). Sequential feature modeling approaches combining convolutional neural networks (CNNs) with temporal models (LSTM/RNN) were the most frequent methodology (40%), followed by direct feature aggregation across timepoints (23%). Most studies targeted classification tasks (56%), while external validation was performed in only 24% of studies. Our findings highlight that deep learning-based longitudinal imaging analysis remains a promising field, though newer temporal architectures and larger datasets may improve success and clinical adoption of these tools.