本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新826篇论文,其中:

  • 自然语言处理104
  • 信息检索22
  • 计算机视觉125

自然语言处理

1. 【2609.20822】Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

链接https://arxiv.org/abs/2609.20822

作者:Bingxin Xu,Yuzhang Shang,Zhen Dong,Emilio Ferrara

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:http URL, language model writes, evaluate coding agent, URL this paradigm, promising paradigm

备注

点击查看摘要

Abstract:Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific this http URL this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch. The agent pursues the goal but collides with the obstacle in most cases, treating task completion as its sole objective while neglecting safety. The agent reasons about the obstacle in its traces, and the prompt already forbids touching it, so neither perception nor instruction is at fault; the fault lies in the planning, where the stated constraint never becomes a priority. By decomposing manipulation into a route phase and a contact-rich moment, we locate the source of the failure. Along the route, the model cannot prioritize the safety constraint, having no notion of a clearing route and none of replanning once a chosen route becomes infeasible. At the contact, it is unaware that contact execution is bounded by the same constraint. To close this gap, we present SafeHarness, which equips the model with two obstacle-aware harnesses that enable it to prioritize the safety constraint. Obstacle-aware route planning grounds the objects as bounding boxes and draws candidate routes over them as sequences of waypoints. The agent then plans a route in advance, verifies it, replans when necessary, and only then executes it. Obstacle-aware contact execution instead selects the contact position so that the contact itself avoids the obstacle. SafeHarness attains 71.9% task success and 87.5% collision avoidance, surpassing the previous SOTA by 6.5% and 27.0%, respectively. These results are $2.3\times$ and $1.5\times$ those of the same agent without harnesses.

2. 【2609.20821】Embedding Models Measure in Peculiar Ways

链接https://arxiv.org/abs/2609.20821

作者:Juri Opitz,Andrianos Michail

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:spaces define notions, Embedding spaces define, distance, define notions, spaces define

备注

点击查看摘要

Abstract:Embedding spaces define notions of semantic similarity and distance. We study whether those embeddings reflect physical measurements of mass, distance, time and volume, which admit a unique, objective notion of semantic equivalence and distance. We find that physical measurement is only weakly modeled in the embedding space, and that instead quite peculiar measurement patterns can be observed. Further analysis indicates that embedding representations of physical measurements are strongly influenced by superficial string similarity, and recalibration of similarity does not substantially improve the alignment.

3. 【2609.20808】Unifying Models of Intergroup Hostility in Online Discourse

链接https://arxiv.org/abs/2609.20808

作者:Patrick Gerard,Julia Mendelsohn,Kristina Lerman

类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:Hostile rhetoric, moderate hostile rhetoric, justify mistreatment, groups can normalize, normalize exclusion

备注: 16 pages

点击查看摘要

Abstract:Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. However, these theories were developed largely in parallel, often propose different and sometimes conflicting accounts of how hostility develops, and have rarely been tested against each other in real discourse. The result is a fragmented understanding of the rhetorical mechanisms of hostility, without a clear sense of how they appear, and relate to each other, in real-world discourse. Using 2.86 million posts from TikTok, Truth Social, and Twitter/X during the 2024 U.S. presidential election, we model the mechanisms of six foundational theories of intergroup hostility -- boundary construction, threat construction, scapegoating, negative evaluation, dehumanization, and action orientation -- within a common empirical framework to recover the broader organization of intergroup hostility rhetoric. Structurally, we find that boundary construction and threat construction anchor the system; temporally, we find that these mechanisms tend to follow a regular ordering: boundary construction, derogation, and action orientation tend to appear early; dehumanization and threat construction later; scapegoating latest. Mapping how these theoretical frameworks actually manifest in discourse bridges longstanding divisions across social science traditions and presents computational social science with a clearer empirical foundation for modeling intergroup hostility rhetoric beyond single-label detection.

4. 【2609.20804】An Empirical Study of Harness Design for Coding Agents

链接https://arxiv.org/abs/2609.20804

作者:Run-Ze Fan,Zihao Zhang,Simin Ma,Yebowen Hu,Shouju Wang,Kaiqiang Song,Fei Liu,Hamed Zamani,Xiaoyang Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Software Engineering (cs.SE)

关键词:Coding harnesses shape, existing work typically, individual components unclear, typically evaluates harnesses, work typically evaluates

备注: 43 pages

点击查看摘要

Abstract:Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 176 matched settings spanning five context-management strategies, four context-window budgets, and targeted ablations of planning and action space. We find that: (1) Context management becomes increasingly valuable as the context-window budget tightens, with most of its benefit coming from preventing context-overflow failures. (2) Staging rule-based elision before LLM-based summarization provides the strongest overall efficiency among the context-management strategies, whereas making elided content recoverable adds machinery that models rarely use and yields no accuracy gain. (3) Planning shifts from an accuracy scaffold for weaker models to a cost saver for stronger models, with little change in accuracy. (4) Predefined tools improve performance for models with weaker bash proficiency, whereas bash-capable models can operate effectively with a bash-only interface and achieve substantially lower cost, especially on command-line-centric tasks. Trajectory-level analysis explains these effects: context management extends execution trajectories without substantially altering agent behavior, planning changes where trajectories stop, and the action space changes the granularity at which code is written. These findings inform model- and budget-aware harness design and provide a modular framework for evaluating future harness components.

5. 【2609.20800】JEPA-Anything: Learning Predictive Models across Different Worlds

链接https://arxiv.org/abs/2609.20800

作者:Taoyong Cui,Zhongyao Wang,Xinyue Xu,Weiyang Liu,Zhaochen Yu,Yuying Zhang,Qiang Gao,Mengyue Yang,Wanli Ouyang,Pheng Ann Heng,Yingcheng Wu,Zhenfei Yin,Ling Yang

类目:Computation and Language (cs.CL)

关键词:modeling enables intelligence, World modeling enables, World modeling, anticipate consequences, enables intelligence

备注: Code: [this https URL](https://github.com/Gen-Verse/JEPA-Anything)

点击查看摘要

Abstract:World modeling enables intelligence to anticipate consequences, guide interventions, and learn from interaction. Yet predictive models remain domain-specific: can a common learning principle support world modeling across radically different systems? We introduce JEPA-Anything, a domain-agnostic framework based on orthogonal predictive factorization (OPF). Extending joint-embedding predictive architectures, OPF decomposes latent targets into complementary factors, learns them through dedicated pathways, and recombines them within a shared predictive design. We evaluate JEPA-Anything across seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather. Experiments span representation learning, intervention prediction, out-of-distribution generalization, and long-horizon dynamics, including 10 matched dynamics tasks, forecasting of over 1,000 clinical events, and 100-step molecular rollouts across four systems. Against matched JEPA baselines, JEPA-Anything improves reported metrics on all 10 dynamics tasks and reduces single-intervention prediction error on Interventional Pong by 34.8%. It achieves the lowest one-step and 100-step molecular errors among compared methods in all four systems. Beyond prediction, a factor-nominated biological intervention receives experimental support in cell co-cultures, patient-derived organoids, tumor fragments, and mice; latent orbital modes recover the Keplerian scaling exponent with a fitted slope of -1.4991. These results support a common factorized predictive principle across heterogeneous worlds, connecting world modeling with intervention and experimentally grounded scientific discovery. Code: this https URL

6. 【2609.20784】RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

链接https://arxiv.org/abs/2609.20784

作者:Yan Yu,Zhengxi Lu,Yizhou Liu,Yichen Pan,Aozhe Wang,Qipeng Chen,Hua Yang,Wenqi Zhang,Weiming Lu,Qianglong Chen,Yongliang Shen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Multi-turn agents trained, supply dense token-level, dense token-level supervision, Multi-turn agents, privileged task skills

备注

点击查看摘要

Abstract:Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.

7. 【2609.20779】Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

链接https://arxiv.org/abs/2609.20779

作者:Sarah Wyer,Sue Black,Noura Al Moubayed

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Safety evaluations, report declining harm, language models rely, large language models, evaluations for large

备注: Accepted at EMNLP 26 Main Conference

点击查看摘要

Abstract:Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($\rho = +0.55$, $p = .034$) while Detoxify does not ($\rho = -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.

8. 【2609.20751】dQwen3.5: Hybrid-Attention Diffusion Language Models

链接https://arxiv.org/abs/2609.20751

作者:Anton Xue,Litu Rout,Aditya Akella,Adam Klivans,Sujay Sanghavi,Sanjay Shakkottai

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:diffusion language model, pretrained autoregressive, language model, cost-efficient route, diffusion language

备注

点击查看摘要

Abstract:Adapting a pretrained autoregressive (AR) model is a cost-efficient route to a diffusion language model (DLM). While nearly all such adaptations start from a full-attention transformer, AR modeling has shifted toward hybrid architectures that interleave attention and RNN layers. This creates an obstacle for adaptation: unlike attention, RNNs are structurally causal and nontrivial to bidirectionalize. Despite this mismatch, we investigate whether such backbones can become effective DLMs by adapting Qwen3.5 at 0.8B, 2B, 4B, and 9B scales, yielding the dQwen3.5 family. We find that hybrid backbones can be efficient starting points for adaptation: against a full-attention control, the hybrid reaches a given training loss in about half the tokens. Across scales, dQwen3.5 resembles full-attention DLMs in any-order decoding behavior and performs strongly under parallel decoding.

9. 【2609.20734】On-Demand Attention: Language Models Know When to Recall

链接https://arxiv.org/abs/2609.20734

作者:Haibo Feng,Ruiqi Liang,Hanyang Peng,Shiqi Yu

类目:Computation and Language (cs.CL)

关键词:agentic workloads increasingly, workloads increasingly demand, increasingly demand efficient, Reasoning and agentic, demand efficient long-context

备注: 28 pages, 5 figures

点击查看摘要

Abstract:Reasoning and agentic workloads increasingly demand efficient long-context inference. Yet full-attention decoding reads the growing history at every step, regardless of its benefit to the next prediction. We show that a pretrained model's decoding states already contain information predictive of this benefit, before the global read. Building on this finding, we introduce On-Demand Attention (ODA), a local-first decoding method that uses a lightweight recall head to selectively invoke global attention as its predicted benefit changes during generation. ODA trains only the recall head, leaving pretrained weights unchanged and the complete historical KV cache available for future recall. We further implement GPU-side conditional execution in vLLM, translating reduced global reads into practical decoding speedups over full attention at long context lengths. Experiments across Qwen and Gemma models, including hybrid-attention backbones, show that selective recall recovers most of the performance lost under local attention while substantially reducing global reads. These findings support long-context inference in which pretrained models guide their own access to the information they retain.

10. 【2609.20715】Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

链接https://arxiv.org/abs/2609.20715

作者:Juzheng Zhang,Disha Makhija,Manoj Ghuhan Arivazhagan,Vinayshekhar Bannihatti Kumar,Rashmi Gangadharaiah

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Agent trajectories record, trajectories record, pass, SFT, Agent trajectories

备注: 29 pages, 9 figures, 11 tables

点击查看摘要

Abstract:Agent trajectories record what an agent does and what happens next. Yet standard supervised fine-tuning (SFT) applies loss only to agent-authored action tokens, using environment observations as context but not as prediction targets. We ask whether this convention provides the best initialization for subsequent reinforcement learning. We introduce ActObs, which also supervises the observation tokens already present in each trajectory. Although deployed agents never generate observations, learning to predict them encourages the policy to model action consequences without adding data, parameters, sequence tokens, or forward passes. The methods perform similarly after SFT but diverge after GRPO. On Qwen3-4B, GRPO from ActObs achieves higher pass@k at every evaluated sampling budget than its action-only counterpart on Terminal-Bench 2.0. On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks. The advantage extends to cross-domain code editing on aider-polyglot (+4.2 pp at pass@1 at 4B), whose tasks are unseen during SFT and RL. ActObs retains more entropy during RL while requiring less policy movement, leaving the final policy closer to its SFT initialization. Our analysis traces this difference to SFT: action and observation gradients rapidly become orthogonal, while action-only training leaves a large residual observation gradient and degrades environment prediction below the base model. Joint supervision prevents this one-sided specialization, preserving consequence prediction and preparing the policy for downstream exploration.

11. 【2609.20712】Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol

链接https://arxiv.org/abs/2609.20712

作者:Levent Bulut

类目:Computation and Language (cs.CL)

关键词:abstract summary label, proposed systematic tendency, reconstructable inferential structure, represent narrative meaning, abstract summary

备注: v1.1. 8 pages. Also archived at Zenodo: [this https URL](https://doi.org/10.5281/zenodo.22817289)

点击查看摘要

Abstract:This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable inferential structure that produces it. Within the Bulut Doctrine, narrative effect is theorized along a told-shown axis: in told mode, emotional and informational content is declared explicitly and requires little reader reconstruction; in shown mode, that content is suppressed at the surface and must be reconstructed from physical cues and indirection (Objective Projection). Shown mode is the higher-load condition the doctrine is designed to measure. The claim is that LLMs fail along this axis in a specific direction. Summarization bias is hypothesized to operate in two regimes: (i) a generative regime, in which a model asked to render an emotion through Objective Projection defaults to declaring it instead; and (ii) an evaluative regime, in which a model judging narrative quality rewards told-mode explicitness and under-detects shown-mode suppression. The evaluative regime is the more consequential, since LLMs increasingly serve as judges and reward models, and a directional bias toward told mode would impose a selection pressure degrading prose toward flat declaration. This report does not claim the bias is validated. It defines the construct, situates it against LLM-as-judge biases, rereads a completed independent reliability study as directional evidence consistent with it, and pre-registers a two-regime test with decision rules under which the construct would be abandoned.

Comments:
v1.1. 8 pages. Also archived at Zenodo: this https URL

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.20712 [cs.CL]

(or
arXiv:2609.20712v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.20712

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
12. 【2609.20684】HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

链接https://arxiv.org/abs/2609.20684

作者:Hassan Saeed Hassan Albattra,Mazen Mohammed Bahgat,Rahatara Ferdousi,Hana Essam Sayed Ahmed Amrya,Mariam Mousa

类目:Computation and Language (cs.CL)

关键词:Large language models, emphasize response quality, evaluations emphasize response, Large language, interpreted correctly

备注: 8 pages, 2 figures, 3 tables. Submitted to the 2026 International Conference on Large Language Models (LLM 2026)

点击查看摘要

Abstract:Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.

13. 【2609.20634】PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations

链接https://arxiv.org/abs/2609.20634

作者:Julian Eggert(Honda Research Institute Europe, Offenbach, Germany)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:exact interval boundaries, qualitative calculus, Allen interval algebra, relations, Allen

备注: 41 pages, 7 figures. Open-source implementation at [this https URL](https://github.com/HRI-EU/probabilistic-allen-algebra)

点击查看摘要

Abstract:Allen's interval algebra is a qualitative calculus for temporal relations, but its thirteen base relations are crisp predicates over exact interval boundaries. This is inadequate for temporal information from language, perception, databases, or uncertain histories, where times, durations, and boundaries are uncertain and expressions such as "just before" or "roughly during" have graded meaning. We develop the probabilistic Allen algebra (PAA): a generative and complete extension in which relation probabilities are derived from distributions over interval boundaries rather than assigned as scores. Time points are Gaussian; intervals have Gaussian midpoints and truncated-Gaussian durations. Every relation is a boundary-ordering predicate in one common probability space: point-point relations reduce to error functions, and point-interval and interval-interval relations to multivariate Gaussian orthant probabilities induced by linear inequalities. Contact relations (meets, starts, finishes, equals) receive positive measure through a tolerance band, and under a single tolerance the thirteen relations form a true partition that recovers crisp Allen as the tolerance vanishes. The construction derives Allen's taxonomy rather than positing it: coarse predicates such as precedence, overlap, and containment are unions of leaves whose probabilities are leaf sums, and this hierarchy is preserved as intervals collapse to points and thirteen relations reduce to five and then three. Each relation further decomposes into correlation-aware temporal primitives in the spirit of CIDOC CRM. The algebra is scale-invariant and separates graded expressions such as "shortly before" from contact relations. All results are Monte-Carlo validated and shipped as an open, tested Python package.

14. 【2609.20630】UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising

链接https://arxiv.org/abs/2609.20630

作者:Kun Yao,Yuhang Zhou,Yichi Zhang,Zeliang Tong,Shengri Xue,Haitao Wang,Siyu Lu,Qianlong Xie,Xingxing Wang

类目:Computation and Language (cs.CL)

关键词:connects user intent, Search advertising connects, advertising connects user, platform monetization, search advertising system

备注: 13 pages, 5 figures, 4 tables

点击查看摘要

Abstract:Search advertising connects user intent with commercial content and plays a critical role in platform monetization. Recent systems typically align pretrained generative models with a single business reward, such as eCPM, or use naive reward fusion for preliminary multi-objective alignment. However, an ideal search advertising system must jointly account for heterogeneous objectives, including relevance, click propensity, and commercial value, to balance user experience and business value while mitigating globally suboptimal performance caused by gradient competition. We propose UniPolicy, an objective-aware multi-policy alignment framework. UniPolicy combines objective-specific prefix tokens, sparse MoE-LoRA routing, and objective-specific residual FFNs to hierarchically decouple parameters within a shared backbone, providing differentiated parameter and policy-expression spaces for different business objectives. It further constructs pairwise preferences from multi-stage behavioral feedback, supplementing the relative preference information in exposed-but-unclicked samples and strengthening the relative advantage of clicked candidates in the generation distribution. At inference, UniPolicy supports parallel, business-customizable multi-policy beam search, flexibly allocating candidate quotas across objectives under a fixed retrieval budget. Large-scale offline experiments show that UniPolicy delivers balanced improvements across multiple metrics while preserving retrieval quality, outperforming single-objective reinforcement learning and naive reward-fusion baselines. In a 7-day online A/B test on a real search advertising system, UniPolicy improves CTR by 0.71%, RPS by 1.58%, and advertising revenue by 1.32%, while maintaining stable serving latency.

15. 【2609.20625】Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

链接https://arxiv.org/abs/2609.20625

作者:Tisha Chawla,Susheem Koul

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language model, read changing state, re-run rarely repeats, Large language, language model responses

备注

点击查看摘要

Abstract:Large language model responses are non-deterministic, so failures in LLM agents are hard to reproduce: a failure depends on inference that is not bitwise reproducible, on tools that read changing state, and on a multi-step trajectory that a re-run rarely repeats. Record-and-replay makes a run reproducible, but existing agent tooling records runs only to trace or score them, not to test a code change against them. We present Chronicle, which records an agent run at its non-deterministic boundaries as immutable envelopes and replays it from the record. Its central operation, cut-point replay, serves a chosen subset of boundaries from the record and executes the complementary subset live with new code, turning a recorded incident into a regression test that runs in continuous integration. On a benchmark of 6 recorded failures with simulated model boundaries, recording adds 23 {\mu}s per crossing (0.008% of an assumed 300 ms model call), full replay issues zero model calls and is bit-stable across 20 repetitions, and cut-point tests fail on faulty code and pass on guarded and benign changes for all 6 incidents. In a mutation study of the guarded tools, cut-point tests catch every mutant that lets the recorded unsafe action through, while a baseline that stubs every boundary, using the same assertion, catches none. Chronicle and the benchmark are publicly available at this https URL.

16. 【2609.20612】What Does Privileged Information Add to On-Policy Self-Distillation?

链接https://arxiv.org/abs/2609.20612

作者:XiuYu Zhang,Wei Chow,Junfeng Fang,Zhenkai Liang,Tat-Seng Chua

类目:Computation and Language (cs.CL)

关键词:On-policy self-distillation, language model learn, language model, frozen copy, model learn

备注

点击查看摘要

Abstract:On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or a worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B's improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Evidence for an additional reference benefit is modest in Qwen, strongest for a polished solution, whereas complete traces add two percentage points in SmolLM3-3B at step 50. These benefits depend on the student being trained. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled rollouts turns gains into losses in both families while the problems, references, and evaluation stay fixed. Teacher profiles and matched loss interventions in Qwen further show that changing token-level supervision can leave student behavior largely unchanged. Together, these findings suggest that OPSD can improve access to existing reasoning capabilities through parameters shared by direct-response and thinking-enabled inference. The value of a privileged reference is what it adds to this cross-mode transfer, not how much of the solution it reveals.

17. 【2609.20593】WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution

链接https://arxiv.org/abs/2609.20593

作者:Yi Zhou,Kiamehr Rezaee,Danushka Bollegala,Mohammad Taher Pilehvar,Jose Camacho-Collados

类目:Computation and Language (cs.CL)

关键词:remains challenging, lexical-semantic tasks, challenging for language, recent progress, progress on lexical-semantic

备注: Accepted to AACL 2026 (main)

点击查看摘要

Abstract:Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also from the absence of an explicit sense inventory that specifies the relevant level of semantic granularity. We evaluate open LLMs on WiC and traditional Word Sense Disambiguation (WSD) under similar settings. We find that providing candidate senses, similar to what is done in traditional WSD, improves WiC performance in all settings. In general, explicit sense information helps models make more consistent and targeted judgements. Human evaluation further shows that many apparent WiC errors reflect label ambiguity or mismatches between model and annotator sense boundaries rather than simple failures of lexical understanding. In particular, results show that LLMs overthink the sense distinction often leading to errors based on overly fine-grained distinctions.

18. 【2609.20584】SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment

链接https://arxiv.org/abs/2609.20584

作者:Chenxi Wu,Zimu Wang,Haiyang Zhang,Wei Wang,Zhijie Xu

类目:Computation and Language (cs.CL)

关键词:Large language models, regulated functional-safety workflows, Large language, workflows remains underexplored, Risk Assessment

备注: Accepted at EMNLP 2026 Industry Track

点击查看摘要

Abstract:Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial benchmark for LLM-assisted automotive Hazard Analysis and Risk Assessment (HARA) under ISO 26262. It contains 3,000 de-identified industrial HARA cases and evaluates two coupled tasks: open-ended hazard analysis and standards-grounded risk assessment. To evaluate open-ended HARA artifacts, we propose the first reference-anchored LLM-as-a-judge protocol with high expert correlation. Experiments with nine frontier LLMs show that models often produce plausible hazard narratives but remain weak at ISO 26262 risk classification, with the best ASIL macro-F1 reaching only 0.261. Chain-of-Thought prompting provides limited benefit and often degrades categorical risk assessment. Error analysis further localizes major failures to scenario-critical context omissions during hazard generation and to controllability misjudgments during risk assessment, indicating where expert oversight should be concentrated. The dataset can be obtained from this https URL.

19. 【2609.20565】Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies

链接https://arxiv.org/abs/2609.20565

作者:Zimu Wang,Yiwen Jiang,Xiangyu Zhao,Yaling Shen,Jiahe Liu,Stephanie Fong,Maxmartwell H Cheng,Guilherme C Oliveira,Anh Nguyen,Robert Desimone,Barnaby Nelson,Dominic Dwyer,Zongyuan Ge

类目:Computation and Language (cs.CL)

关键词:Cognitive Behavioral Therapy, Behavioral Therapy, Cognitive Behavioral, large language models, Recent advancements

备注: Accepted at EMNLP 2026

点击查看摘要

Abstract:Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic decision-making informed by the client's real-time mental state, this aspect has often been overlooked in current research, limiting both flexibility and therapeutic outcomes. In this paper, we introduce StratCBT, a dataset specifically designed for psychological counseling conversations with CBT Strategies, consisting of 9,688 sessions and around 256K utterances, with each counselor's response aligned with one of eight distinct strategies. The creation of StratCBT involves modeling clients based on their negative thoughts and generating high-quality counseling conversations through self-chat, incorporating realistic sessions as guidance, thereby significantly surpassing existing datasets in both general counseling and CBT-specific skills. We conduct extensive experiments to demonstrate the effectiveness of strategy-aligned generation and evaluate its efficacy in delivering professional and effective counseling with LLM-simulated clients to reflect real-world scenarios. The dataset can be obtained from this https URL.

20. 【2609.20543】Language-model groups overstate consensus when replaying human deliberation on a reasoning task

链接https://arxiv.org/abs/2609.20543

作者:Tengfei Shao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Multiagent Systems (cs.MA)

关键词:Full-consensus rates, states are operationalized, treated as indicators, participation and final, final states

备注: 37 pages, 4 figures. Preregistration: [this https URL](https://osf.io/5jp7s) . Code and data: [this https URL](https://doi.org/10.5281/zenodo.21318346)

点击查看摘要

Abstract:Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.

21. 【2609.20541】An Analysis of Training-Free Self-Reported Confidence in Language Models

链接https://arxiv.org/abs/2609.20541

作者:Lukas Meyer,Sofia Rossi,Wei Chen,Thomas Laurent,Yiming Li

类目:Computation and Language (cs.CL)

关键词:Large language models, Large language, generated content, calibrated rhetoric, language models

备注: workshop

点击查看摘要

Abstract:Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4\% to 9\% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.

22. 【2609.20530】Relational Attention for Data-Efficient Language Modeling

链接https://arxiv.org/abs/2609.20530

作者:Adrian Brasoveanu,Ece Takmaz,Jakub Dotlačil

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:single decoder-only Transformer, cognitively motivated inductive, motivated inductive biases, Dual Attention Transformer, decoder-only Transformer

备注: BabyLM Workshop, EMNLP 2026. Source code: [this https URL](https://github.com/abrsvn/babylm_dat_2026)

点击查看摘要

Abstract:We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level ("sensory") lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM's data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT's three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPE-based, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100M-word) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard's NLP-task subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.

23. 【2609.20484】Edustories: A Collection of Real-world Case Studies from Classroom Practices

链接https://arxiv.org/abs/2609.20484

作者:Michal Štefánik,Jan Nehyba,Jirina Karasova,Martin Fico,Lucie Škarková,Markéta Košatková,David Kosatka

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:widely recognized potential, individualized student assistance, widely recognized, prior work, work has focused

备注

点击查看摘要

Abstract:Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce Edustories, a dataset of 1,492 teacher-written case studies describing real elementary and high-school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. Among many other applications, Edustories enables evaluating LLMs' ability to predict the success of teacher interventions, crucial for providing practicing teachers with useful feedback. Comparing the latest models from four language-model families against expert assessments, we find that current models fall short of human expertise in predicting classroom outcomes; the strongest models reach 58% accuracy compared to 64% of human experts. This gap highlights both the limitations and the emerging potential of AI as assistants for practicing teachers.

24. 【2609.20412】Stress-testing Alignment Midtraining

链接https://arxiv.org/abs/2609.20412

作者:Sid Baines,Jonathan Bostock,Maria Angelica Martinez,Andrew Draganov,David Africa,Daniel Tan

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:deployment environments, aligning frontier models, AMT, midtraining, post-training techniques

备注

点击查看摘要

Abstract:When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One proposed solution is alignment midtraining (AMT), which continues pretraining on large volumes of alignment-relevant documents to encourage generalisation in later stages of training. Despite the prominence of AMT as an alignment approach, there is limited public evidence for its effectiveness. To resolve this, we identify several assumptions around midtraining and evaluate them across scale: up to 110 billion-parameter models and 1 billion midtraining tokens. For instance, we study a scenario where post-training data is ambiguous between two possible motivations. We find that midtraining can steer the model's motivation in simple versions of this setting. However, the presence of a tiny fraction of finetuning data which suggests a competing motivation erases the effects of AMT. We also study scenarios in which we want an AI to follow a number of rules, but only demonstrate a subset of them. We find that demonstrations must be present either in midtraining or post-training datasets for these rules to be robustly learned. Based on these and other findings, we do not believe that there is sufficient public evidence for us to confidently state that midtraining can address the core difficulties inherent in aligning powerful AI systems.

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2609.20412 [cs.CL]

(or
arXiv:2609.20412v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.20412

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
25. 【2609.20408】Xeno-Interpretability: Investigating the Alien Minds of LLMs

链接https://arxiv.org/abs/2609.20408

作者:F. Pierucci,M. Bracale Syrnikov,M. Prandi,M. Galisai,F. Giarrusso,P. Bisconti

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, Large language, related categories, Large, language models

备注

点击查看摘要

Abstract:Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.

26. 【2609.20398】Schema-Anchored Latent Reasoning for Semantic Parsing-Based Knowledge Base Question Answering

链接https://arxiv.org/abs/2609.20398

作者:Guangze Gao,Zixuan Li,Sikui Zhang,Chunfeng Yuan,Wenjuan Li,Bing Li,Xiaolong Jin,Weiming Hu

类目:Computation and Language (cs.CL)

关键词:based knowledge base, executable logical forms, answer natural language, Semantic parsing, knowledge base question

备注

点击查看摘要

Abstract:Semantic parsing (SP)-based knowledge base question answering aims to answer natural language questions by generating executable logical forms (LFs) over knowledge bases (KBs). When applying Large Language Models (LLMs) to this task, a key challenge over large, heterogeneous KBs is selecting question-related schema elements (i.e., relations and classes) and composing them into complex LFs. Recent LLM-based methods often make early discrete commitments to schema elements during intermediate reasoning, allowing incorrect intermediate schema decisions to propagate and finally result in incorrect LFs. To overcome this limitation, we propose SALR, a schema-anchored latent reasoning method for LF construction. It performs multi-step reasoning by generating continuous thoughts in the model's hidden states, thereby delaying the explicit commitment to LF decisions. To ground this latent reasoning process in the corresponding KB schema, SALR aligns continuous thoughts with a codebook of KB schema elements through an alignment objective supervised by schema traces deterministically derived from gold LFs. It then incorporates the aligned schema codes into inputs for subsequent reasoning steps. This schema-mediated feedback guides LF generation without requiring the model to emit an explicit textual reasoning trajectory. Experiments on GrailQA and WebQSP show that SALR achieves consistent overall gains over strong baselines. Notably, on compositional questions from GrailQA, SALR outperforms TIARA, a strong SP-based baseline, by 2.86 F1 points. Further analyses show that schema-mediated feedback affects LF generation and that schema information is recoverable from the latent states.

27. 【2609.20303】Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

链接https://arxiv.org/abs/2609.20303

作者:Tamal Maharaj

类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)

关键词:culturally sensitive material, philosophical corpora pose, English Complete Works, Classical philosophical corpora, contemporary readers

备注

点击查看摘要

Abstract:Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must be verifiably grounded. We present Viveka-Insight, a bilingual resource and open-source pipeline for the works of Swami Vivekananda (1863-1902): the nine-volume English Complete Works and the ten-volume Bengali Vani o Rachana, two related but non-parallel corpora of about 15 million characters. Four layers are released: (i) a structure-preserving parse (32,694 paragraphs, 168,842 sentences) with per-paragraph anchors deep-linking into the published editions; (ii) a cross-lingual concept graph of 8,362 language-agnostic concepts with 87,518 relation-typed paragraph-concept and 55,872 concept-concept edges, in which canonical English labels act as a string-equality key linking Bengali and English passages with no parallel data; (iii) a bilingual alias inventory of 60,850 surface forms (30,053 English, 30,797 Bengali); and (iv) a human-annotated set of 200 paragraph-concept edges judged by three annotators, released with all per-annotator judgments. We report known-item cross-lingual retrieval over 194 verified rendered lecture pairs (Recall@10 0.86 in both directions), a 30-question audit of citation integrity and modern-question bridging, and a human study placing concept-extraction precision at 0.60 under strict two-annotator consensus (Cohen's kappa = 0.61). The extractor's confidence weight is calibrated: restricting to weight = 0.8 raises precision to 0.71 while retaining 98% of concept-bearing paragraphs. Precision is markedly lower in Bengali than English (0.54 vs 0.68), locating the weakness in exactly the half that cross-lingual access depends on. The design transfers to other multilingual classical corpora.

28. 【2609.20252】Lens: Bringing the Right Semantic Perspective into Focus for Training-Free Multimodal Representation Learning

链接https://arxiv.org/abs/2609.20252

作者:Xinran Liu,Shouqian Shi,Yixian Chen,Ruizhi Chen,Xin-Wei Yao,Sheng Zhong

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:large language models, High-quality representations, wide range, multimodal large language, language models

备注: 9 pages, 3 figures, 4 tables

点击查看摘要

Abstract:High-quality representations are essential for a wide range of downstream tasks. Dedicated embedding models are explicitly optimized for representation learning, yet their training data are often more limited in scale and diversity than the massive corpora used to pretrain modern large language models and multimodal large language models. Large-scale pretraining and instruction following enable autoregressive models to select relevant evidence, integrate multimodal information, and infer semantics under different task perspectives, creating a distinctive opportunity for training-free representation learning. However, our analysis reveals that existing semantic-elicitation methods do not reliably orient the extracted states toward the semantic perspective required by the downstream task. Consequently, the resulting representations often remain dominated by salient input content. We characterize this problem as semantic perspective misalignment and propose Lens, a training-free framework that makes representation readout task-directed. Semantic Perspective Anchoring associates the task-required perspective with a task-specific readout phrase, specifying the interpretive role of the positions later used for extraction. Contextualized Phrase Readout places the same phrase after the complete input and aggregates its token states, combining full-context access with the anchored perspective. The resulting representation reflects task-conditioned evidence integration and inference rather than a generic summary of salient content. Without parameter updates, architectural modification, or reranking, Lens achieves an overall Precision@1 of 63.9 across all 36 MMEB datasets, outperforming the closest same-backbone training-free embedding baseline by 10.2 points.

29. 【2609.20232】he Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation

链接https://arxiv.org/abs/2609.20232

作者:Bo Chen

类目:Computation and Language (cs.CL)

关键词:Public Discourse Corpus, Public Discourse, public-figure interview speech, interview speech jointly, Discourse Corpus

备注: 16 pages, 1 figure

点击查看摘要

Abstract:We introduce the \textbf{Public Discourse Corpus (PDC)}, the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality. The corpus contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words) after sentence segmentation and filtering. To ensure that all retained videos contain analyzable speech from the intended speaker, we introduce \textbf{Target Speaker Participation (TSP)}---a five-category annotation taxonomy with documented inter-annotator reliability ($\kappa = 0.616$)---as a key methodological contribution that any corpus construction project can adopt. Target-speaker turns are separated from interviewer and third-party speech through an \textbf{audio-first diarization pipeline} combining local Whisper ASR with pyannote speaker separation, released as an open-source implementation. We release the annotated corpus, the annotation tools, the cross-provider validation sample, and the complete processing pipeline. The dataset is available at this https URL.

30. 【2609.20223】Before the Warning Comes Too Late: Incremental Phone-Scam Detection from Speech

链接https://arxiv.org/abs/2609.20223

作者:Khang Nhat Hoang Vo,Anh Trac Duc Dinh,Tai Tien Ta,Tho Quan

类目:Computation and Language (cs.CL)

关键词:study weakly supervised, weakly supervised incremental, supervised incremental telecom, call ends, incremental telecom fraud

备注

点击查看摘要

Abstract:We study weakly supervised incremental telecom fraud detection from raw telephone audio, where training provides only conversation-level labels and predictions must be updated before a call ends. We introduce StreamFraudNet, which processes incoming audio through overlapping bounded-context windows using a frozen self-supervised speech encoder, recurrent temporal modeling, and learned aggregation of latent window scores. On a controlled English benchmark, StreamFraudNet achieves a ROC--AUC of \(0.9953\), significantly outperforming acoustic and mean-pooling baselines while remaining competitive with strong global temporal models. The model produces its first prediction after 10 seconds of audio, updates every 2 seconds, and operates faster than real time on the evaluated server hardware. Ablations identify recurrent temporal context as the principal contributor to performance. These results demonstrate that fraud risk can be scored incrementally from raw speech without transcripts or temporal annotations, while highlighting the need for latency-aware training to improve early prediction.

31. 【2609.20207】Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices

链接https://arxiv.org/abs/2609.20207

作者:Matthew F Dixon

类目:Computation and Language (cs.CL); Category Theory (math.CT); Probability (math.PR); Machine Learning (stat.ML)

关键词:scientific systems require, systems require uncertainty, Large language models, models produce prompt-dependent, produce prompt-dependent probabilities

备注

点击查看摘要

Abstract:Large language models produce prompt-dependent probabilities over words, whereas scientific systems require uncertainty over meaningful states that can be updated as evidence arrives. We develop an observable framework for determining when language-derived probabilities support such a sequential state representation. Theoretically, we define typed measurable transformations of contextual language, construct a minimal closed representation, and give necessary and sufficient conditions for semantic updates to exist uniquely. We bound irreducible nonclosure and accumulated error, and under average contraction prove existence, uniqueness and stability of an external random recursion on a probability simplex. These results define a stochastic lexical calculus without attributing an internal calculus to the language model. Empirically, frozen experiments test the observable implications. Raw prompt-conditioned probabilities fail the prespecified invariance gate; after prompt-specific calibration, a common three-state representation passes the stability gates and covers 28 of 30 untouched eight-step paths, or 0.933 at nominal level 0.90. Accordingly, language probabilities support a stochastic state only conditionally on verified closure, stability and coverage within a declared operating domain.

32. 【2609.20186】o Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

链接https://arxiv.org/abs/2609.20186

作者:Roy Eisenstadt,Ido Cohen,Edo Cohen-Karlik,Lior Wolf,Itamar Zimerman

类目:Computation and Language (cs.CL)

关键词:accelerated Large Language, Large Language Model, significantly accelerated Large, Large Language, accelerated Large

备注

点击查看摘要

Abstract:Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level n-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce SwitchSD, an adaptive framework that treats copying as a latent control signal of the LLM. By training lightweight probes on the target model's internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC 0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning copying from a noisy heuristic into a principled, model-aware decoding regime.

33. 【2609.20169】Fine-Tuning Models for Biomedical Relation Extraction

链接https://arxiv.org/abs/2609.20169

作者:Claudiu Creanga,Liviu P. Dinu,Daniela Gifu

类目:Computation and Language (cs.CL)

关键词:enabling large-scale investigations, Next-Generation Sequencing, Sequencing has revolutionized, genetic mutations, enabling large-scale

备注

点击查看摘要

Abstract:Next-Generation Sequencing has revolutionized the study of genetic mutations, enabling large-scale investigations into their roles in disease development. However, extracting meaningful insights from the vast amount of biomedical literature remains a complex challenge that cannot be addressed manually. In this paper, we present pre-trained models (PTMs) for the automatic extraction of relations from biomedical text, specifically targeting the variant-phenotype domain. Our evaluation on the SNPPhenA corpus demonstrates that fine-tuning small BERT-based models, particularly DeBERTa, yields strong performance, approaching the current state-of-the-art (SOTA). Additionally, our results indicate that carefully fine-tuning Google's Gemini Pro 1.0 outperforms the existing SOTA for both sentence-level tasks (where the model processes only the target sentence) and abstract-level tasks (where the model processes the entire abstract).

34. 【2609.20131】hink Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking

链接https://arxiv.org/abs/2609.20131

作者:Lijun Liu,Zhengzong Chen,Wenyan Li,Yuanyuan Zhao,Fei Huang

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, shown promising improvements, Reasoning-based reranking, Language Models

备注

点击查看摘要

Abstract:Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking. However, current methods predominantly rely on a single reasoning trajectory, resulting in rankings that are susceptible to reasoning errors and inherently constrained in modeling the multifaceted signals underlying document relevance. To resolve this dilemma, we propose MERIT-Rank(Multi-perspective Evidence and Reasoning Integration for Text Reranking), a framework that models complementary reasoning trajectories to improve reranking robustness. MERIT-Rank formulates a Multi-Trajectory Reasoning Space (MTRS) that evaluates query-document relevance from multiple perspectives and introduces a joint reranker that consolidates these reasoning paths into a unified ranking decision. We further develop Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives. Experiments on both reasoning-intensive and traditional retrieval benchmarks show that MERIT-Rank consistently achieves superior performance over competitive baselines. The 4B model notably outperforms most 7B and even 32B rerankers on BRIGHT.

35. 【2609.20104】Design of the IBM Granite 5.0 TurboCTC ASR Model

链接https://arxiv.org/abs/2609.20104

作者:Brian Kingsbury,George Saon,Masayuki Suzuki,Hong-Kwang J. Kuo,Takashi Fukuda,Samuel Thomas,Vishal Sunder,Avihu Dekel

类目:Computation and Language (cs.CL)

关键词:million parameter encoder-only, Turbo CTC, excellent speed-accuracy tradeoff, parameter encoder-only model, million parameter

备注: 5 pages, 2 figures, submitted to ICASSP 2027

点击查看摘要

Abstract:We describe the architecture, training methodology and inference speedups of Granite 5.0 Turbo CTC, a 470 million parameter encoder-only model with an excellent speed-accuracy tradeoff. The architecture uses pyramidal temporal subsampling within Conformer blocks using strided depthwise convolutions, block-diagonal (chunk-wise) self-attention, and conditioning on intermediate predictions from the middle layer. Training highlights are the use of only publicly available data, the novel use of a Muon optimizer, and balanced data sampling. Inference speedups include replacing 1 x 1 convolutions with linear layers and optimizing the attention computation in the Conformer blocks. Collectively, these result in a model that is on the speed-accuracy Pareto frontier of the Open ASR leaderboard for English short-form ASR while being twice as fast as the fastest competitor. The model can be used under a permissive license and downloaded from this https URL.

36. 【2609.20082】MATCH: Model-Aware Tool Learning with Curriculum Scheduling and Hierarchically Gated Rewards

链接https://arxiv.org/abs/2609.20082

作者:Shihao Liu,Hao Yin,Lijun Liu,Zhengzong Chen,Yuanyuan Zhao,Fei Huang

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:enables large language, learning enables large, Tool learning enables, large language models, parametric knowledge

备注

点击查看摘要

Abstract:Tool learning enables large language models (LLMs) to use external tools for tasks beyond parametric knowledge. Reinforcement learning can optimize tool-call behavior from feedback, but current methods still face two problems: fixed-threshold curricula can become misaligned with the policy's evolving capability boundary, and additive rewards can leak argument-level credit when the predicted tool is wrong. To address these problems, we propose MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards. Model-Aware Curriculum Learning (MACL) maintains reward-derived sample difficulty that co-evolves with the policy, and each epoch selects samples near the current capability boundary together with a top-k pool of harder cases. Hierarchical Tool-call Gated Reward (HTGR) scores tool name, argument key, and argument value as a gated chain, granting credit at each level only when prerequisites hold. The same HTGR rewards drive both GRPO updates and MACL's difficulty refresh, closing the loop between policy optimization and sample scheduling. On API-Bank and BFCL V3, MATCH reaches 72.19% and 62.87% overall accuracy, outperforming the main supervised and RL-based baselines. Backbone experiments further show consistent improvements across four backbones from two model families.

37. 【2609.20081】Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

链接https://arxiv.org/abs/2609.20081

作者:Hasindri Watawana,Sergio Burdisso,Esaú Villatoro-Tello,Manjunath K E,Kadri Hacioglu,Petr Motlicek,Andreas Stolcke

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)

关键词:favors frequent classes, shown strong potential, frequent classes, shown strong, strong potential

备注

点击查看摘要

Abstract:SpeechLLMs have shown strong potential for emotion recognition, yet they read the predicted emotion off a generative decoder not suited for classification: it can emit labels outside the target set and favors frequent classes. We propose a discriminative adaptation that reads the final prompt token's hidden state through a classification head, producing a label in one forward pass without modifying the backbone. Because this readout starts from the hidden state the model would otherwise decode, it gives a controlled comparison of generative and discriminative inference in an otherwise identical speechLLM. We keep the head a single linear layer, trading little accuracy for interpretability: each emotion becomes one direction in the LLM output token space, revealing associated tokens. On IEMOCAP, across two speechLLM architectures, it improves Macro F1 and removes hallucinations, with largest gains on realistic ASR transcripts. Our analysis reveals that these emotion directions encode indirect associations mirroring biases in web-scale text.

38. 【2609.20068】Marginal utility, matrix factorization, and the Key-Value (KV) cache: a unified information-economic framework for sovereign geo-mining inference

链接https://arxiv.org/abs/2609.20068

作者:Caroline Gans Combe(INSEEC)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:marginal utility schedule, marginal utility, transformer language models, diminishing marginal utility, utility schedule

备注: Version 11, 14 septembre 2026. 49 pages, 9 tables. Les valeurs de l'architecture souveraine sont projet{é}es et non mesur{é}es ; le calcul {à} grande {é}chelle est en cours. Soumission pr{é}vue {à} IEEE Transactions on Artificial Intelligence

点击查看摘要

Abstract:This paper builds a theoretical bridge between the economic notion of marginal utility and two machine-learning constructs, matrix factorization and the Key--Value cache of transformer language models. The singular value spectrum of a rating matrix is shown to be a diminishing marginal utility schedule for latent factors, the eigenvalue spectrum of the projected covariance operator to be the marginal utility schedule of a model's learned representation, and cache eviction and low-rank cache compression to be instances of constrained utility maximization under a memory budget. The three collapse into a single allocation rule: retain the top dimensions whose eigenvalue exceeds the shadow price of the binding constraint. The framework is applied to the automated extraction of structured information from geo-mining documents, where it motivates a multi-pass inference protocol, a layer-wise TIES model merging procedure, and a selection policy combining extraction quality, localization drift and energy, scalarized with a Conditional Value-at-Risk term on drift. Two empirical contributions are reported. An 11.2-million-parameter hierarchical classifier, trained in about five minutes on a single GPU, reaches 90.0 per cent level-1 accuracy on a held-out test set from a 973-document uranium-exploration corpus, against 92.0 per cent for a proprietary model on a fifty-document human audit of the same corpus, at a latency of 2.62 ms per card against approximately 2,000 ms for the API and at negligible cost. A diagnostic of uniform-density TIES merging exposes a reproducible degenerate mode in which the merged model returns token-identical outputs across five geographically distinct districts while declaring high confidence; re-executing the merge under layer-wise calibrated densities removes that signature on the diagnostic sample. The full-scale extraction benchmark, including LoRA fine-tuning, is reported as projected rather than measured and remains an empirical extension of this work.

39. 【2609.20059】AI Should Facilitate Democratic Deliberation at Scale

链接https://arxiv.org/abs/2609.20059

作者:José Ramón Enríquez,Jiaxin Pei,Alex Pentland

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:addressing cognitive, market-driven frictions, scale by addressing, preserving human agency, strengthen democracy

备注: 15 pages, 2 figures, ICML 2026

点击查看摘要

Abstract:AI systems can strengthen democracy by supporting deliberation at scale by addressing cognitive, social, platform-design, and market-driven frictions, while preserving human agency. Unlike proposals such as liquid democracy that restructure representation through vote delegation, in this position paper, we argue that AI-assisted deliberation offers a more promising path by lowering barriers to meaningful engagement without substituting machine judgment for human choice. Drawing on evidence from online deliberation platforms and experimental research, we identify four guiding principles: preserving agency and autonomy, encouraging mutual respect, promoting equality and inclusiveness, and augmenting rather than substituting active citizenship. We also address critical challenges, including alignment, sycophancy, training bias, and over-reliance on AI systems. We call on the machine learning community to develop deliberation-focused AI systems evaluated not on engagement metrics but on their capacity to facilitate informed, representative, and friction-robust discourse.

40. 【2609.20050】he Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

链接https://arxiv.org/abs/2609.20050

作者:Zhexi Feng,Ruiyi Zhang,Yongbo Yang,Pengtao Xie

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:retriever ranks highest, coding agent halfway, ranks highest, retriever ranks, coding agent

备注: 32 pages, 3 figures. Benchmark and evaluation resources: [this https URL](https://github.com/LordTARN1SHED/SERBench)

点击查看摘要

Abstract:A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination that supplies the support its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the current decision was annotated to require. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone reaches 66.6%, placing the gain in the set-level policy, not the computation. From frozen repository source with no gold-derived pool, the lead is 5.0 points. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark's own memory agent. Removing one required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.

41. 【2609.20005】Geopolitical Divisions Across Languages in Large Language Models

链接https://arxiv.org/abs/2609.20005

作者:Maxim Chupilkin

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:People increasingly turn, People increasingly, world events, explanations of world, People

备注

点击查看摘要

Abstract:People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.

42. 【2609.19989】Benchmarking LLM Compliance with China AI Generated Content Regulations

链接https://arxiv.org/abs/2609.19989

作者:Chenrui Cui,Hongye Fang,Lisha Song,Weichao Chen,Yue Zhu,Gang Xu

类目:Computation and Language (cs.CL); Multiagent Systems (cs.MA)

关键词:escalating content compliance, widespread adoption, led to escalating, content compliance risks, escalating content

备注: 5 pages, 3 figures, with appendix still improving

点击查看摘要

Abstract:The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China's current AI-Generated content compliance requirements and provides evaluation results on 20 notable LLMs, offering insight into China's regulatory landscape. We design a novel framework to assess the compliance and refusal rates with 2303 questions spanning six distinct dimensions, including 203 self-constructed constitutional questions. The framework employs several judges to generate verdicts independently based on their hierarchical alignment memory. Our findings show that international models also exhibit high levels of compliance despite the use of standard Chinese questions, and the main differences may stem from dimensions closely related to ideological alignment. We establish a regulatory benchmark that enables the global AI community to evaluate both Chinese and non-Chinese LLMs under a unified set of legally grounded compliance requirements.

43. 【2609.19969】DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

链接https://arxiv.org/abs/2609.19969

作者:DeepSeek-AI:Anyi Xu,B. Li,Bangcai Lin,Bing Xue,BingCheng Xian,Bingzheng Xu,Bochao Wu,Bowei Zhang,Boyi Deng,C.C. Yu,Chao Jin,Chaofan Lin,Chen Dong,Chenbing Wang,Chenfan Feng,Chengda Lu,Chenggang Zhao,Chengqi Deng,Chengyuan Zhang,Chenhao Xu,Chenqi Zhao,Chenze Shao,Chuhao Wang,Chuqi Zhang,Damai Dai,Dejian Yang,Deli Chen,Di Huang,Di Wu,Donghao Li,Erhang Li,Eric Fu,F. Zhou,Fangwei Zhou,Fangyun Lin,Fangzhou Yuan,Feiyu Xia,Fucong Dai,Guangbo Hao,Guanglin Li,Guanting Chen,Guoai Cao,Guofan Fan,Guolai Meng,Guowei Li,Haichuan Zhang,Haiyang Ma,Haiyang Shen,Han Li,Han Yu,Han Zhang,Hangyuan Deng,Hanwei Xu,Hanxiang Xu,Hanxun Zhong,Hao Guo,Hao Jiang,Hao Li,Hao Qin,Haodong Wen,Haofen Liang,Haofeng Huang,Haohua Liu,Haoling Zhang,Haoming Luo,Haoran Yang,Haotian Xu,Haotian Yuan,Haoting Huang,Haowen Luo,Haoyang Cai,Haoyu Chen,Haozhe Ji,Hengran Zhang,Hengrui Wang,Hengxu Wu,Honghui Ding,Hongxuan Tang,Huadong Wang,Huanqi Cao,Huazuo Gao,Hui Qu,Hui Zeng,J. Yang,J.H. Jin,J.H. Zhang,J.X. Zou,Jia Yu,Jiahui Zhou,Jiajun Chen,Jialiang Huang,Jialin Zhao,Jiamin Tang,Jian Zhou,Jianan Tong,Jianwen Li,Jiaqi Zhu,Jiarui Wang,Jiasheng Ye

类目:Computation and Language (cs.CL)

关键词:workloads increasingly input-heavy, increasingly input-heavy, widespread adoption, adoption of long-horizon, long-horizon agents

备注

点击查看摘要

Abstract:The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at this https URL.

44. 【2609.19965】Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

链接https://arxiv.org/abs/2609.19965

作者:Yutong Yao,Yanjie Cao,Guanhua Chen,Xu Yang,Junchao Wu,Zeyu Wu,Lidia S. Chao,Derek F. Wong

类目:Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:Large Language Models, Large Language, Language Models, existing work focuses, evidence largely unexplored

备注: Accepted by EMNLP 2026 Findings. Codes are available at: [this https URL](https://github.com/NLP2CT/PIJ-benchmark)

点击查看摘要

Abstract:Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely unexplored. To fill this gap, we introduce the Profiling, Investigation, and Judgment (PIJ), comprising 2,500 real homicide cases from five countries. PIJ evaluates LLMs across three tasks that span the entire criminal investigation pipeline: criminal profiling, which requires abductive reasoning to infer suspect attributes from fragmentary scene evidence, crime process reconstruction, which tests structured information extraction, and sentence prediction, which demands legal deductive reasoning. We evaluate 9 powerful LLMs and find that performance degrades systematically as tasks shift from explicit fact extraction to implicit reasoning over unknown suspect profiles. Categories requiring inferential reasoning, such as motivation and victim-offender relationships, remain the primary bottlenecks. Further analysis reveals substantial gaps between LLMs and human experts, along with pervasive biases in gender, age, and motive attribution. Our findings indicate that pre-arrest inference from incomplete evidence remains an open challenge.

45. 【2609.19942】Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

链接https://arxiv.org/abs/2609.19942

作者:Gunwoo Lee,Changmin Sung,Sang-Hwan Gwak,Ina Kim,Ji-Young Choi,Kyong-Ha Lee

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:extractive document question, document question answering, document question, question answering, confidence-driven mechanisms

备注: 26 pages main text + 26 pages supplementary (Online Resource 3). Submitted to Applied Intelligence. Code and data: doi: [https://doi.org/10.5281/zenodo.22710121](https://doi.org/10.5281/zenodo.22710121) , doi: [https://doi.org/10.5281/zenodo.22721044](https://doi.org/10.5281/zenodo.22721044)

点击查看摘要

Abstract:In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode -- area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion -- before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.

46. 【2609.19916】KoNeoBench: A Curated Evaluation Dataset for LLM Understanding of Korean Neologisms

链接https://arxiv.org/abs/2609.19916

作者:Soha Lee,Soojin Lee,Heesung Yang,Hyunju Song,Hyunji Lee,Jinsan An,Jeongwan Shin,Jin Hyun Park,Jun Lee,Hyeyoung Park,Kilim Nam

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, natural language constantly, language constantly evolves, Large language, newly emerging words

备注: Accepted to Findings of EMNLP 2026. Code and data are available at the project repository

点击查看摘要

Abstract:Large language models (LLMs) are typically evaluated on static benchmarks, even though natural language constantly evolves through newly emerging words and meanings. Existing Korean benchmarks are centered on established vocabulary and therefore provide limited coverage of such recent lexical change, and their English-oriented design makes it difficult to assess the typological properties of Korean, in which content words combine productively with functional morphemes. In this paper, we introduce KoNeoBench, a benchmark for evaluating LLMs' understanding of Korean neologisms. KoNeoBench is built on 1,785 Korean neologisms attested in online news since 2020 and curated through expert lexicographic review. Each entry provides usage examples, word-formation analyses, and dictionary-style definitions. Based on this resource, we define four tasks and report results on recent models, together with a human baseline. Our experiments show that current LLMs exhibit clear limitations in recovering source components, distinguishing semantic categories, and generating accurate definitions. These results reveal specific aspects of recent Korean lexical change that remain challenging for current LLMs. KoNeoBench is available at this https URL .

47. 【2609.19887】Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

链接https://arxiv.org/abs/2609.19887

作者:Giuseppe Samo,Vivi Nastase,Paola Merlo

类目:Computation and Language (cs.CL)

关键词:provide sufficient information, functional words, grammatical number, provide sufficient, replace concrete nouns

备注: 16 pages, 11 figures

点击查看摘要

Abstract:Pronouns, adverbs and other functional words (such as they, her, somewhere, there) are often used in language to replace concrete nouns or phrases, when their properties - such as gender, grammatical number - provide sufficient information for the given context. Do pretrained transformer models encode such functional words in a manner that allows them to be used like humans do? Can language models recognize the syntactic and semantic parallelism of sentences such as "The researchers wrote the paper" and "They wrote it", which relies on such lexical abstraction? We map these linguistic questions into the embedding space of a pretrained transformer model, and compare representations of nouns, with the representations of the pronouns and adverbs that can replace these nouns, in isolation and in parallel lexicalized and functional sentences. We then probe for shared syntactic and semantic structure in the embeddings of parallel lexicalized and functional sentences. We find that functional words are located centrally compared to nouns, but are also distinct, which is congruent with their behaviour as place-holders in a wide variety of contexts. The analysis of the embeddings of parallel (lexicalized and functional) sentences show them inhabiting different subspaces of the embedding space. Experiments that distil the structural information of the sentence show that training on either type of data does not reveal the shared structure - because of the over-consistency of the vocabulary (in case of the functional data), and the too much variety (in case of the lexicalized versions). However, training with a mix of functional and lexicalized sentences, the shared structure emerges.

Comments:
16 pages, 11 figures

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2609.19887 [cs.CL]

(or
arXiv:2609.19887v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2609.19887

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Giuseppe Samo [view email] [v1]
Thu, 17 Sep 2026 08:35:09 UTC (2,495 KB)

48. 【2609.19885】Evaluating Communicative Success in Machine-Translated Conversation

链接https://arxiv.org/abs/2609.19885

作者:Faiz Ghifari Haznitrama,Alice Oh

类目:Computation and Language (cs.CL)

关键词:increasingly mediate live, Interpreter agents built, mediate live conversation, increasingly mediate, isolated sentences

备注: 32 Pages, 11 Figures, 11 Tables

点击查看摘要

Abstract:Interpreter agents built on machine translation (MT) increasingly mediate live conversation between people who do not share a language, yet we still evaluate them with metrics built for isolated sentences, which measure fidelity rather than whether communication succeeds. We introduce a reusable three-layer checklist-and-judge framework that evaluates interpreter-mediated conversation across semantic, pragmatic, and cultural-social dimensions, covering the naturalness, intent, and social appropriateness that fidelity metrics leave unmeasured. It runs in both single-turn and interactive multi-turn settings, where simulated users reply to translated messages as the conversation unfolds and each turn is scored alongside the conversation as a whole. We extensively validate it through controlled perturbations, cross-judge comparisons, and human annotations. Our main single-turn benchmark evaluates 10 interpreter setups across Arabic, Bengali, Indonesian, and Korean from 5,624 OpenSubtitles-derived scenarios spanning 12 translation directions, and our multi-turn study covers all 6 language pairs in scripted and live modes. Results show a consistent decline from semantic to pragmatic and cultural-social success, while conventional MT metrics overlook failures among stronger interpreters, and prompt ablations show that scenario context, structured instructions, and cultural context improve communicative success, although gains vary across setups. Our work thus provides an evaluation framework and benchmark for interpreter agents in conversation, and highlights the importance of communicative success alongside existing translation metrics.

49. 【2609.19883】PetriBench: Benchmarking LLM Reasoning over Dynamic State Spaces

链接https://arxiv.org/abs/2609.19883

作者:Pyrros Koussios,Benjamin Jäger,John Hua Yao,Ajay Sridhar,Violet Xiang,Chenhao Li

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Logic in Computer Science (cs.LO)

关键词:Characterizing LLM reasoning, existing benchmarks isolate, benchmarks isolate specific, Characterizing LLM, specific reasoning skills

备注

点击查看摘要

Abstract:Characterizing LLM reasoning remains an open challenge, as many existing benchmarks isolate specific reasoning skills, rely on external knowledge, or are costly to extend. We introduce PetriBench, a compact, fully self-contained, and scalable benchmark for evaluating LLM reasoning over dynamic state spaces using Petri nets, a mature formalism for modeling real-world concurrent and distributed systems. PetriBench organizes reasoning into four task families varying by scope and temporal horizon, with Easy, Medium, and Hard levels generated by increasing structural complexity and evaluated against exact ground truth. Across a diverse set of proprietary and open-weight models, accuracy decreases consistently with difficulty, while harder instances expose increasingly distinct task-specific capability profiles. Additional analyses show that test-time compute improves performance but interacts differently with different reasoning tasks, and that procedural generation yields smooth scaling with structural complexity. Together, these results show that PetriBench provides a unified and extensible setting for probing the strengths, limits, and scaling behavior of LLM reasoning.

50. 【2609.19880】D-Quant: Driftable Entropy Coding for KV Cache Quantization

链接https://arxiv.org/abs/2609.19880

作者:Yi Su,Hong Liu,Guanghua Yu,Jianchen Zhu

类目:Computation and Language (cs.CL)

关键词:imposing substantial pressure, footprint grows linearly, memory footprint grows, deploying LLMs, batch size

备注

点击查看摘要

Abstract:The KV cache has become a major bottleneck in deploying LLMs, as its memory footprint grows linearly with sequence length and batch size, imposing substantial pressure on both memory capacity and bandwidth. Among various KV cache compression techniques, quantization is particularly attractive due to its effectiveness and ease of deployment. However, most existing methods rely on fixed-width quantization, where a $b$ bit representation is inherently limited to $2^b$ quantization levels. As the bit width decreases, the number of available levels shrinks exponentially, leading to severe information loss and rapid performance degradation. We further observe that fixed-width quantization fails to exploit the highly non-uniform distribution of KV cache. After rotation and normalization, KV values approximately follow a normal distribution, with most values concentrated near the center and only a small fraction appearing in the tails. Nevertheless, fixed-width coding allocates the same number of bits to frequent and rare symbols. Entropy coding naturally exploits such non-uniformity by assigning shorter codewords to frequent symbols and longer ones to rare symbols, substantially reducing the average number of bits required for representation. However, its variable-length output is not suited to highly parallel attention kernels, where efficient dequantization and computation rely on regular memory layouts and fixed-stride accesses. To bridge this gap, we propose \textbf{D-Quant}, a flexible KV cache quantization framework that introduces a \textbf{drift} mechanism to convert entropy-coded representations of each token into fixed-size bitstreams, enabling regular memory access and parallel dequantization within attention kernels.

51. 【2609.19879】VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

链接https://arxiv.org/abs/2609.19879

作者:Bhavana Akkiraju,Ravi Sastry Kolluru,Sri Charan D,Srihari Bandarupalli,Santosh Kesiraju,Anil Vuppala

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:Spoken question answering, Question answering, advanced rapidly, rapidly with large, predominantly for high-resource

备注: Paper is accepted in IEEE SLT 2026

点击查看摘要

Abstract:Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.

52. 【2609.19878】Uni-LaDiR: Latent Diffusion Unifies Multimodal Reasoning

链接https://arxiv.org/abs/2609.19878

作者:Haoqiang Kang,Yizhe Zhang,Nikki Lijing Kuang,Yian Ma,Lianhui Qin

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Multimodal reasoning requires, Multimodal reasoning, thought tokens, reasoning requires models, Unified Latent Diffusion

备注

点击查看摘要

Abstract:Multimodal reasoning requires models to draw on information from multiple modalities throughout the reasoning process. Yet existing methods often concatenate modality-specific thought tokens in a single sequence, leaving the model to bridge representational differences as it reasons across modalities. We introduce Uni-LaDiR (Unified Latent Diffusion Reasoner), a framework that brings these thoughts into a shared latent space for reasoning. A unified encoder maps teacher reasoning steps from different modalities into shared thought tokens, trained to preserve the information needed for later reasoning steps and the final answer or action. Because the same context can support multiple valid next steps, we use diffusion to predict the next block of thought tokens from the input and preceding blocks. Jointly training the encoder and diffusion reasoner with shared model weights encourages thought tokens to be both useful for the task and predictable from the available context. At inference, the model generates these tokens without teacher observations. Across eleven vision-language model (VLM) benchmarks and two vision-language-action (VLA) suites, Uni-LaDiR achieves relative gains over the strongest evaluated baselines of 7.3% on visual reasoning tasks and 6.1% on robot manipulation tasks.

53. 【2609.19877】JustMem: Just-Enough Memory Access for Long-Term Conversations

链接https://arxiv.org/abs/2609.19877

作者:Guanhua Chen,Yanting Wang,Wenjing Zhi,Lei Sha

类目:Computation and Language (cs.CL)

关键词:Efficient long-term conversational, Efficient long-term, requires retrieving sufficient, retrieving sufficient evidence, long-term conversational memory

备注: 12 pages, 8 tables, 3 figures. Includes appendix

点击查看摘要

Abstract:Efficient long-term conversational memory requires retrieving sufficient evidence without indiscriminately expanding the context presented to the language model. This is challenging because relevant evidence may be distributed across multiple sessions, while compression may discard details needed for answering. Different queries therefore require different forms of memory access. To capture these demands, we formulate memory access along two dimensions: discovery breadth, which controls how broadly evidence is searched, and reading fidelity, which controls whether evidence is read in compact form or recovered from the original conversation. Based on this formulation, we introduce JustMem, which stores conversation history as compact atomic memories and adapts memory access along these two dimensions to each query. Specifically, LOOKUP handles local evidence, COMPOSE broadens discovery for distributed evidence, and REPLAY increases reading fidelity for fidelity-sensitive evidence. On LoCoMo and LongMemEval-S, JustMem achieves the highest mean accuracy and retrieval recall among the compared memory systems while using substantially fewer generative-model tokens for memory construction and inference.

54. 【2609.19868】Zarya: A Hybrid Autoregressive--Masked Diffusion Language Model with Flexible Training and Dual-Mode Inference

链接https://arxiv.org/abs/2609.19868

作者:Leonid Sinev,Ilya Koziev,Vladislav Leshchuk

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:incoherent generation arising, high computational overhead, computational overhead due, incoherent generation, generation arising

备注: Preprint. Work in progress. Please cite peer-reviewed version when published

点击查看摘要

Abstract:Autoregressive language models (ARMs) are constrained by sequential, left-to-right generation, while masked diffusion models (MDMs) enable parallel decoding but suffer from high computational overhead due to the inability to reuse Key-Value (KV) cache and from incoherent generation arising from learning dependencies over an intractable space of token combinations. We introduce Zarya, a family of hybrid language models that jointly optimizes an autoregressive (AR) objective and a masked-diffusion objective within a single architecture. Zarya structures training data into variable-size slots and employs a curriculum that gradually increases slot granularity, enabling a smooth transition from fine-grained AR learning to coarse-grained diffusion learning. At inference, Zarya provides two distinct decoding paradigms through a unified interface: (i) MDM sampling with first-hitting denoising, and (ii) slotted speculative decoding that interleaves inter-slot diffusion-based selection with intra-slot autoregressive infilling, achieving full KV cache reuse. The training and inference regimes are fully decoupled, allowing a model trained with any configuration to be deployed in either mode. Extensive configurability --- including grouped noise patterns (Prefix Completion, Fill-In-the-Prefix, Fill-In-the-Middle), ordered sampling schedules, and noise-level permutation strategies --- enables flexible research exploration. We release Zarya models publicly in sizes 0.6B, 1.7B, and 4B, demonstrating performance on standard benchmarks while offering a principled integration of autoregressive and diffusion paradigms.

55. 【2609.19866】Reproducibility is not construct validity: LLM measurement of institutionally situated communication

链接https://arxiv.org/abs/2609.19866

作者:Veronika Batzdorfer(KIT),Carlo Romano Marcello Alessandro Santagiustina(ALMAnaCH, médialab, Sciences Po)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Risk Management (q-fin.RM)

关键词:High annotation reproducibility, High annotation, necessarily imply, Commission AI Act, High

备注

点击查看摘要

Abstract:High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.

56. 【2609.19827】F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

链接https://arxiv.org/abs/2609.19827

作者:Bojian Xiong(Tianjin University)Wentao Ding(Baidu Inc.)Yujing Lu(Baidu Inc.)Shaowei Zhang(Tianjin University)Ling Shi(Tianjin University)Jing Liao(Baidu Inc.)Yan Wang(Baidu Inc.)Yueyang Zhang(Baidu Inc.)Long Xia(Baidu Inc.)Zhiyuan Sun(Baidu Inc.)Daiting Shi(Baidu Inc.)Jingzhou He(Baidu Inc.)Yuqi Ren(Tianjin University)Deyi Xiong(Tianjin University)

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, complex user queries, widespread industrial deployment, resolving complex user

备注

点击查看摘要

Abstract:With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.

57. 【2609.19805】Dictionary-Constrained Grapheme-to-Phoneme for Unsegmented Languages from LLM-Annotated Data

链接https://arxiv.org/abs/2609.19805

作者:Rui Hu,Zhenpeng Zhan,Xiaolong Lin

类目:Computation and Language (cs.CL)

关键词:conversion turns raw, automatic speech recognition, turns raw text, conversion turns, speech recognition

备注: Submitted to ICASSP 2027

点击查看摘要

Abstract:Grapheme-to-phoneme (G2P) conversion turns raw text into its phonemic form and is an essential part of both text-to-speech (TTS) and automatic speech recognition (ASR) systems. It is required to be fast, stable and context-aware. For unsegmented languages such as Japanese, G2P additionally couples word segmentation with highly context-dependent polyphone disambiguation, and the scarcity of accurately annotated data remains a bottleneck. In this paper, we present a context-aware neural G2P method that scores paths of a discriminative conditional random field (CRF) over a word lattice constructed from dictionaries. To tackle data scarcity, we utilize large language models (LLMs) to generate more than 2 million sentences. Experimental results demonstrate that our method strongly outperforms conventional morphological analyzer-based methods and neural sequence models. On the Joyo-Kanji-Yomi benchmark, our method reaches 99.62% target word reading accuracy, 0.32% target word phoneme error rate (PER) and 0.14% sentence PER.

58. 【2609.19799】Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search

链接https://arxiv.org/abs/2609.19799

作者:Tal Oved,Roi Pony,Oshri Naparstek,Udi Barzelay

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:search finds programs, LLM-driven evolutionary search, evolutionary search finds, finds programs, programs by launching

备注

点击查看摘要

Abstract:LLM-driven evolutionary search finds programs by launching seeds and iterating each one. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point. We show this is not enough. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results. We run the analysis over a full grid of seeds and iterations. Our findings suggest that the best way to split a fixed budget between more seeds (width) and more iterations (depth) changes with the strategy, the task, and the total budget. Furthermore, we observe that the ranking of strategies also changes with the budget. On one task the strategy that looks worst at one seed is best at forty seeds. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score. We provide a measurement protocol that reports the seeds-by-iterations frontier and practical guidance for using it.

59. 【2609.19778】Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

链接https://arxiv.org/abs/2609.19778

作者:Bo Xu,Chenyuan Wang,Xinyu Chen,Quanhao Zhu,Rui Lin,Liang Zhao,Hongfei Lin,Feng Xia

类目:Computation and Language (cs.CL)

关键词:spread abusive content, memes spread abusive, hateful meme, Hateful memes spread, images and text

备注: 26 pages, 16 figures, 7 tables

点击查看摘要

Abstract:Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection results. However, we find that existing explain-then-detect methods often couple explanation generation and label prediction within the same training process. This coupling causes interference between task objectives, leading to limited detection performance and even worse results than simple SFT baselines. To address these challenges, we propose ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection. Inspired by the human annotation training process, ProKDA first uses an agentic background knowledge construction pipeline to obtain external knowledge related to meme understanding. It then adopts a three-stage training strategy that sequentially performs background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment. Unlike prior explain-then-detect methods that jointly optimize both tasks, ProKDA focuses on a single training objective at each stage. This design reduces interference between the two tasks and progressively transforms background knowledge into robust detection decisions. Experiments on three public hateful meme benchmarks show that ProKDA achieves state-of-the-art detection performance and provides accurate, explainable, and evidence-supported decisions for hateful meme moderation. Project page: this https URL.

60. 【2609.19754】AutoData: Agentic Search for Pre-training Data Selection

链接https://arxiv.org/abs/2609.19754

作者:Yan Meng,Dhruv Srikanth,Bingchen Zhao,Zhengyao Jiang,Yuxiang Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:recently shown promise, LLM agents, recently shown, shown promise, promise in automating

备注

点击查看摘要

Abstract:LLM agents have recently shown promise in automating machine learning engineering by editing model and training code under execution feedback. Data, however, remains largely outside this agentic optimisation loop. We frame pre-training data selection as heuristic engineering over per-document features, i.e., lexical statistics, categorical labels, and perplexity. We introduce AutoData, an agent that searches directly over executable selection algorithms. Unlike prior data mixture methods that optimise weights over a fixed set of domains, AutoData searches a richer program space of scoring, stratification, and stochastic selection rules, discovering feature interactions automatically by iteratively refining algorithms with validation feedback from a proxy model. Within an overnight search, AutoData discovers a selection algorithm that outperforms existing human-designed curation pipelines. Despite being searched only on this small proxy, the discovered recipe transfers to larger scales and improves the downstream metric CORE. These results suggest that data engineering can be treated as an agentic machine learning problem, extending autonomous research from model and training-code optimization to the data.

61. 【2609.19736】A Phonemically Comprehensive, ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design

链接https://arxiv.org/abs/2609.19736

作者:Zijie Zhang,Tan Lee

类目:Computation and Language (cs.CL)

关键词:Thai and Lao, ASCII-only romanization scheme, cross-lingual design problem, closely related languages, unified cross-lingual design

备注: Accepted by O-COCOSDA 2026

点击查看摘要

Abstract:This paper proposes a phonemically comprehensive, ASCII-only romanization scheme for Thai and Lao, treating the two closely related languages as a unified cross-lingual design problem. The scheme represents segmental contrasts, vowel length, and lexical tone while maintaining one-symbol-one-phoneme transparency and systematic correspondence between Thai and Lao. The scheme prioritizes synchronic phonetic correspondence, including correspondence with Pinyin and Jyutping where applicable, while preserving historical-phonological correspondence where it does not conflict with phonetic transparency. Tone uses a compact single-digit default notation, supplemented by optional tone-value and historical tone-category representations. The resulting scheme provides a readable, keyboard-friendly, and machine-processable phonemic representation for language learning and cross-lingual speech processing.

62. 【2609.19717】Learn Your Own Thoughts: Abstract Token Curriculum

链接https://arxiv.org/abs/2609.19717

作者:Khashayar Gatmiry,Avrajit Ghosh,Parsa Mirtaheri,Jason D. Lee,Nika Haghtalab,Emmanuel Abbe,Peter Bartlett

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:Large Language Models, Large Language, achieved remarkable reasoning, remarkable reasoning capabilities, Language Models

备注

点击查看摘要

Abstract:Large Language Models (LLMs) have achieved remarkable reasoning capabilities by utilizing chain-of-thought (CoT) as a scratchpad for intermediate stages of thinking. However, CoT techniques require explicit supervision on thinking tokens, which requires rich, task-specific data. In this work, we propose Abstract Token Curriculum (ATC), a novel curriculum learning framework that elicits effective continuous intermediate representations without direct supervision or manual scratchpad design. ATC gradually increases problem complexity through a sequence of distributions, training the model to develop internal abstract ``thoughts'' in the continuous representation space. This paper provides both theoretical and experimental evidence for the benefits of ATC and its advantages over previous methods for training continuous thoughts. Theoretically, we show that for learning parity functions with single-layer softmax attention using ATC, attention naturally focuses on the CoT tokens in the context that provide the ``easiest path'' to predicting the next token. Experimentally, we show ATC's effectiveness on graph reachability and arithmetic learning tasks.

63. 【2609.19650】Improving Cross-Lingual Transfer for Sequential Sentence Classification in Research Papers via Structural Similarity

链接https://arxiv.org/abs/2609.19650

作者:Kazuhiro Yamauchi,Marie Katsurai

类目:Computation and Language (cs.CL); Digital Libraries (cs.DL)

关键词:Sequential sentence classification, structuring scientific publications, extending SSC research, Sequential sentence, English can improve

备注: Accepted at JCDL 2026 (ACM/IEEE Joint Conference on Digital Libraries), Frisco, TX, USA, October 13-16, 2026. 12 pages, 5 figures, 9 tables. DOI: [https://doi.org/10.1145/3805696.3846040](https://doi.org/10.1145/3805696.3846040)

点击查看摘要

Abstract:Sequential sentence classification (SSC) is an essential task for structuring scientific publications, and extending SSC research to languages other than English can improve accessibility to scientific knowledge in multilingual digital libraries. Cross-lingual transfer is a promising approach to address the scarcity of training data in non-English languages. Prior work on other natural language processing tasks has shown the benefits of capturing linguistic similarity between source and target languages. However, SSC inherently depends on patterns at the discourse level, such as label sequences and positional regularities, which appear consistently across languages regardless of linguistic differences. To examine the factors that determine transfer success in SSC, we constructed a multilingual SSC dataset covering 13 non-English languages collected from five academic databases. Our cross-lingual transfer experiments, using both encoder-based and generative models, show that linguistic proximity has no consistent predictive power for transfer performance, whereas structural similarity in rhetorical organization shows a weak but consistent positive correlation across models. After controlling for source-language performance, the similarity of label distributions is the most consistent predictor. Building on this finding, we propose a set of three methods that explicitly leverage structural information using generative models. In the in-domain evaluation, the best combination reaches parity with the strongest encoder baselines, and in transfer to languages unseen during training, it outperforms the strongest encoder baseline.

64. 【2609.19634】Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

链接https://arxiv.org/abs/2609.19634

作者:Yinuo Zhang,Bingshuo Liu,Zhiying Tu,Dianhui Chu,Qingbin Liu,Xi Chen,Jiang Bian,Xiaoyan Yu,Dianbo Sui

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:image quality assessment, scientific image quality, Retrieval-Augmented Generation, complex scientific images, quality assessment

备注

点击查看摘要

Abstract:This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.

65. 【2609.19630】From Intent to Action: Benchmarking LLM Safety in Vehicle Voice Command Authorization

链接https://arxiv.org/abs/2609.19630

作者:Diba Afroze,Xingli Zhang,Yazhou Tu,Xiali Hei

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Robotics (cs.RO)

关键词:Large language models, Large language, vehicle voice assistants, voice assistants, increasingly integrated

备注

点击查看摘要

Abstract:Large language models (LLMs) are increasingly integrated into vehicle voice assistants. But linking natural-language requests to vehicle functions creates a safety-critical authorization problem. Before executing a command, the system must choose whether to execute, refuse, clarify, require confirmation, defer to manual control, trigger an emergency response, or make no tool call. To our knowledge, prior evaluations do not isolate this pre-action decision across speaker role, authentication status, vehicle state, and tool availability. We introduce a 202-scenario benchmark with Reference Decisions under a seven-class taxonomy. We evaluate two local open-weight models and three API-based LLMs using Decision Alignment and safety-specific error metrics. Alignment ranges from 40.1% for Llama 3.2 3B to 89.1% for Gemini 3.1 Pro Preview. The API-based models score between 83.2% and 89.1%, with no statistically significant differences among them. Even these models produce two to three False Executes among 161 non-execution scenarios, and persistent errors remain in confirmation and manual-control decisions. A controlled Llama 3.2 3B ablation increases alignment to 40.1% under the structured authorization policy, versus 28.2-29.2% under schema-only and generic-safety baselines, but it does not eliminate False Executes. Structured LLM decisions are therefore insufficient as a standalone safety mechanism, and deployment requires an independent enforcement layer that verifies tool permissions and vehicle-state constraints before invoking any vehicle function.

66. 【2609.19615】Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction

链接https://arxiv.org/abs/2609.19615

作者:Yuanzhe Jia,Ali Anaissi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Software Engineering (cs.SE)

关键词:Modern applications generate, heterogeneous event streams, generate massive volumes, applications generate massive, actionable business insights

备注

点击查看摘要

Abstract:Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge. Data engineers and analysts expend substantial effort reconciling semantic discrepancies, hand-crafting parsing logics, and maintaining fragile mappings between raw data and business KPIs. In this paper, we present an end-to-end framework that fully automates the construction of a business semantic layer from application raw logs. Our approach introduces a two-stage semantic abstraction: first, high-level business features are identified via LLM inference augmented with domain-specific industry knowledge; second, fine-grained business nodes are derived through a structured pipeline comprising data refinement, hybrid retrieval, multi-stage filtering, semantic clustering, and canonical naming. Evaluation on production-scale telemetry demonstrates that our system improves human-assessed semantic quality from 50 to 80+ on a 100-point scale, reduces maintenance effort by 80%, filters out 74% of noise, and achieves 0.87 Cohen's kappa via an integrated LLM-as-Judge evaluation, enabling continuous, scalable quality assurance. Overall, our work distinguishes itself from prior work by addressing the novel problem of business semantic layer induction from raw telemetry, operating without labeled training data or manual rule engineering.

67. 【2609.19606】Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction

链接https://arxiv.org/abs/2609.19606

作者:Theodore O. Cochran

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:dissociation Zhao reported, dissociation Zhao, Zhao reported, shape signal, model

备注

点击查看摘要

Abstract:This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model's chain-of-thought entropy trajectory predicts whether the final answer is correct, while the magnitude of its total entropy drop does not. The dissociation merits reproduction because the magnitude half rests on a single 300-problem run with one model at one seed, while the shape half was reported at full scale on both benchmarks and on a second model family. Registered at OSF before any confirmatory run, the reproduction crosses the complete GSM8K and MATH-500 benchmark test sets with four open-weight models including one reasoning-distilled model of a kind the original did not test. The shape signal replicates. The magnitude signal divides by setting. On the anchor model the accuracy gap between monotone and non-monotone chains is +9.6 percentage points on GSM8K and +27.5 on MATH-500, while the rank correlation of the total entropy drop with correctness is -0.018 on GSM8K and +0.414 on MATH-500. On the reasoning-distilled model the binary form of the shape signal fires on about one chain in a hundred, too few to estimate the registered contrast, while the graded violation count remains predictive there. In an exploratory comparison the final-step entropy alone outperforms the binary shape flag in all eight model-by-benchmark cells by ROC area, and in six or seven by the risk-coverage area the original reports, depending on an integration range the original does not state. The study contributes a reproduction of the shape signal at full test-set scale under seven documented protocol differences, a map of the settings where the magnitude signal holds and fails, and measurements of four protocol dependencies the original does not report.

68. 【2609.19596】Full-Duplex Speech Models Take the Floor When Asked, Not When Needed

链接https://arxiv.org/abs/2609.19596

作者:Linkai Peng,Baorian Nuchged,Kaiqi Fu,Yuyang Yao

类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)

关键词:promising always-on assistants, promising always-on, always-on assistants, speech models listen, models listen

备注: 5 pages

点击查看摘要

Abstract:Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2\,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.

69. 【2609.19589】Form Over Content In Gradient-Based Data Attribution Methods

链接https://arxiv.org/abs/2609.19589

作者:Sunwoo Kim,Seokwon Jung,Sohyung Kim,Seong Joon Oh,Alice Oh

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:select training data, large language models, answer format, measures is debated, format

备注

点击查看摘要

Abstract:Data attribution methods using gradient similarity are widely used to analyze and select training data for large language models, but what gradient similarity actually measures is debated. Some interpret it as identifying task-relevant skills, while other work reports that surface form is the main factor. We resolve this debate for supervised fine-tuning examples by varying task and answer format independently. Specifically, we render benchmarks in different answer formats, such that datasets can share a task without a format or a format without a task. We find that gradient alignment follows the answer format, as benchmark pairs sharing an answer format align strongly (disattenuated cosine near 0.4), while same benchmarks rendered with different answer format classes show no alignment (near 0.0). We demonstrate that this ordering holds from the earliest pretraining checkpoints through post-training, and across model scales and families. We then analyze the released selections of LESS, a gradient-based data selection method for instruction tuning, and find that each target's selections over-represent the target's own answer format. Hence, we demonstrate that gradient-based attribution methods track format similarity more than task semantics, meaning that such methods, as well as the semantic interpretation of the gradient, should be tested on data where answer format and task vary independently for greater robustness and reliability.

70. 【2609.19587】Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents

链接https://arxiv.org/abs/2609.19587

作者:Alex Remedios,Simon Storf,Fabien Roger,John Hughes

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:OpenAI Codex, Auto Mode, Claude Code, systems now review, review each proposed

备注

点击查看摘要

Abstract:To keep coding agents from going off the rails, production systems now review each proposed action with a blocking monitor that can reject it before it runs (Auto Mode in Claude Code, Guardian in OpenAI's Codex). Prior evaluations of such monitors largely measure robustness to accidental harm or prompt injections from untrusted sources looking to hijack the agent. Less understood is how they hold up when the agent they monitor is persistently misaligned. To understand this risk, we task an adversarial agent with evading production blocking monitors and causing catastrophic harm, e.g. by exfiltrating its own weights. We find that when instructed with high-level attack strategies, adversarial agents can succeed through several distinct mechanisms, such as agent-generated prompt injection against the monitor, multi-agent attacks, and malicious compaction. In particular we find that in 79% of trials, the agent can use an injection attack against Auto Mode and Guardian to run arbitrary bash commands. We also find that it is possible to greatly improve Auto Mode through design changes like enhancements to tool coverage, transcript formatting and an agentic monitor stage. Even with these improvements, preventing multi-context attacks at an acceptable cost remains an open problem. By detailing our red-teaming methodology and highlighting new attack vectors, we aim to help defenders evaluate their mitigations against the possibility of persistent malign coding agents. Code is available at this https URL.

71. 【2609.19585】CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

链接https://arxiv.org/abs/2609.19585

作者:Aiwei Ivy Zhang,Nimra Ishfaq,Mohit Chandra,Santiago Alvarez Lesmes,Adam Coscia,Khatiya Chelidze Moon,Xiaohan Ding,Munmun De Choudhury

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:reasoning over patient, task for clinicians, mental health care, key task, patient journeys

备注

点击查看摘要

Abstract:In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making temporal recovery challenging. We present CliniCIRCA, a multi-stage LLM framework for Calendar-anchored, Imprecision-aware Reconstruction of Clinical Annals. To our knowledge, CliniCIRCA is the first to temporally classify clinical events across unstructured discharge summaries without event-level timestamps. From 14,882 MIMIC-III mental health admissions, we first construct a benchmark of 52 discharge summaries on which CliniCIRCA produces 15,891 temporally tagged events. After correcting 629 errors based on a clinician-in-the-loop evaluation, we produce verified gold-standard labels. Finally, the corrected timelines drive a temporally grounded summarization stage that compresses each source 1.52 times into a date-grouped chronological record. We then scale the framework to generate 1,000 silver-standard timelines and evaluate them as training data. Compared with zero- and few-shot prompting, instruction tuning generally improves five open-weight models on event extraction, temporal tagging, and summarization across silver and clinician-verified evaluations.

72. 【2609.19553】From Parameters to Behaviors: A Survey of Model Fusion for Large Language Models

链接https://arxiv.org/abs/2609.19553

作者:Shuo Cai,Yanggan Gu,Zihao Wang,Yuanyi Wang,Yibo Yan,Wenjun Wang,Yuhang Liu,Guanghao Zhu,Sirui Huang,Ming Li,Hongxia Yang

类目:Computation and Language (cs.CL)

关键词:single target model, Hugging Face hosts, integrates the capabilities, capabilities from source, single target

备注: 25 pages, 4 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026

点击查看摘要

Abstract:Model fusion integrates the capabilities from source models into a single target model. As of June 2026, Hugging Face hosts more than 2M models. This growing pool provides a rich base for model reuse and capability integration. Yet existing surveys often cover only separate parts of this space, and they do not provide a unified definition or a systematic taxonomy. This survey defines model fusion and organizes prior work into three levels: parameter-level, representation-level, and behavior-level fusion. We also review related metrics, benchmarks, and applications, summarize current challenges, and identify future directions. Our goal is to provide a clear map of this area and support future work on model fusion. A comprehensive list of papers about model fusion is available at this https URL.

73. 【2609.19549】Finding Common Ground: Graded Communal Knowledge in Bluesky Starter Packs

链接https://arxiv.org/abs/2609.19549

作者:Sagar Kumar,Lawrence Swaminathan Xavier Prince,Julia Mendelsohn,Brooke Foucault Welles,Nicholas W. Landry

类目:ocial and Information Networks (cs.SI); Computation and Language (cs.CL)

关键词:Communication is made, common ground, common, online or offline, ground

备注

点击查看摘要

Abstract:Communication is made possible by common ground---the unspoken knowledge that people share and presuppose of one another, whether that be online or offline. In his conception of common ground, Clark (1996) distinguishes between personal and communal common ground, and asserts that the latter is graded: the more community affiliations two people share, the more common ground they share as well. Social media research has invoked this mechanism to explain how users connect, but it has gone largely untested because community memberships are rarely visible and, where they are, they are coupled to user interactions in a way that leads to conflating effects. To circumvent these challenges, this study repurposes Bluesky starter packs (SPs) as user-curated community affiliation labels. Across 191,648 pairs of users, we show that shared lexical repertoire---our proxy for common ground---grows monotonically with the number of SPs that users share, with users sharing a single pack being roughly twice as similar as equally connected strangers. A semantic renormalization of SP co-membership shows furthermore that it is more so the number of topically \emph{distinct} communities, rather than the raw count, in which common ground is graded. Finally, we show that community co-membership adds to common ground independently of proximity in the Bluesky follow network. These results lead to the conclusion that community membership is a measurable, separable, and semantically structured carrier of common ground. Reading it as such makes common ground observable before an exchange rather than inferred from it, and thus opens the door for large-scale observational approaches to a set of questions that have so far only been posed in the laboratory.

74. 【2609.19530】When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent Résumé Screening

链接https://arxiv.org/abs/2609.19530

作者:Jian Gao,Hang Jiang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:employers assess fit, employers assess, assess fit, candidates present, present and defend

备注: 9 pages, 5 tables, 1 figure. Accepted to the REALM Workshop at EMNLP 2026

点击查看摘要

Abstract:Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet résumé screening, the first gate, is commonly automated as a static, one-call judgment over a résumé-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed résumé-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%, respectively. This is not a uniform relaxation: two-agent screening rejects applications one-call advances, changing decisions in both directions. At similar pass volumes, the procedures advance different applications, and no one-call threshold recovers applications consistently selected by two-agent screening. Among discovery-selected cases re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a separate one-call follow-up shows no comparable decline. As hiring becomes agent-mediated on both sides, the screening procedure, not only the model behind it, shapes who reaches human review and how reliably that access recurs.

75. 【2609.19523】EconSkills: Studying Skill Transfer and Retrieval for Web Agents on Live Economic Data

链接https://arxiv.org/abs/2609.19523

作者:Yinzhu Quan,Zefang Liu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:earlier successful interactions, successful interactions, learned in earlier, earlier successful, evaluations discard

备注

点击查看摘要

Abstract:Web agents often revisit the same sites, yet most evaluations discard the procedures learned in earlier successful interactions. We introduce EconSkills, a skill library and evaluation framework that distills verified EconWebArena trajectories into parameterized standard operating procedures for retrieving live economic data. Each skill records its scope, navigation procedure, site-specific guidance, verification checks, and recovery steps while replacing source-instance values with placeholders. EconSkills separates two questions: whether a known relevant procedure transfers to a held-out task, and whether an agent can retain that benefit when selecting from a library. In controlled transfer, matched skills improve success over no-skill prompting and require fewer steps on paired successes, while abstraction is substantially more effective than replaying raw trajectories. At library scale, retrieval is competitive with the no-skill baseline overall and performs best on directly covered tasks; coverage-stratified outcomes show that approximate matches on uncovered tasks offset these gains. Browser trajectories further identify when procedural guidance shortens portal-specific navigation and when semantic verification remains necessary. These results establish that reusable economic web procedures can transfer across task instances and provide a concrete design target for coverage-aware selection and context delivery.

76. 【2609.19504】For Your Eyes Only: Evaluating Coordination Between Isolated Language Model Instances

链接https://arxiv.org/abs/2609.19504

作者:Alexander Shirnin,Aleksey Kudelya

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:important question arises, practically important question, automated workflows, question arises, task instructions

备注

点击查看摘要

Abstract:As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any shared memory or coordination-specific training? We introduce For Your Eyes Only, a cooperative signalling game designed to evaluate this directly. A Sender produces free-form descriptions for two words, one of which is a hidden target; an isolated Receiver must identify it. We evaluate seven contemporary models from four architectural families on 300 word pairs from established psycholinguistic corpora, using the Double-Pass Success Rate to control for output biases. We find that most models struggle to maintain coordination once they are required to avoid detectable signals, while one frontier model retains near-perfect performance even after such filtering. We further show that models can direct this capability toward deliberate misdirection, and that coordination is consistently weaker across architectures than within them.

77. 【2609.19472】Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

链接https://arxiv.org/abs/2609.19472

作者:Alizishaan Khatri,Chiquita Prabhu,Omkar Neogi

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)

关键词:Autonomous systems increasingly, Large Language Models, Large Language, systems increasingly rely, safety infrastructure surrounding

备注

点击查看摘要

Abstract:Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.

78. 【2609.19445】From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning

链接https://arxiv.org/abs/2609.19445

作者:Pan Wang,Siwei Song,Hui Ji,Siqi Cao,Heng Yu,Zhijian Liu,Huanrui Yang,Yingyan Celine Lin,Beidi Chen,Mohit Bansal,Xiaoming Liu,Pengfei Zhou,Ming-Hsuan Yang,Tianlong Chen,Jingtong Hu

类目:Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:surfaced formidable bottlenecks, pivotal research frontier, bottlenecks in computation, catalyzing the rise, Efficient Multimodal Learning

备注: TMLR

点击查看摘要

Abstract:The rapid expansion of multimodal models has surfaced formidable bottlenecks in computation, memory, and deployment, catalyzing the rise of Efficient Multimodal Learning (EML) as a pivotal research frontier. Despite intensive progress, a cohesive understanding of what, how, and where efficiency is manifested across the learning stack remains fragmented. This survey systematizes the EML landscape by introducing the first structured, model-to-system taxonomy. We distill insights from over 300 seminal works into three hierarchical levels--model, algorithm, and system--addressing architectural parsimony, execution refinement, and hardware-aware orchestration, respectively. Moving beyond a purely categorical review, we offer a methodological synthesis of the vertical synergies between these layers, elucidating how cross-layer co-design contributes to the fundamental "Efficiency-Utility-Privacy" trade-off. Through an integrative case study of Multimodal Large Language Models (MLLMs), we trace the field's evolutionary trajectory from initial structural adjustments to modern full-stack resource orchestration. Furthermore, we provide a holistic discussion and application-specific optimization blueprints for diverse domains and posit a paradigm shift toward self-regulating intelligence, where efficiency is an intrinsic, emergent property of the model's fundamental design rather than a post-hoc constraint. Finally, we present open challenges and future directions that will define the trajectory of EML research. This survey establishes a structured framework for multimodal systems that are not only high-performing and generalizable but natively efficient and ready for ubiquitous deployment. A continuously updated version is available at this https URL.

79. 【2609.19422】BurnRiSc: Toward Non-Invasive Burnout Screening in Open Source from Public Repository Signals

链接https://arxiv.org/abs/2609.19422

作者:Timofey Sanko,Yuan Tian,Mariam Guizani

类目:Computers and Society (cs.CY); Computation and Language (cs.CL)

关键词:chronic occupational syndrome, absorb unbounded demand, maintainers absorb unbounded, occupational syndrome, notice decline

备注: 8 Pages, Submitted to the JAWs 2 Workshop

点击查看摘要

Abstract:Burnout is a chronic occupational syndrome, and open source is close to a worst case for it: maintainers absorb unbounded demand with no manager to reallocate work and no organization to notice decline. The cost is not only personal. Burnout precedes withdrawal, and in projects sustained by a handful of maintainers, one departure can break infrastructure that thousands of downstream systems depend on. Yet the field has no way to see it coming: self-report inventories, the only existing measure, miss exactly the contributors most in need of detection and cannot be applied retroactively, so the field cannot even ask how common burnout is or what helps. We present BurnRiSc, a framework that operationalizes the Oldenburg Burnout Inventory's two dimensions, exhaustion and disengagement, as 14 behavioral and linguistic signals computed from GitHub activity and scored against each contributor's own history. The signals aggregate into two weighted dimension scores, with weights learned from labeled cases, and average into a monthly Burnout Risk Score (BRS). In a preliminary evaluation across 68 contributors in ten repositories (ten disclosed burnout cases, twelve comparable-volume collapses, and 46 comparison contributors), sustained BRS elevation precedes 6 of 10 disclosures by 6-15 months, 8 of 10 when adding peak BRS as a second criterion, and 10 of 10 over any prior time frame. We thus present BurnRiSc as evidence that burnout is screenable from public data.

Comments:
8 Pages, Submitted to the JAWs 2 Workshop

Subjects:

Computers and Society (cs.CY); Computation and Language (cs.CL)

Cite as:
arXiv:2609.19422 [cs.CY]

(or
arXiv:2609.19422v1 [cs.CY] for this version)

https://doi.org/10.48550/arXiv.2609.19422

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
80. 【2609.19417】Less Is More: Graph-free Multimodal RAG via Multi-signal Late Fusion

链接https://arxiv.org/abs/2609.19417

作者:Tithi Rakshit,Hongkuan Zhou,Lavdim Halilaj,Yuqicheng Zhu

类目:Computation and Language (cs.CL)

关键词:Graph-based retrieval-augmented generation, cross-document question answering, retrieval-augmented generation, RAG, Graph-based retrieval-augmented

备注

点击查看摘要

Abstract:Graph-based retrieval-augmented generation (RAG) is widely used for multimodal, cross-document question answering. However, building corpus-level graphs is expensive, slow to query, and difficult to maintain. We present TrioRAG, a graph-free multimodal framework that integrates evidence from three complementary signals: the question, the anchor image, and a VLM-enhanced query generated from both. Each signal retrieves independently over a shared multi-vector index of page text and page images, and the results are combined through late fusion. Further, we introduce AutoQA, a multimodal automotive benchmark whose questions are grounded in noisy, web-sourced images rather than clean document-sourced figures. Its questions require reasoning across manuals. We position it as a model-curated testbed rather than a human-validated gold standard. Across three benchmarks, TrioRAG matches or outperforms graph-based systems while reducing total cost and accelerating per-query inference by 1.6-2.3 times. By construction, AutoQA grounds its questions in out-of-corpus web images. In this setting image retrieval reaches only 19.3% document-level recall, while text-derived signals, especially the VLM-enhanced query, keep retrieval robust.

81. 【2609.19398】A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech

链接https://arxiv.org/abs/2609.19398

作者:Roksana Khanom,Raghib Asfak Tasnim,Bodrun Nahar Bithi,Shafia Shirin Supty,Saiful Islam Raju,Ashok Agrawala,Nirupam Roy

类目:ound (cs.SD); Computation and Language (cs.CL)

关键词:respiratory health assessment, language-specific phonetic variation, languages remain challenging, Language Invariance Score, offers a scalable

备注: Under review

点击查看摘要

Abstract:Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by language-specific phonetic variation. We present CL-DAF, a Cross-Lingual Disease-Alignment Framework that identifies acoustic dimensions whose disease effects remain consistent across languages. Using 201 English and 75 newly collected Bangla speakers, we construct a common 272-dimensional acoustic representation and quantify disease alignment using signed rank-biserial effects and the Language Invariance Score. We first show that spontaneous Bangla speech separates COPD from controls (AUC 0.85); however, 133 features reverse their disease direction across languages and the full representation transfers poorly (AUC 0.49 from Bangla to English). CL-DAF isolates 26 disease-aligned features that raise AUCs to 0.825 and 0.722 from English to Bangla and Bangla to English, respectively. These findings provide a foundation for multilingual clinical speech models emphasizing pathology over language-dependent variation.

82. 【2609.19384】Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models

链接https://arxiv.org/abs/2609.19384

作者:Badri N. Patro,Vijay S. Agneeswaran

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)

关键词:Scaling deep learning, exponential training costs, faces critical bottlenecks, deep learning faces, learning faces critical

备注

点击查看摘要

Abstract:Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogeneous merging setting that retains both architectures while aligning parameter groups by semantic role. Our proposed Riemannian--Lorentz Parameter Fusion (RLPF) method projects aligned groups to common coordinates, lifts selected coordinates to the Lorentz hyperboloid model of hyperbolic space, computes a regularized geodesic barycenter, and decodes the result into the two branches. A learned gate then combines branch logits for each input. Component groups use fixed curvature values, with normalization parameters treated as Euclidean. In the results available in this manuscript, the fine-tuned system obtains 82.37\% on CIFAR-10, 75.04\% on Oxford-IIIT Pet, and 78.58\% top-1 accuracy on ImageNet-1K; the corresponding best-parent accuracies are 76.54\%, 71.42\%, and 76.42\%. On ImageNet-1K, the reported pre-fine-tuning initialization reaches 77.80\%. These results support further study of geometry-aware heterogeneous fusion, but not a training-free single-checkpoint merge: RLPF is a two-branch hybrid whose gate and reported final models are trained.

83. 【2609.19366】he Role of Fine-grained Harm Signals in LLM Safety

链接https://arxiv.org/abs/2609.19366

作者:Soyeon Park(1),Seogyeong Jeong(1),Sunwoo Kim(1),Alice Oh(1) ((1) KAIST)

类目:Computation and Language (cs.CL)

关键词:Prior work, general harm representation, common general harm, general harmfulness representation, large language models

备注: 9 pages, 6 figures

点击查看摘要

Abstract:Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm representation in LLM safety. To answer this question, we isolate the category-specific component by removing shared general harmfulness representation from each categorical harmfulness representation, yielding a category residual that is orthogonal to general harmfulness at every layer. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, we find that whether category residuals encode harmfulness varies across categories, and that this category-wise pattern is similar across models. Whether category residuals induce refusal also varies across categories, but this category-wise pattern is more model-dependent. We also find that category residuals increase LLMs' downstream internal alignment with shared general harmfulness representation. Together, these findings demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety. More broadly, our findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification.

84. 【2609.19334】A frontend-backend architecture for tool calls in full-duplex speech models

链接https://arxiv.org/abs/2609.19334

作者:Ke Hu,Slyne Deng,Chen Chen,Elena Rastorgueva,Edresson Casanova,Punit Kumar,Dharmendra Choudhary,Nikhil Srihari,Ameya Sunil Mahabaleshwarkar,Viet Anh Trinh,Slim Essid,Oluwatobi Olabiyi,Zhehuai Chen

类目:Computation and Language (cs.CL)

关键词:complete voice-agent tasks, voice-agent tasks, complete voice-agent, low-latency conversational interaction, forwards streaming ASR

备注

点击查看摘要

Abstract:Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend learns to emit a delegation token and forwards streaming ASR transcripts to a text-based backend LLM for tool calls. Tool-call results from the backend are injected back into the frontend through a lightweight prefill-and-repeat mechanism and then synthesized using streaming TTS to the user. Our approach largely preserves regular duplex turn-taking, interruption handling, and low-latency interaction as it requires minimal modifications to the frontend model. In a single-turn tool-call evaluation, our system achieves 92-97% tool-call recall, competitive tool-call prediction performance, and 81.2% accuracy in rejecting irrelevant calls. When equipped with a larger backend (e.g., Qwen3-235B-A22B), our system achieves competitive results on Full-Duplex-Bench-V3 compared to open and closed source models, and significantly outperforms GPT-realtime-mini and Qwen3-Omni-30B-A3B-Instruct on EVA-Bench. These results demonstrate that backend delegation is an effective and modular approach for combining natural duplex speech interaction with strong agentic tool-call capabilities.

85. 【2609.19325】AUDITPLAN: Commit, Then Answer for Auditable Safety Alignment

链接https://arxiv.org/abs/2609.19325

作者:Sai Sri Pushpa Jampani,Kshitij Mishra,Asif Ekbal

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:tuning pipelines judge, Safety tuning pipelines, distinguish robust refusal, blanket refusal, undesirable shortcuts

备注

点击查看摘要

Abstract:Safety tuning pipelines judge only the final answer, which makes it difficult to distinguish robust refusal from two undesirable shortcuts: blanket refusal on benign requests and polished but unfaithful safety rationales that do not actually constrain the answer. We propose AUDITPLAN, a single-model plan-then-answer approach where the model first emits a compact structured safety plan and then answers conditioned on it. The plan records a threat label, intended action, and explicit constraints, enabling machine-checkable auditing while remaining hidden from users at deployment. We train this behavior with supervised fine-tuning followed by reinforcement learning with FAITHGATE, a reward-gating objective that grants answer reward only when the safety plan is correct. This discourages safe-looking but unfaithful behavior and promotes tighter plan-answer coupling. Across Qwen backbones, AUDITPLAN improves both robustness and auditability: on Qwen2.5-3B-Instruct, FAITHGATE reduces ASR from 24.0% to 11.6%, LSR from 1.0% to 0.36%, and over-refusal from 11.0% to 2.0%, outperforming answer-only RL, free-form explanation, and weighted-sum structured rewards. Similar trends hold for Qwen2.5-1.5B-Instruct. Larger-model confirmation runs on Qwen-3-4B-Instruct and Qwen2.5-7B-Instruct preserve the same trend suggesting that explicit internal commitments can make safety alignment more faithful, robust, and auditable.

86. 【2609.19291】Why Pretraining Fails to Share Cross-Lingual Knowledge

链接https://arxiv.org/abs/2609.19291

作者:Adam Gaber,Uriel Dolev,Elisabeth Fittschen,Bobby Cheng,Yuval Marton,Leshem Choshen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Large Language Models, made remarkable progress, Large Language, Language Models, made remarkable

备注

点击查看摘要

Abstract:Large Language Models (LLMs) have made remarkable progress in the processing and modeling of many languages. Yet, unlike human multilinguals, they exhibit surprisingly limited cross-lingual knowledge transfer. While this limitation is well documented, its origins during multilingual training remain unclear. We pretrain 360M- and 7B-parameter LLMs and show that poor cross-lingual knowledge generalization emerges during pretraining and persists under standard interventions. To isolate its cause, we employ a controlled bilingual pretraining setting using two copies of the same language, sharing identical text and token segmentation, but mapped to disjoint token spaces. We find that disjoint tokens alone are enough to induce knowledge compartmentalization, even between identical copies of the same language, establishing disjoint token spaces as a fundamental barrier to cross-lingual knowledge generalization. Guided by this understanding, we suggest mapping languages into a shared token space by simple word-wise translation and find it substantially improves cross-lingual knowledge generalization, recovering up to 12.6\% of native-language learning efficiency --- 14$\times$ the baseline.

87. 【2609.19238】YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers

链接https://arxiv.org/abs/2609.19238

作者:Hao Yang,Jin Wang,Xuejie Zhang

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Bridging the Gap, Robustly Optimized BERT, Optimized BERT Approach, YNU-HPCC team, Gap in Text-Based

备注

点击查看摘要

Abstract:This paper describes the participation of the YNU-HPCC team in subtask A of task 11, Bridging the Gap in Text-Based Emotion at SemEval-2025. Our best-performing system employs the RoBERTa (Robustly Optimized BERT Approach) model, an improved version of BERT that utilizes the Transformer encoder architecture. We enhanced the output head to allow the model to process one emotion simultaneously. We obtained the official ranking score (0.44), including results from all languages. The entire dataset was translated into English using Google Translate to facilitate subsequent processing. Through probabilistic and attention analyses, we found that (I) a single prediction head performs better than six heads predicting six emotions simultaneously, and (II) training on a uniformly translated English dataset yields better results than using the original dataset. The code is available at: this https URL.

88. 【2609.19189】CovR: Coverage-Aware Hardware Verification via Reasoning-Guided Reinforcement Learning

链接https://arxiv.org/abs/2609.19189

作者:Manar Abdelatty,Maryam Nouh,Sherief Reda

类目:Hardware Architecture (cs.AR); Computation and Language (cs.CL)

关键词:total design effort, Design verification remains, design effort, total design, Large Language Models

备注

点击查看摘要

Abstract:Design verification remains one of the most resource-intensive stages of hardware development, often consuming up to 70% of the total design effort. While recent work has explored using Large Language Models (LLMs) to automate testbench generation, most existing approaches focus narrowly on functional correctness, overlooking the critical aspect of coverage quality. To bridge this gap, we present CovR, an agentic framework for automated testbench generation that combines self-reflection loops with simulation-based feedback to maximize coverage. Using this pipeline, we construct a large-scale dataset of 16,514 natural specification RTL reasoning testbench tuples with a strong teacher model, enabling coverage-aware supervision. Building on this, we propose a reinforcement learning (RL) framework tailored for coverage-driven testbench generation, leveraging tool-derived rewards from simulation and coverage feedback to optimize a student model. Experimental results show that the CovR finetuned model achieves 93.81% cov@10 on VerilogEval and RTLLM V2.0, and 87.76% cov@10 on CVDP, outperforming state-of-the-art approaches by 7.97% and 3.59%, respectively. Furthermore, deploying the finetuned model back into the agentic refinement pipeline further improves cov@10 to 94.27% on VerilogEval and RTLLM V2.0 and 91.39% on CVDP. Moreover, when integrated as a plug-in stimulus engine for full verification workflows, CovR improves coverage by 18.95% and mutation detection score by 1.19%, while revealing 4.46% undetected failures, highlighting the importance of optimizing for coverage in LLM-based hardware verification.

89. 【2609.19183】Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

链接https://arxiv.org/abs/2609.19183

作者:Makoto Fukushima

类目:Multiagent Systems (cs.MA); Computation and Language (cs.CL); Adaptation and Self-Organizing Systems (nlin.AO); Physics and Society (physics.soc-ph)

关键词:large language model, bounded by cognition, human or large, large language, claim wording sets

备注

点击查看摘要

Abstract:Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others' messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model's judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network's degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim's wording sets for the agent's answer before any message is read: the experimental claims' fields lay below the calibration mean, and with each claim's own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective's fate is largely set by two single-agent measurements: the threshold a claim's wording sets, and the message capacity that sets the transition point.

90. 【2609.19182】What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

链接https://arxiv.org/abs/2609.19182

作者:Chao Wang(Independent Researcher)

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, assessed and communicated, progress in large, large language, language models

备注: 15 pages, 5 figures, 7 tables. Data and code: [this https URL](https://github.com/xxcg322/LLM-Bench-Map)

点击查看摘要

Abstract:Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?

91. 【2609.19167】o Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives

链接https://arxiv.org/abs/2609.19167

作者:Wenqi Zhou,Zhuorui Yu,Kaiao Wen,Hao Zheng,Xinyi Zheng,Peiran Wu,Enmin Zhou,Chi-Hao Wu,Junxiao Shen

类目:Computation and Language (cs.CL)

关键词:storing past events, tracking longitudinal experiences, long-term personal history, past events, Real-world Long-term Multimodal

备注

点击查看摘要

Abstract:As AI systems evolve into personalized digital companions, a central capability is reasoning over a user's long-term personal history: not merely storing past events, but tracking longitudinal experiences and evolving preferences. Progress here is bottlenecked by evaluation, existing long-term memory benchmarks are largely synthetic and text-only, they overlook the visual records that anchor everyday human memory, lack the authentic and causally connected longitudinal data that real personalization demands, and consequently remain confined to shallow factual recall. We introduce ReaLMem (Real-world Long-term Multimodal Memory), the first benchmark built from authentic multi-year personal visual archives, paired with first-person subjective annotations. ReaLMem evaluates models across three cognitive tiers of increasing difficulty: factual recall, persona inference, and predictive personalization. We further propose ChronoProfiler, a temporal-weighting profiling module that computes temporal stability scores for user attributes and applies them as a salience prior, resolving conflicts among temporally inconsistent preferences and helping models compound multiple co-active preferences in complex personalized decisions. Extensive evaluation of frontier multimodal large language models (MLLMs) and memory systems on ReaLMem reveals predictive personalization as a consistent ceiling, exposes clear performance gaps and bottlenecks between MLLMs and memory systems, and shows that high-quality, temporally informed representations substantially improve personalization. Together, ReaLMem and ChronoProfiler provide an authentic testbed and a simple, effective mechanism for long-term personalization, laying a foundation for future research on lifelong AI companions.

92. 【2609.19164】Advantage Scale Calibration Imbalance in Group-Relative Optimization under Low-Variance Rewards: Diagnosis and Bounded Recovery

链接https://arxiv.org/abs/2609.19164

作者:Fei Ding

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:verifier-style RLVR, treats advantage scale, group-relative optimization, implementation detail, optimization often treats

备注: 10 pages,2 figures

点击查看摘要

Abstract:In verifier-style RLVR, group-relative optimization often treats advantage scale as an implementation detail. This paper separates two low-variance cases: sub-resolution jitter that should not become a preference signal, and credible but small cardinal gaps that should be learned without distorting KL calibration. We propose an advantage-scale three-way calibration interface: the same within-group scale denominator simultaneously determines the reward-branch strength, prompt-level batch weight, and the effective KL calibration induced when the reward branch is re-expressed on the original cardinal scale. This interface explains why RLOO / this http URL can let credible small gaps become KL dominated, whereas GRPO's standard-deviation denominator can amplify tiny gaps without bound. Based on this interface, we further introduce the Reward-Resolution Protocol and MaxNorm-AC, respectively filtering sub-resolution gaps and providing bounded cardinal recovery on credible nonzero gaps. Across dense / MoE architectures and math / code reasoning, MaxNorm-AC improves over the strongest robust-scale baseline while truncating the low-variance inverse-scale tail.

93. 【2609.19158】VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

链接https://arxiv.org/abs/2609.19158

作者:Yixin Peng,Er Jin,Shiwei Luo,Diego Collarana,Stefan Decker

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:graph neural network, retrieved knowledge graphs, Knowledge graphs, neural network, network and fusing

备注

点击查看摘要

Abstract:Knowledge graphs are usually integrated into question answering by encoding a retrieved subgraph with a graph neural network and fusing it with the language model in the online inference path. The same subgraph is therefore re-encoded from scratch every time a pair is scored, across training epochs, seeds, and evaluation runs, even though the knowledge graph never changes. We ask whether the retrieved knowledge graphs can instead be compiled once, offline, and then accessed as read-only memory. VisKG-LM shows that it can, by decoupling graph encoding from language reasoning. It serializes each retrieved candidate-specific subgraph as Relation-Labeled Paths and renders the result as an image whose two-dimensional layout preserves the branching structure of the paths. Each image is encoded once, offline, and cached for reuse. At inference, the language model contextualizes the question and candidate from text alone, and only its final layer consults the cached visual memory, reading both its global layout and its local relational detail. The graph information thus enters only after the text has been understood. On the test sets of CommonsenseQA, OpenBookQA, and MedQA-USMLE, VisKG-LMimproves over GreaseLM by $1.2$, $0.8$, and $4.3$ points, respectively, while matching or surpassing GraphVis, a $7$B vision-language model, with only about $400$M online parameters. Against a matched text-only control that receives the identical Relation-Labeled Paths, it gains $4.2$, $6.5$, and $5.1$ points across the three benchmarks. These gains show that the complete visual-memory interface adds value beyond path textualization alone and support compiled visual memory as an alternative to online graph propagation.

94. 【2609.19156】Reflective Recovery: A Self-Supervised Method for Reasoning by Learning from Mistakes

链接https://arxiv.org/abs/2609.19156

作者:Qirui Chen,Renjie Pi,Jiahui Gao,Lingpeng Kong

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, Data-driven fine-tuning, Language Models, simplicity and efficiency

备注

点击查看摘要

Abstract:Data-driven fine-tuning is widely adopted to enhance reasoning in Large Language Models (LLMs) due to its simplicity and efficiency. However, mainstream imitation learning methods that rely exclusively on perfect reasoning trajectories suffer from a Scaling Collapse: when the problem set is limited, increasing positive examples fails to yield continuous improvement. However, during inference, an LLM can not guarantee that every intermediate step is correct and is therefore prone to errors. Once such errors arise, the LLM often struggles to recover and may be further misled by the accumulation of previous mistakes. To address this, we propose Reflective Recovery, a simple yet effective self-supervised approach that transforms failed reasoning attempts into recovery training data. Specifically, we extract initial segments of failed trajectories, concatenate them with prompts, and use them to guide the LLM toward valid solutions. Because these segments from failed trajectories are likely to contain errors, this process teaches models to recognize and correct mistakes during reasoning, enabling recovery from erroneous states without relying on external critics or reward models. Evaluated on extensive benchmarks, Reflective Recovery significantly improves performance. On DeepSeek-R1-Distill-Qwen-7B, it boosts accuracy from 30.0% to 37.5% on AIME 2025 and from 37.6% to 47.8% on Minerva. More importantly, analyses demonstrate that it breaks the scaling collapse barrier and enables models to develop emergent self-correction behaviors, representing a paradigm shift from outcome-oriented memorization to process-oriented reflective reasoning.

95. 【2609.19155】owards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue

链接https://arxiv.org/abs/2609.19155

作者:Jinqiang Wang,Tao Zhu,Huansheng Ning

类目:Computation and Language (cs.CL)

关键词:earlier intents, generate inappropriate responses, user-side, user-side conflict detection, user-side conflicts

备注: 24 pages, 13 figures

点击查看摘要

Abstract:In Human-LLM dialogue, follow-up user utterances may implicitly conflict with earlier intents, leading the LLM to misinterpret user needs and generate inappropriate responses. A reliable dialogue system should proactively detect user-side conflicts before generating a response and seek clarification when necessary. However, prior work has largely focused on LLM-side conflicts, leaving user-side conflicts underexplored. To fill this gap, we construct UC-Bench, a human-annotated benchmark for evaluating user-side conflict detection. Preliminary experiments show that existing LLMs struggle with this task, especially when conflicts arise from implicit incompatibilities grounded in dialogue history. To improve lightweight LLMs with limited training data, we investigate data synthesis for user-side conflict detection. Existing synthesis methods do not explicitly model the implicit incompatibilities between historical and current user utterances, making it difficult to capture the evolution of conflicts and to generate reliably labeled implicit conflict samples. We propose SynUC, a constraint-guided synthesis method that represents user-side conflicts in a constraint space and uses the SPEAKING framework to guide traceable constraint transformations. Applying SynUC to WildChat, we construct UC-Data, a user-side conflict training set containing 2,487 samples. On UC-Bench, Qwen3.5-4B trained on UC-Data outperforms larger general-purpose LLMs such as Claude Opus 4.8, as well as the same backbone trained on data synthesized by existing methods.

96. 【2609.19154】Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry

链接https://arxiv.org/abs/2609.19154

作者:Han Zhang,Zihan Gu,Zhiyuan Wang,Tianyi Ma,Jiacheng Lu,Xinyan Zhang,Yuhao Wei,Cheng Hua

类目:Computation and Language (cs.CL)

关键词:established Classical Chinese, Classical Chinese Poetry, Large Language Models, Large Language, Classical Chinese

备注: Published in the Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026

点击查看摘要

Abstract:While Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. To address this issue, we introduce Neo-Classic, an evaluation benchmark that combines a constructionist Out-of-Sample (OOS) dataset with a suite of reverse understanding probes. Unlike traditional benchmarks that rely on verification or generation over historical corpora, Neo-Classic comprises strictly metrical poetry authored by contemporary experts, reducing the possibility of direct retrieval. We evaluate state-of-the-art models, including Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2, across five behavioral probes designed to test hierarchical constraint satisfaction. Our results reveal two primary limitations. First, a performance gap of 20 to 50 percent emerges when models transition from historical to contemporary texts. Second, models exhibit substantial difficulties in discourse-level ordering tasks, with standard accuracy remaining low (0 to 13 percent). Although expert-level guidance improves the performance of reasoning-enhanced models to 36 percent, a notable gap with human experts persists. These findings suggest that while current LLMs capture local formal patterns, they struggle with global hierarchical planning required for robust Linguistic-Aesthetic Reasoning.

97. 【2609.19153】Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

链接https://arxiv.org/abs/2609.19153

作者:Gregory M. Dickinson

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:scholarship increasingly treats, increasingly treats judicial, Empirical legal scholarship, treats judicial text, legal scholarship increasingly

备注

点击查看摘要

Abstract:Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines -- TF-IDF features and linear classifiers -- because the textual feature is often the object of study, not merely a means to a prediction. Yet these pipelines inherit a chain of preprocessing defaults from mid-century information retrieval that were never validated against classification accuracy, the most entrenched being stopword removal. This study introduces an exhaustive single-word ablation that measures a preprocessing step's effect directly against the downstream objective, and applies it to stopword removal as the hardest case to dislodge. Matching Supreme Court Database labels to Caselaw Access Project opinion texts, it examines two binary tasks that bracket F1 headroom, ideological direction (no-removal baseline F1 ~ 0.68) and constitutional versus non-constitutional law type (~ 0.92), across 7,668 and 7,001 opinions. For each task the analysis approximates the best stoplist any expert could build, removing each of roughly 18,500 candidate words and measuring the effect directly. Three findings follow: generic stoplists in common use fall below the no-removal baseline in every test; even optimized stoplists are statistically indistinguishable from removing nothing; and meta-models trained on word-level features cannot predict which removals help, so list curation has nothing to target. The method generalizes to any inherited preprocessing default, and the result is a caution specific to interpretable legal text-as-data: a step that silently reshapes which features a model sees can distort the very doctrinal and ideological signal such research exists to recover. Leaving stopwords in place is a question of measurement validity.

98. 【2609.19152】FakeSpotter: A content and strategy agnostic Viral Misinformation Detection Tool

链接https://arxiv.org/abs/2609.19152

作者:Giovanni Spitale,Federico Germani

类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:misleading narratives emerge, Misinformation detection tools, limiting their usefulness, rely on binary, binary true

备注

点击查看摘要

Abstract:Misinformation detection tools often rely on binary true and false classifications or models trained on historical examples, limiting their usefulness when novel misleading narratives emerge. Here, we present FakeSpotter, a content- and strategy-agnostic tool designed to estimate the viral misinformation risk of textual content by measuring structural fingerprints of misinformation rather than directly adjudicating truthfulness. FakeSpotter operationalizes a theory-driven framework across linguistic, narrative, logical, and critical-thinking dimensions, using repeated LLM assessments and domain-specific logistic regression classifiers for short and long texts. In a labelled corpus of 764 texts from social media and FakeNewsNet, FakeSpotter achieved macro F1 scores of 0.788 for short texts and 0.793 for long texts on a held-out test set. FakeSpotter's interpretive layer provides explainable outputs through feature-based scores, signal agreement, and a caution index, and can be used for social listening. These findings suggest that identifying the structural fingerprints of misinformation can support early, explainable, and human-supervised assessment of potentially viral misinformation.

99. 【2609.19151】What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

链接https://arxiv.org/abs/2609.19151

作者:Md Jafrin Hossain,Umme Nusrat Jahan,Shouvaggo Sharif Shammo

类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)

关键词:examines user-perceived quality, large-scale research examines, research examines user-perceived, Apple App Store, achieved rapid consumer

备注: Submitted to Array (Elsevier); currently under peer review

点击查看摘要

Abstract:Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store. We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a stratified sample of 300 reviews. Results show that negative sentiment concentrates in advertising (91%), authentication (89%), server reliability (83%), and subscription pricing (73%). Sentiment differs significantly across applications, with Claude exhibiting the highest negative sentiment (47.7%) alongside a strongly enthusiastic user base, indicating statistically significant polarization. These findings are robust despite unequal review counts across applications. As exploratory observations, a subset of DeepSeek reviews raised geopolitical and data privacy concerns related to its Chinese origin, while a proposed Trust Friction Score summarizes application-specific trust and usability barriers into interpretable dimensions. The study provides validated and actionable evidence on user trust, usability, and adoption barriers in consumer generative AI applications.

100. 【2609.19150】Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations

链接https://arxiv.org/abs/2609.19150

作者:Ajit Mallavarapu,Ziwei Gu

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:supervised contrastive data, Large language models, typically requires supervised, requires supervised contrastive, Large language

备注

点击查看摘要

Abstract:Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis (PCA) to the pooled hidden activations, and label the resulting axes automatically from the pole generations. We validate the discovered axes against 245 human-elicited stylistic annotations in a two-phase study. On our strongest model (Qwen-3.5-4B-Instruct), the top two axes match spontaneously requested human dimensions with 72.8% precision and 43.6% macro-recall, and 75.6% of validity ratings judge the axes' polar generations accurate to their labels, with 90.9% adjacent inter-annotator agreement. Discoverability is strongly model-dependent: both Qwen models and Llama-3.2-3B expose human-salient axes, while DeepSeek-7B-Chat drops to 35.3% precision, its leading components dominated by structural rather than stylistic variance. Simple PCA over a model's own decoding variance is thus an effective, low-cost probe of stylistic structure in LLM representations, one that also exposes sharp cross-model differences in how that structure is organized.

101. 【2609.19149】Subliminal Prompting Beyond Static Geometry: Causal Depth and Multi-Token Confounds

链接https://arxiv.org/abs/2609.19149

作者:Barath Velmurugan

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Subliminal learning shows, Subliminal learning, Subliminal, transmit a hidden, language models

备注: 7 pages, 3 figures, 5 tables. Preprint

点击查看摘要

Abstract:Subliminal learning shows that language models can transmit a hidden trait through outputs that appear unrelated to it. One proposed explanation, token entanglement, links animal and number tokens through the model's output vocabulary. Yet existing measurements answer different questions: whether outputs co-vary, fixed output vectors align, an answer can be read from a hidden state, or that state causally controls the answer. We measure each separately in a fixed animal-number prompting protocol. From Llama-3.1-8B to 70B, fixed output-vector similarity predicts behavior less well: the paired mean correlation change is -0.080 (95% CI [-0.127, -0.035]). A fixed output-head readout shows no resolved change in normalized depth AUC. To test control, we copy the temporary answer-position state from one number prompt into another at five depths and measure which prompt the final animal score follows. Donor-control AUC rises from 0.254 to 0.540, a paired change of +0.286 (95% CI [+0.272, +0.300]), with increases for all 18 concepts. The contrast remains with exactly eight transformer blocks remaining, while specificity and identity controls remain small or exact. In two Qwen models, scoring every digit in sequence does not recover the positive one-token association. Per-token averaging instead creates a positive pooled association that disappears after controlling number width, revealing a length confound. Thus, fixed geometry, observational readability, causal timing, and multi-token measurement are distinct properties of this frozen prompting channel. They constrain token-level explanations but do not identify the mechanism of training-time trait transfer.

102. 【2609.19148】Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition

链接https://arxiv.org/abs/2609.19148

作者:Shiyu Luo,Yu Wang,Jiawen Huang,Zhaoxiang Xiao,Chenxi Huang,Qi Zhang,Bin Liu

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:individuals express contradictory, express contradictory signals, Ambivalence and hesitancy, linguistic channels, affective states

备注: 10 pages

点击查看摘要

Abstract:Ambivalence and hesitancy (A/H) are affective states in which individuals express contradictory signals across facial, vocal, and linguistic channels. Automatically recognising A/H in clinical videos requires detecting cross-modal disagreement -- the signal that standard fusion methods suppress. Based on the conflict-aware multimodal fusion framework of Bekhouche et al., we present the Modality Discrepancy Transformer (MDT). MDT enriches the original 6-token design to a 9-token representation comprising three modality embeddings, three absolute-difference features, and three Hadamard-product discrepancy features learned through linear projections. These nine tokens undergo Transformer self-attention, with FiLM-based text-conditioned modulation and LoRA fine-tuning as core architectural components. A text-guided late fusion branch blends a text-only auxiliary head with the full multimodal output at inference. On the BAH dataset from the 3rd ABAW Challenge, MDT achieves 0.7408 Macro F1 on the labelled test split and 0.7368 on the private leaderboard, outperforming the strongest published baseline by over 10 points while training in under 20 minutes on a single GPU.

103. 【2609.20504】Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain

链接https://arxiv.org/abs/2609.20504

作者:Aakash Singh,Lakshmi Pedapudi,Chandrashekar M S,Sanyam Singh,Naga Ganesh,Vineet Singh

类目:Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)

关键词:Digital Green AI-powered, Digital Green, Green AI-powered agricultural, Green AI-powered, AI-powered agricultural advisory

备注: 20 tables, 11 figures, 23 pages

点击查看摘要

Abstract:FarmerChat is Digital Green's AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel for this population, yet field-recorded speech is challenging for general-purpose automatic speech recognition (ASR) because recordings frequently contain machinery noise, background media, competing speakers, and domain-specific agricultural vocabulary. These conditions disproportionately affect crop, pest, chemical, and quantity terms that carry the meaning of a farmer's query. We present a modular, model-agnostic pipeline for improving ASR quality in FarmerChat without fine-tuning or replacing the underlying ASR model. The pipeline combines gated audio enhancement, speaker diarization and target-speaker selection, ASR, domain-aware correction using a weighted agricultural lexicon, and a quality gate for detecting unreliable transcripts. Only the diarization stage is fine-tuned; all other stages use off-the-shelf models behind common interfaces. We evaluate the pipeline on human-annotated FarmerChat recordings in Hindi, Telugu, and Odia using word error rate (WER) and a domain-weighted error rate that gives greater importance to agricultural terminology. The largest improvements occur on multi-speaker recordings, where target-speaker selection prevents competing speech from entering the transcript. Across the full corpus, the pipeline reduces WER by 16-23% relative on three cloud ASR models and by 5% on an on-device model. On multi-speaker recordings, the reductions are 32-42% for the cloud models and 16% for the on-device model. All reported reductions are statistically significant. These results show that targeted preprocessing, speaker selection, and domain-aware post-processing can substantially improve agricultural speech transcription while preserving the underlying ASR model.

Comments:
20 tables, 11 figures, 23 pages

Subjects:

Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Sound (cs.SD)

Cite as:
arXiv:2609.20504 [eess.AS]

(or
arXiv:2609.20504v1 [eess.AS] for this version)

https://doi.org/10.48550/arXiv.2609.20504

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
104. 【2609.19569】Large Language Model Agents for Evidence Based Genetic Disease Severity Classification

链接https://arxiv.org/abs/2609.19569

作者:Tohid Ghasemnejad,Ahmadreza Argha,Mark Grosser,John Wang,Min Yang,Thantrira Porntaveetus,Tony Roscioli,Nigel H. Lovell,Mahmoud Aarabi,Hamid Alinejad-Rokny

类目:Genomics (q-bio.GN); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Human Phenotype Ontology, Disease severity classification, commercial panels vary, panels vary widely, subjective and labor-intensive

备注

点击查看摘要

Abstract:Disease severity classification for genetic conditions is subjective and labor-intensive, creating bottlenecks in genomic screening, where commercial panels vary widely in size and overlap. We developed an autonomous AI agent integrating Reasoning and Acting (ReAct) with Retrieval-Augmented Generation (RAG) to classify 10,211 Human Phenotype Ontology terms. It uses American College of Medical Genetics (ACMG)-endorsed severity guidelines and American College of Obstetricians and Gynecologists (ACOG) quality-of-life criteria to retrieve PubMed literature, generate interpretable reasoning chains, and independently verify claims. At the phenotype level, using expert-curated cohorts, the agent achieved 93.55% accuracy (MCC 0.9237) with 82.6% to 91.4% of claims supported by direct evidence or valid inferences. Gene-level severity was aggregated across 8,738 pairs, identifying 3,283 autosomal recessive pairs with severe or profound presentations. External validation showed 95.2% concordance with Mackenzie's Mission gene list. This system enables standardized panel design by providing reliable, automated classification supported by direct evidence.

信息检索

1. 【2609.20563】Reasoning Quality Matters: Combating Reasoning Collapse in LLM-based Embedding Learning

链接https://arxiv.org/abs/2609.20563

作者:Zihan Gong,Xiaohan Ye,Jiangchao Yao,Jinsong Lan,Xiaoyong Zhu,Xu Chen

类目:Information Retrieval (cs.IR)

关键词:Large Language Models, Large Language, recently shown strong, shown strong potential, producing context-rich text

备注: 30 pages, 8 figures

点击查看摘要

Abstract:Large Language Models (LLMs) have recently shown strong potential for producing context-rich text embeddings for retrieval. Most existing methods either treat embedding learning as passive feature extraction or exploit LLM reasoning through instruction following for better embedding optimization. However, specialization toward embedding objectives can suppress useful reasoning generation or produce retrieval-irrelevant text. We refer to these two forms of degradation as reasoning collapse. To address this issue, we propose CoFree (Collapse-Free Reasoning Embedding), a two-stage framework that progressively integrates LLM reasoning into query and document embedding optimization while preserving reasoning quality. At the first stage, CoFree applies reference-guided supervised fine-tuning to restore the reasoning ability and retain representational strength of the foundation embedding model. At the second stage, we introduce dual rewards, an embedding-oriented reward and a reasoning-oriented reward, to guarantee fine-grained reasoning of the relevance toward the embedding goal in reinforcement learning. This endpoint-coupled optimization transforms embedding learning from static alignment into a high-quality reasoning-guided search process for retrieval. Extensive experiments demonstrate the effectiveness of CoFree, with CoFree-4B achieving an average absolute improvement of 2.8 nDCG@10 points over Qwen3-Embedding-4B across 22 datasets from MTEB and BRIGHT. Online experiments in a real-world retrieval system further show consistent gains. Code, RTED, and model checkpoints will be made publicly available.

2. 【2609.20303】Viveka-Insight: a cross-lingual concept graph and citation-grounded retrieval resource over the complete works of Swami Vivekananda in English and Bengali

链接https://arxiv.org/abs/2609.20303

作者:Tamal Maharaj

类目:Computation and Language (cs.CL); Digital Libraries (cs.DL); Information Retrieval (cs.IR)

关键词:culturally sensitive material, philosophical corpora pose, English Complete Works, Classical philosophical corpora, contemporary readers

备注

点击查看摘要

Abstract:Classical philosophical corpora pose three compounding challenges for language resources: they exist in several languages without parallel alignment, their vocabulary is remote from that of contemporary readers, and generated text over culturally sensitive material must be verifiably grounded. We present Viveka-Insight, a bilingual resource and open-source pipeline for the works of Swami Vivekananda (1863-1902): the nine-volume English Complete Works and the ten-volume Bengali Vani o Rachana, two related but non-parallel corpora of about 15 million characters. Four layers are released: (i) a structure-preserving parse (32,694 paragraphs, 168,842 sentences) with per-paragraph anchors deep-linking into the published editions; (ii) a cross-lingual concept graph of 8,362 language-agnostic concepts with 87,518 relation-typed paragraph-concept and 55,872 concept-concept edges, in which canonical English labels act as a string-equality key linking Bengali and English passages with no parallel data; (iii) a bilingual alias inventory of 60,850 surface forms (30,053 English, 30,797 Bengali); and (iv) a human-annotated set of 200 paragraph-concept edges judged by three annotators, released with all per-annotator judgments. We report known-item cross-lingual retrieval over 194 verified rendered lecture pairs (Recall@10 0.86 in both directions), a 30-question audit of citation integrity and modern-question bridging, and a human study placing concept-extraction precision at 0.60 under strict two-annotator consensus (Cohen's kappa = 0.61). The extractor's confidence weight is calibrated: restricting to weight = 0.8 raises precision to 0.71 while retaining 98% of concept-bearing paragraphs. Precision is markedly lower in Bengali than English (0.54 vs 0.68), locating the weakness in exactly the half that cross-lingual access depends on. The design transfers to other multilingual classical corpora.

3. 【2609.20175】FacetCRS: Multi-Faceted Preference Learning for Pricking Filter Bubbles in Conversational Recommender System

链接https://arxiv.org/abs/2609.20175

作者:Yongsen Zheng,Ziliang Chen,Jinghui Qin,Liang Lin

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)

关键词:filter bubbles, Conversational Recommender System, Recommender System, bubbles, Pricking Filter Bubbles

备注

点击查看摘要

Abstract:The filter bubble is a notorious issue in Recommender Systems (RSs), which describes the phenomenon whereby users are exposed to a limited and narrow range of information or content that reinforces their existing dominant preferences and beliefs. This results in a lack of exposure to diverse and varied content. Many existing works have predominantly examined filter bubbles in static or relatively-static recommendation settings. However, filter bubbles will be continuously intensified over time due to the feedback loop between the user and the system in the real-world online recommendation. To address these issues, we propose a novel paradigm, Multi-Facet Preference Learning for Pricking Filter Bubbles in Conversational Recommender System (FacetCRS), which aims to burst filter bubbles in the conversational recommender system (CRS) through timely user-item interactions via natural language conversations. By considering diverse user preferences and intentions, FacetCRS automatically model user preference into multi-facets, including entity-, word-, context-, and review-facet, to capture diverse and dynamic user preferences to prick filter bubbles in the CRS. It is an end-to-end CRS framework to adaptively learn representations of various levels of preference facet and diverse types of external knowledge. Extensive experiments on two publicly available benchmark datasets demonstrate that our proposed method achieves state-of-the-art performance in mitigating filter bubbles and enhancing recommendation quality in CRS.

4. 【2609.20131】hink Thrice Before Reranking: Multi-perspective Evidence and Reasoning Integration for Text Reranking

链接https://arxiv.org/abs/2609.20131

作者:Lijun Liu,Zhengzong Chen,Wenyan Li,Yuanyuan Zhao,Fei Huang

类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, shown promising improvements, Reasoning-based reranking, Language Models

备注

点击查看摘要

Abstract:Reasoning-based reranking with Large Language Models (LLMs) has shown promising improvements in text ranking. However, current methods predominantly rely on a single reasoning trajectory, resulting in rankings that are susceptible to reasoning errors and inherently constrained in modeling the multifaceted signals underlying document relevance. To resolve this dilemma, we propose MERIT-Rank(Multi-perspective Evidence and Reasoning Integration for Text Reranking), a framework that models complementary reasoning trajectories to improve reranking robustness. MERIT-Rank formulates a Multi-Trajectory Reasoning Space (MTRS) that evaluates query-document relevance from multiple perspectives and introduces a joint reranker that consolidates these reasoning paths into a unified ranking decision. We further develop Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives. Experiments on both reasoning-intensive and traditional retrieval benchmarks show that MERIT-Rank consistently achieves superior performance over competitive baselines. The 4B model notably outperforms most 7B and even 32B rerankers on BRIGHT.

5. 【2609.20050】he Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

链接https://arxiv.org/abs/2609.20050

作者:Zhexi Feng,Ruiyi Zhang,Yongbo Yang,Pengtao Xie

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:retriever ranks highest, coding agent halfway, ranks highest, retriever ranks, coding agent

备注: 32 pages, 3 figures. Benchmark and evaluation resources: [this https URL](https://github.com/LordTARN1SHED/SERBench)

点击查看摘要

Abstract:A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination that supplies the support its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the current decision was annotated to require. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone reaches 66.6%, placing the gain in the set-level policy, not the computation. From frozen repository source with no gold-derived pool, the lead is 5.0 points. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark's own memory agent. Removing one required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.

6. 【2609.19942】Intrinsic Sequence-Likelihood Confidence in Retrieval-Dominated Extractive QA: Two Pre-Specified Negatives, and What They Do and Do Not Attribute

链接https://arxiv.org/abs/2609.19942

作者:Gunwoo Lee,Changmin Sung,Sang-Hwan Gwak,Ina Kim,Ji-Young Choi,Kyong-Ha Lee

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:extractive document question, document question answering, document question, question answering, confidence-driven mechanisms

备注: 26 pages main text + 26 pages supplementary (Online Resource 3). Submitted to Applied Intelligence. Code and data: doi: [https://doi.org/10.5281/zenodo.22710121](https://doi.org/10.5281/zenodo.22710121) , doi: [https://doi.org/10.5281/zenodo.22721044](https://doi.org/10.5281/zenodo.22721044)

点击查看摘要

Abstract:In extractive document question answering whose questions were generated from the passages that contain their answers -- so that retrieval recovers 92-99.8% of what any mode combination could reach, whatever its absolute accuracy -- confidence-driven mechanisms have little to gain. Fine-tuning an open language model on a specialized domain corpus yields a model whose own confidence is a tempting control signal: it could decide which queries warrant further adaptation, and which answers to trust. We evaluate both uses under criteria fixed before the runs were executed, across four 7-9B model families whose adaptation moved closed-book F1 by at most +0.03, and both fail: a distillation trigger on all four families, under its pre-specified three-step transfer budget, and a routing-and-abstention policy in its single-model pilot. Retrieval alone recovers 92-99.8% of best-case combined accuracy under every correctness criterion we test, leaving routers no meaningful gain. The sequence-likelihood signal is insufficient relative to that mode -- area under the receiver operating characteristic curve 0.65-0.81 under the registered criterion -- before adaptation as well as after, unchanged by scalar recalibration and not consistently improved by token-level temperature rescaling. And the finer diagnostics depend on the correctness criterion and on answer length; on the three adapted combinations where we could test it, selector ablations show no statistically detectable downstream benefit from the confidence term on any seed; on Gemma, removing it changes the selector from failing to passing both registered criteria. The usable product is a set of pre-specified negatives with their dependencies made explicit.

7. 【2609.19844】rust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies

链接https://arxiv.org/abs/2609.19844

作者:Hang Xiao,Chuhong Xu,Kainan Zhou,Gangzhen Qian,Lu Yi

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:AI-generated RTL verification, RTL verification plans, AI-generated RTL, RTL verification, verification plans

备注: Cyber-AI

点击查看摘要

Abstract:AI-generated RTL verification plans can satisfy a provider schema yet fail at the boundary to trusted execution. We present SecTB-RTL, an auditable framework covering 31 tasks and 124 authored hardware-security regressions. A deterministic non-AI baseline killed 36, 75, and 78 mutants at increasing resource limits. The first confirmatory run (C1-R2) failed before model execution because the provider rejected its response schema. After a schema-only repair made without viewing outcomes, a separately frozen follow-up run (C1-R3) completed 1,860 calls. The provider accepted 1,857 responses, but only nine passed the production semantic validator. The generation and execution rules did not match. We therefore preserve the run as an instrument-validation incident and report no prompt-effect estimate. This incident shows that provider or schema acceptance does not establish execution validity. Compilation and coverage are only diagnostics; the exact saved artifact must pass the full production path. A subsequent follow-up is excluded because it did not satisfy the preregistered evidence-completeness gate and is treated only as future work. We release the benchmark, failure-preserving contract, incident provenance, and governance controls needed to prevent infrastructure behavior from being misreported as model behavior.

8. 【2609.19831】Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles

链接https://arxiv.org/abs/2609.19831

作者:Noah Mamié,Laurin van den Bergh

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:recommender systems enhanced, incorporating generated natural-language, represent user preferences, generated natural-language user, natural-language user profiles

备注: Accepted at BlackBoxNLP@EMNLP'26 (The 9th BlackboxNLP Workshop Special Track: Reproducibility and Reliability in Interpretability Analyses)

点击查看摘要

Abstract:In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preferences. The original paper explores the synthesis of user profiles from raw user-generated review text across domains such as movies and accommodations (Amazon Movies TV, TripAdvisor). Crucially, these natural-language user profiles enable direct user interaction and intervention, allowing users to customize recommendations by correcting misattributed preferences or addressing cold-start settings. We successfully reproduce the core findings of the original study. Additionally, we extend the evaluation by conducting systematic context ablation experiments, multi-seed stability across five distinct random seeds to establish statistical reliability, and a mechanistic interpretability analysis using the nnsight framework to probe internal model representations under counterfactual profile perturbations. Our findings verify the original paper's claim that User Profile Recommendation (UPR) achieves competitive performance under its test-set reranking protocol and makes recommendations more transparent. Perturbing the natural-language profiles does change predictions, but it shifts predicted ratings uniformly across genres with no detectable genre-selective effect, leaving rankings unchanged even under direct activation steering. We trace this back to the rating-regression objective rather than the profile interface, with ranking-objective models clearly exceeding in this task.

9. 【2609.19787】Dense Feature Representation over Sequence Modeling: A Solution to the KDD Cup 2026 UniRec Challenge

链接https://arxiv.org/abs/2609.19787

作者:Yi Zhang,Weiliang Ji

类目:Information Retrieval (cs.IR)

关键词:Tencent UniRec Challenge, KDD Cup, Tencent UniRec, UniRec Challenge, move held-out AUC

备注: 6 pages, 1 figure, 4 tables. KDD Cup 2026 Tencent UniRec Challenge Workshop

点击查看摘要

Abstract:We describe our 10th-place solution to the KDD Cup 2026 Tencent UniRec Challenge, industrial click-to-conversion (CVR) prediction over 34.82M records, and we ask which mechanisms actually move held-out AUC. Starting from the official PCVRHyFormer baseline, a 15-step single-variable chain raises test AUC from 0.813237 to 0.827816, and our final submission reaches 0.828535. A leave-one-out ablation from the full model attributes the gain: removing the dense-feature representation stack costs 0.0095 AUC and removing the orthogonalized optimizer costs 0.0028, while no sequence-modeling component (merged single-stream backbone, polarity channel, auxiliary head, per-token FFN) costs more than 0.0005, within or adjacent to a $\pm$0.0004 seed band. We also report a generalization hazard: the row-group train/validation split shares one time window, so validation AUC overstates the leaderboard by about 0.014; anti-memorization and high-cardinality-ID changes even invert sign against it, a divergence that traces to dump-to-dump distribution shift and survives a time-ordered re-split. Dense representation and optimization, not finer sequence modeling, drive CVR AUC at this scale, and verdicts must come from the held-out leaderboard.

10. 【2609.19680】FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

链接https://arxiv.org/abs/2609.19680

作者:Yanzhang Ma,Zhenghan Tai,Hanwei Wu,Sizhe Guan,Jianliang Lei,Hailin He,Chaolong Jiang,Jijun Chi,Tung Sum Thomas Kwok,Bohuai Xiao,Jingrui Tian,Xinlu Wu,Xingao Zhan,Peng Lu,Muzhi Li,Yihong Wu,Liheng Ma,Sicheng Lyu,Tianshuo Yan,Junhao Zhu,Yaqian Xu,Lei Ding,Yufei Cui,Ziquan Liu,Boyu Han,Hengli Liu,Ling Zhou,Xinyu Wang

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multiagent Systems (cs.MA); Software Engineering (cs.SE)

关键词:reliability behavior fixed, agent coordination, leaving their reliability, typically improved, reliability behavior

备注

点击查看摘要

Abstract:Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-improvement methods can turn failures into new behaviors, but offer limited control over where a correction should apply or which previously correct answers it may break. We therefore frame post-deployment improvement as controlled behavioral maintenance: recurring failures should become scoped skill patches, and each patch should earn deployment with- out introducing regressions. We instantiate this view in FINSKILLOPS, a multi-agent system for SEC filing QA. FINSKILLOPS derives reusable skills from evidence-grounded, typed failure diagnoses and governs them through targeted validation, protected-case regression checks, negative controls, and versioned replacement or retirement. Across six financial QA benchmarks, a single frozen skill registry achieves the highest verdict-weighted correctness and reference consistency among the evaluated systems. Evolved skills raise correctness from 3.70 to 4.55 on our enhanced benchmark. In a separate 12-round operational study, only six of 33 proposed skills are promoted, while the monitoring non-correct rate falls from 20.0% to 12.5%. These results establish controlled skill scope, admission, and lifecycle management as the foundation for reliable self-improvement.

11. 【2609.19656】Self-Evolving Search Index

链接https://arxiv.org/abs/2609.19656

作者:Sangam Lee,Wonjae Lee,Sunghwan Kim,Deogyong Kim,Jaehoon Kim,Daye Nam,SeongKu Kang,Dongha Lee

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:LLM agents tackle, tackle complex tasks, complex tasks involving, involving diverse information, important as LLM

备注: Work in progress

点击查看摘要

Abstract:Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.

12. 【2609.19622】Beyond Similarity through Zero-Token Geometric Graphs for Multi-Hop RAG

链接https://arxiv.org/abs/2609.19622

作者:Zeliang Li,Xiaofen Xing,Kailing Guo,Xiangmin Xu

类目:Information Retrieval (cs.IR)

关键词:bridge semantic gaps, Large Language Model, Multi-hop retrieval-augmented generation, costly Large Language, retrieval-augmented generation

备注

点击查看摘要

Abstract:Multi-hop retrieval-augmented generation (RAG) requires evidence that remains relevant to a query while introducing enough novelty to bridge semantic gaps. Dense retrieval tends to concentrate on semantically similar documents, whereas graph-based alternatives often depend on costly Large Language Model (LLM) entity extraction and may propagate through noisy connections. We introduce Geometric Gain Graph RAG (G$^3$RAG), a document-only framework whose offline graph construction uses no LLM calls or generated tokens. G$^3$RAG assigns each edge a geometric gain score, $\cos\theta \cdot \sin\theta$, that jointly captures directional consistency and orthogonality between document representations. A density-aware topological penalty suppresses highly connected hubs, while single-step controlled diffusion expands from filtered query seeds toward complementary evidence. We evaluate G$^3$RAG on MusiQue, 2WikiMultiHopQA, and HotpotQA using Nv-embed-v2 and Qwen3-8B-embed. G$^3$RAG obtains the best average F1 and answer-document hit rate among the evaluated graph-based baselines in both embedding settings, with gains of up to 4.26 F1 points in average performance and 5.76 points on MusiQue. It also removes the graph-construction token cost incurred by entity-based graph methods. These results show that geometric structure can support efficient multi-hop evidence discovery without LLM-based graph construction. Code is available at this https URL

13. 【2609.19615】Semantic Layer Induction from Raw Telemetry via Hierarchical LLM and RAG Abstraction

链接https://arxiv.org/abs/2609.19615

作者:Yuanzhe Jia,Ali Anaissi

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Software Engineering (cs.SE)

关键词:Modern applications generate, heterogeneous event streams, generate massive volumes, applications generate massive, actionable business insights

备注

点击查看摘要

Abstract:Modern applications generate massive volumes of raw telemetry data, but translating those noisy, heterogeneous event streams into actionable business insights remains a fundamental challenge. Data engineers and analysts expend substantial effort reconciling semantic discrepancies, hand-crafting parsing logics, and maintaining fragile mappings between raw data and business KPIs. In this paper, we present an end-to-end framework that fully automates the construction of a business semantic layer from application raw logs. Our approach introduces a two-stage semantic abstraction: first, high-level business features are identified via LLM inference augmented with domain-specific industry knowledge; second, fine-grained business nodes are derived through a structured pipeline comprising data refinement, hybrid retrieval, multi-stage filtering, semantic clustering, and canonical naming. Evaluation on production-scale telemetry demonstrates that our system improves human-assessed semantic quality from 50 to 80+ on a 100-point scale, reduces maintenance effort by 80%, filters out 74% of noise, and achieves 0.87 Cohen's kappa via an integrated LLM-as-Judge evaluation, enabling continuous, scalable quality assurance. Overall, our work distinguishes itself from prior work by addressing the novel problem of business semantic layer induction from raw telemetry, operating without labeled training data or manual rule engineering.

14. 【2609.19601】FootprintRAG: Visual Analytics for Evidence Context Refinement in RAG-based Scientific Literature Exploration

链接https://arxiv.org/abs/2609.19601

作者:Xingyu Liu,Yu Dong,Qizhen Yu,Shiyu Cheng,Zhe Wang,Guan Li,Guihua Shan,Dong Tian,Christy Jie Liang,Quang Vinh Nguyen

类目:Graphics (cs.GR); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)

关键词:large language model, ground large language, evidence, language model, evidence context

备注

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) is increasingly used to ground large language model (LLM) outputs in scientific literature. However, in open-ended literature exploration, the evidence context used for generation is often produced through hidden retrieval, reranking, assessment, and filtering steps. Users may receive retrieval summaries without knowing how the system constructed the evidence context, which evidence units were retained or discarded, or whether potentially useful evidence was excluded before synthesis. We present FootprintRAG, an LLM-agent-powered visual analytics system for evidence context refinement in RAG-based scientific literature exploration. The core idea is to treat the RAG evidence context as an explicit, inspectable, and revisable analytical object before generation. FootprintRAG parses scientific literature into text and figure evidence units, expands an initial query into parallel query variants, retrieves and assesses evidence across iterative rounds, and surfaces ERS-ranked supplementary candidates from the corpus-level evidence space. Through coordinated views, the system connects retrieval trajectories, evidence-state revision, and provenance-aware summary generation into a user-steerable workflow. We evaluate FootprintRAG through two case studies, a user study, and a workflow-level comparison with representative RAG systems. The results show that FootprintRAG helps users compare retrieval directions, revise candidate evidence, recover potentially overlooked evidence, and trace generated summaries back to supporting evidence units. FootprintRAG is available at this https URL.

15. 【2609.19585】CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

链接https://arxiv.org/abs/2609.19585

作者:Aiwei Ivy Zhang,Nimra Ishfaq,Mohit Chandra,Santiago Alvarez Lesmes,Adam Coscia,Khatiya Chelidze Moon,Xiaohan Ding,Munmun De Choudhury

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:reasoning over patient, task for clinicians, mental health care, key task, patient journeys

备注

点击查看摘要

Abstract:In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making temporal recovery challenging. We present CliniCIRCA, a multi-stage LLM framework for Calendar-anchored, Imprecision-aware Reconstruction of Clinical Annals. To our knowledge, CliniCIRCA is the first to temporally classify clinical events across unstructured discharge summaries without event-level timestamps. From 14,882 MIMIC-III mental health admissions, we first construct a benchmark of 52 discharge summaries on which CliniCIRCA produces 15,891 temporally tagged events. After correcting 629 errors based on a clinician-in-the-loop evaluation, we produce verified gold-standard labels. Finally, the corrected timelines drive a temporally grounded summarization stage that compresses each source 1.52 times into a date-grouped chronological record. We then scale the framework to generate 1,000 silver-standard timelines and evaluate them as training data. Compared with zero- and few-shot prompting, instruction tuning generally improves five open-weight models on event extraction, temporal tagging, and summarization across silver and clinician-verified evaluations.

16. 【2609.19483】SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features

链接https://arxiv.org/abs/2609.19483

作者:Abdarahmane Traoré,Andy Couturier,Éric Hervet

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:Text-based person retrieval, synthetic training data, Text-based person, frozen text encoder, synthetic training

备注: 16 pages, 4 figures, 3 tables. Accepted at the ECCV 2026 Workshop on AI City Challenge (Track 4). Code and annotations: [this https URL](https://github.com/abtraore/SCOUT-ECCV)

点击查看摘要

Abstract:Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman $\rho = 1.0$); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ($\rho = 0.8$) but not for a linear probe ($\rho = -0.2$). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors' fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: this https URL

17. 【2609.19482】Algebraic Retrieval: Composable Search for Agents

链接https://arxiv.org/abs/2609.19482

作者:Damian Delmas

类目:Information Retrieval (cs.IR)

关键词:compose search strategies, agents compose search, Algebraic Retrieval, compose search, search strategies

备注: 5 pages, 1 figure. Code and reproducible examples: [this https URL](https://github.com/algebraicretrieval/algebraicretrieval)

点击查看摘要

Abstract:Algebraic Retrieval lets AI agents compose search strategies at query time. Relevance criteria, eligibility constraints, and ranking preferences can be expressed together in a mathematical query. The query surface exposes available operations, so an agent can combine them for the question at hand and revise a program after inspecting results. We evaluate execution parity, not agent behavior or retrieval quality. Building on Programmatic Embedding Modulation (PEM), which exposes vector and score arithmetic during retrieval, we demonstrate contrastive scoring, candidate-pool reranking, and weighted ranking as composable queries, alongside executable SQL and PyTerrier counterparts. On the public 11,429-document Vaswani fixture, each program's implementations select the same document set with score differences below 1e-6; one tied pair orders differently across scoring paths.

18. 【2609.19456】Beyond Private Training: The New Landscape of AI Privacy

链接https://arxiv.org/abs/2609.19456

作者:Sean Culatana,Kang Li

类目:Cryptography and Security (cs.CR); Information Retrieval (cs.IR)

关键词:Retrieval-augmented systems increasingly, systems increasingly rely, Retrieval-augmented systems, retain deleted items, systems increasingly

备注

点击查看摘要

Abstract:Retrieval-augmented systems increasingly rely on vector indexes that may retain deleted items in their search graph. Existing deletion interfaces can prevent deleted identifiers from appearing in returned results while still computing distances to their embeddings during graph traversal. We formalize this distinction as output safety versus traversal safety, and introduce TSD-AUDIT, a framework for auditing and enforcing traversal-safe deletion in graph-based approximate nearest-neighbor retrieval. On Faiss IndexHNSWFlat, native filtering leaves the number of distance computations unchanged relative to unfiltered search; at a 70% deletion rate, trace-faithful replay detects deleted-vector scoring in all 100 audited queries. Code inspection of hnswlib's mark_deleted path reveals the same scoring-before-liveness pattern. TSD-AUDIT enforces an alive-before-scoring invariant, repairs connectivity using only live candidates, and emits per-query scored-trace certificates that an independent verifier can check against the deletion snapshot. Under region-targeted deletion, TSD-AUDIT improves Recall@10 over native filtering by 4.3--42.2 percentage points across deletion fractions from 0.5 to 0.9, while remaining comparable under random deletion. These results show that output-only deletion audits can miss process-level exposure: auditing deletion in vector retrieval requires accounting for the vectors scored during search, not only the identifiers returned.

19. 【2609.19244】Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

链接https://arxiv.org/abs/2609.19244

作者:Mahsa Amani,Seungeon Lee,Abhisek Dash,Asmaa El Fraihi,Yunah Jang,Elisabeth Kirsten,Qinyuan Wu,Krishna P. Gummadi,Manish Gupta,Abhilasha Ravichander,Muhammad Bilal Zafar,Soumi Das

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:remains poorly understood, LLM agents increasingly, Web search, Conversational LLM agents, search remains poorly

备注

点击查看摘要

Abstract:Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.

20. 【2609.19158】VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering

链接https://arxiv.org/abs/2609.19158

作者:Yixin Peng,Er Jin,Shiwei Luo,Diego Collarana,Stefan Decker

类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:graph neural network, retrieved knowledge graphs, Knowledge graphs, neural network, network and fusing

备注

点击查看摘要

Abstract:Knowledge graphs are usually integrated into question answering by encoding a retrieved subgraph with a graph neural network and fusing it with the language model in the online inference path. The same subgraph is therefore re-encoded from scratch every time a pair is scored, across training epochs, seeds, and evaluation runs, even though the knowledge graph never changes. We ask whether the retrieved knowledge graphs can instead be compiled once, offline, and then accessed as read-only memory. VisKG-LM shows that it can, by decoupling graph encoding from language reasoning. It serializes each retrieved candidate-specific subgraph as Relation-Labeled Paths and renders the result as an image whose two-dimensional layout preserves the branching structure of the paths. Each image is encoded once, offline, and cached for reuse. At inference, the language model contextualizes the question and candidate from text alone, and only its final layer consults the cached visual memory, reading both its global layout and its local relational detail. The graph information thus enters only after the text has been understood. On the test sets of CommonsenseQA, OpenBookQA, and MedQA-USMLE, VisKG-LMimproves over GreaseLM by $1.2$, $0.8$, and $4.3$ points, respectively, while matching or surpassing GraphVis, a $7$B vision-language model, with only about $400$M online parameters. Against a matched text-only control that receives the identical Relation-Labeled Paths, it gains $4.2$, $6.5$, and $5.1$ points across the three benchmarks. These gains show that the complete visual-memory interface adds value beyond path textualization alone and support compiled visual memory as an alternative to online graph propagation.

21. 【2609.19153】Stop Removing Stopwords: How an Inherited Preprocessing Default Distorts Legal Text-as-Data

链接https://arxiv.org/abs/2609.19153

作者:Gregory M. Dickinson

类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Information Retrieval (cs.IR); Machine Learning (cs.LG)

关键词:scholarship increasingly treats, increasingly treats judicial, Empirical legal scholarship, treats judicial text, legal scholarship increasingly

备注

点击查看摘要

Abstract:Empirical legal scholarship increasingly treats judicial text as data, and much of it still runs on sparse, interpretable pipelines -- TF-IDF features and linear classifiers -- because the textual feature is often the object of study, not merely a means to a prediction. Yet these pipelines inherit a chain of preprocessing defaults from mid-century information retrieval that were never validated against classification accuracy, the most entrenched being stopword removal. This study introduces an exhaustive single-word ablation that measures a preprocessing step's effect directly against the downstream objective, and applies it to stopword removal as the hardest case to dislodge. Matching Supreme Court Database labels to Caselaw Access Project opinion texts, it examines two binary tasks that bracket F1 headroom, ideological direction (no-removal baseline F1 ~ 0.68) and constitutional versus non-constitutional law type (~ 0.92), across 7,668 and 7,001 opinions. For each task the analysis approximates the best stoplist any expert could build, removing each of roughly 18,500 candidate words and measuring the effect directly. Three findings follow: generic stoplists in common use fall below the no-removal baseline in every test; even optimized stoplists are statistically indistinguishable from removing nothing; and meta-models trained on word-level features cannot predict which removals help, so list curation has nothing to target. The method generalizes to any inherited preprocessing default, and the result is a caution specific to interpretable legal text-as-data: a step that silently reshapes which features a model sees can distort the very doctrinal and ideological signal such research exists to recover. Leaving stopwords in place is a question of measurement validity.

22. 【2609.19151】What Users Think of Generative AI: A Cross-Platform NLP Analysis of Trust and Friction in App Store Reviews

链接https://arxiv.org/abs/2609.19151

作者:Md Jafrin Hossain,Umme Nusrat Jahan,Shouvaggo Sharif Shammo

类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)

关键词:examines user-perceived quality, large-scale research examines, research examines user-perceived, Apple App Store, achieved rapid consumer

备注: Submitted to Array (Elsevier); currently under peer review

点击查看摘要

Abstract:Generative AI (GenAI) applications have achieved rapid consumer adoption, yet little large-scale research examines user-perceived quality, trust, and adoption barriers. We present one of the first cross-application analyses of app store reviews for six major GenAI applications (ChatGPT, Gemini, Microsoft Copilot, Claude, DeepSeek, and Perplexity), comprising 17,012 English-language reviews from Google Play and the Apple App Store. We combine BERTopic topic modeling with RoBERTa sentiment classification and evaluate cross-application differences using chi-square, Kruskal-Wallis, and multinomial logistic regression with Bonferroni correction. Both components are validated against human coding using a stratified sample of 300 reviews. Results show that negative sentiment concentrates in advertising (91%), authentication (89%), server reliability (83%), and subscription pricing (73%). Sentiment differs significantly across applications, with Claude exhibiting the highest negative sentiment (47.7%) alongside a strongly enthusiastic user base, indicating statistically significant polarization. These findings are robust despite unequal review counts across applications. As exploratory observations, a subset of DeepSeek reviews raised geopolitical and data privacy concerns related to its Chinese origin, while a proposed Trust Friction Score summarizes application-specific trust and usability barriers into interpretable dimensions. The study provides validated and actionable evidence on user trust, usability, and adoption barriers in consumer generative AI applications.

计算机视觉

1. 【2609.20822】Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

链接https://arxiv.org/abs/2609.20822

作者:Bingxin Xu,Yuzhang Shang,Zhen Dong,Emilio Ferrara

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:http URL, language model writes, evaluate coding agent, URL this paradigm, promising paradigm

备注

点击查看摘要

Abstract:Coding agents have emerged as a promising paradigm for robot manipulation: a language model writes the robot controller as a program, and agents built in this way now operate robots without robot-specific this http URL this paradigm is also safe, however, has not been asked. We evaluate coding agent under a safety constraint, where each task pairs a manipulation goal with an obstacle the robot must not touch. The agent pursues the goal but collides with the obstacle in most cases, treating task completion as its sole objective while neglecting safety. The agent reasons about the obstacle in its traces, and the prompt already forbids touching it, so neither perception nor instruction is at fault; the fault lies in the planning, where the stated constraint never becomes a priority. By decomposing manipulation into a route phase and a contact-rich moment, we locate the source of the failure. Along the route, the model cannot prioritize the safety constraint, having no notion of a clearing route and none of replanning once a chosen route becomes infeasible. At the contact, it is unaware that contact execution is bounded by the same constraint. To close this gap, we present SafeHarness, which equips the model with two obstacle-aware harnesses that enable it to prioritize the safety constraint. Obstacle-aware route planning grounds the objects as bounding boxes and draws candidate routes over them as sequences of waypoints. The agent then plans a route in advance, verifies it, replans when necessary, and only then executes it. Obstacle-aware contact execution instead selects the contact position so that the contact itself avoids the obstacle. SafeHarness attains 71.9% task success and 87.5% collision avoidance, surpassing the previous SOTA by 6.5% and 27.0%, respectively. These results are $2.3\times$ and $1.5\times$ those of the same agent without harnesses.

2. 【2609.20819】Can 4D Foundation Models Remember?

链接https://arxiv.org/abs/2609.20819

作者:Guangzhao He,Hadar Averbuch-Elor,Wei-Chiu Ma

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:world is fundamental, fundamental to navigating, navigating and interacting, reconstruct dynamic environments, Perceiving and remembering

备注: Project Page: [this https URL](https://guangzhaohe.com/persistbench)

点击查看摘要

Abstract:Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360° videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating various models across diverse categories reveals that current models can only maintain short-term consistency that degrades significantly once objects leave the field of view. Our findings highlight the gap between current model capabilities and robust visual memory ("seeing is not remembering"), providing guidance for future development of 4D foundation models. Dataset and code are available on the project page: this https URL.

3. 【2609.20818】SplashSplat: Reconstructing Splashing Liquids from Real-World Multi-View Videos

链接https://arxiv.org/abs/2609.20818

作者:Peiyu Liu,Dingxi Zhang,Federico Tombari,Marc Pollefeys,Christina Tsalicoglou,Daniel Barath

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:sheets tear, ligaments and droplets, appearance is view-dependent, splash lives, tear into ligaments

备注: 18 pages (11 main + 7 supplementary), 14 figures, 12 tables. Project page: [this https URL](https://niko-creater.github.io/splashsplat-web/)

点击查看摘要

Abstract:A splash lives for a fraction of a second: sheets tear into ligaments and droplets, appearance is view-dependent and nearly textureless, and little persists long enough to track. Reconstruction research has consequently focused on smoke, synthetic liquids, or gently deforming surfaces. To our knowledge, no synchronized multi-view dataset of splashing liquids exists. We therefore introduce a benchmark of 20 real scenes, from coherent streams to violent splashes, captured by seven synchronized, calibrated 4K cameras at 60 fps, with manually refined per-view liquid and container masks and fixed evaluation splits. We further present SplashSplat, built on a single principle: impose physical structure only where the observations can constrain it. Per-frame liquid SDFs fused from the masks provide the geometry, level-set transport between consecutive SDFs yields a coarse velocity field, and Lagrangian carriers advected along this flow, corrected against each new observation and reseeded where coverage is lost, decode local Gaussians for differentiable rendering. SplashSplat outperforms state-of-the-art dynamic Gaussian splatting methods on our real captures and on a synthetic benchmark, with physically more plausible motion and a lower training cost. The same representation supports temporal interpolation and style transfer without re-optimization.

4. 【2609.20817】FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

链接https://arxiv.org/abs/2609.20817

作者:Kevin Qu,Tao Sun,Massimiliano Viola,Liyuan Zhu,Zhizhuo Zhou,Sayan Deb Sarkar,Konrad Schindler,Iro Armeni

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:Modeling articulated objects, Modeling articulated, articulated objects, sparse monocular views, Modeling

备注: Project page: [this https URL](https://kevinqu7.github.io/famos)

点击查看摘要

Abstract:Modeling articulated objects from sparse monocular views is challenging because each observation reveals only partial geometry and motion evidence. Most feed-forward methods infer articulation from a single observation and therefore rely heavily on learned category-level shape priors. We present FAMOS, a feed-forward model that predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds. Our model jointly reasons over multiple observations and naturally supports a variable number of inputs, including a single view. To aggregate articulation cues across observations, we introduce a Multi-state Articulation Transformer with alternating state-wise and global attention. We further propose an observed articulation span objective that supervises the motion range each part exhibits across the input observations, encouraging the model to leverage the full observation set. To overcome the limited scale and diversity of existing datasets, we introduce a procedural data generator that synthesizes self-annotated assets during training. Experiments on PartNet-Mobility, ACD, and ArtiCraft-10K demonstrate consistent improvements over both feed-forward and optimization-based baselines. Project page: this https URL

5. 【2609.20816】Paint-Anything: Unified Any-Color Control for Image Generation and Editing

链接https://arxiv.org/abs/2609.20816

作者:Ji Xie,Dewei Zhou,Xinyu Huang,Zhennan Chen,Xun Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Professional design requires, requires any-color control, design requires any-color, Professional design, any-color control

备注: 29 pages, Seed Technical Report

点击查看摘要

Abstract:Professional design requires any-color control: the ability to specify an object's target color with any 24-bit hex value for image generation and editing. Prior work has explored color generation, editing, and colorization, but often relies on dedicated color representations or specialized inference procedures. Advances in large language models offer a simpler starting point: even compact models can associate hex values with color semantics. We present Paint-Anything, which learns a shared hex-prompt interface for generation and editing through object-level color supervision. We develop a data pipeline that constructs Paint-500K from real images through object grounding, perceptual color labeling, and editing-pair synthesis. Since shadows make real-image labels only approximate colors, we complement this supervision with pure-color anchors whose pixels exactly match their paired hex values. These anchors are used only at high-noise timesteps, leaving low-noise training to natural images. We further introduce Any Color Benchmark (ACBench), comprising ACBench-T2I and ACBench-Edit, to measure object-level hex color fidelity across both tasks. On FLUX.2-4B, Paint-Anything improves ACBench-T2I and ACBench-Edit scores by 85.3% and 28.3%, respectively, relative to the base model, with ablations supporting the training recipe. It also achieves the highest average CompColor score among the compared methods.

6. 【2609.20815】ERCPMP-Gx: Endoscopic Image and Video Dataset for Morphological, Histopathological, and Genomic Characterization of Colorectal Polyposis

链接https://arxiv.org/abs/2609.20815

作者:Zahra Ghaffari,Massih Bahar,Mojgan Forootan,Ali Darvishi,Hamidreza Bolhasani

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Hereditary polyposis syndromes, extracolonic tumors, broad spectrum, spectrum of extracolonic, polyposis syndromes

备注

点击查看摘要

Abstract:Hereditary polyposis syndromes can be precursor lesions to colorectal cancer and are associated with a broad spectrum of extracolonic tumors. Early identification and accurate classification of these syndromes are essential for timely diagnosis, individualized patient management, and targeted surveillance strategies for affected families. However, public endoscopic datasets are largely organized around the individual sporadic polyp, and none links the polyposis phenotype to histopathology and germline findings at the patient level. Here, we present ERCPMP-Gx, an endoscopic, histopathological, and genomic dataset developed to support the application of artificial intelligence (AI) in the recognition, characterization, and classification of colorectal polyposis. Most procedures were performed using the Olympus EVIS X1 system with white-light endoscopy (WLE), narrow-band imaging (NBI), magnifying NBI (M-NBI), and NBI with near focus modes, yielding 160 images and accompanying video clips. Approximately eighty percent of cases represent clinically and/or genetically confirmed hereditary polyposis syndromes (PG), including familial adenomatous polyposis (FAP), Peutz-Jeghers syndrome (PJS), juvenile polyposis syndrome (JPS), and ganglioneuroma syndrome (GNS), while the remaining twenty percent comprise non-hereditary polyps and polyp-mimicking lesions with overlapping morphological features (Non-PG), included to support differential classification. Each released record is linked, where available, to standardized endoscopic annotations, representative histopathology, and clinically reported germline findings, forming an AI-ready, patient-level annotation framework. The dataset is publicly accessible at Mendeley (this https URL). For the latest updates and further information, readers are referred to the DataBioX website: this https URL.

7. 【2609.20769】FlowSGS: Improving Flow Matching Priors for Inverse Imaging with Stochastic Interpolants

链接https://arxiv.org/abs/2609.20769

作者:Tianao Li,Xinhui Qian,Emma Alexander

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Split Gibbs Sampling, computational imaging, matching has emerged, prior step, solve inverse problems

备注

点击查看摘要

Abstract:Flow matching has emerged as the state-of-the-art generative model and has been used for plug-and-play (PnP) priors to solve inverse problems in computational imaging. However, existing flow-based inverse solvers assume linear forward models and/or make simplifying approximations in posterior sampling. To circumvent these problems, we introduce FlowSGS, a flow-based posterior sampling method using Split Gibbs Sampling (SGS) to decompose the posterior into a likelihood step and a prior step. Specifically, we sample from the likelihood step using Langevin dynamics and leverage the Stochastic Interpolants (SI) framework to integrate a pretrained flow model into the prior step. We provide a form for the prior step that uses SI's reverse-time SDE, and show connections to previous PnP methods. Moreover, with the aid of the flow prior's straight probability paths and a novel timestep correction technique for the reverse-time SDE, FlowSGS requires fewer network evaluations in its prior step than plug-and-play diffusion samplers. Our experiments show state-of-the-art performance on a range of inverse problems. For the first time, we provide an experiment on a nonlinear inverse problem (Fourier phase retrieval) for flow-based inverse solvers.

8. 【2609.20756】OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

链接https://arxiv.org/abs/2609.20756

作者:Damiano Da Col,Maximilian Igl,Peter Karkus,Kashyap Chitta,Boris Ivanovic,Marco Pavone,Konrad Schindler,Christos Sakaridis

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:yields diminishing returns, scaling pre-training data, diminishing returns, scaling pre-training, yields diminishing

备注: 9 pages, 5 figures

点击查看摘要

Abstract:As scaling pre-training data alone yields diminishing returns, post-training is becoming increasingly important across physical AI domains such as autonomous driving. End-to-end driving policies are pre-trained in open loop with behavior cloning on human demonstrations. However, compounding errors during closed-loop deployment can take the vehicle outside the training data distribution, increasing the risk of safety-critical incidents. Closed-loop post-training can mitigate this risk but requires costly simulation for sensor-based policies. We propose OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes). This teacher then provides supervision to the pre-trained student during closed-loop post-training. We apply OPTED to two camera-based models, TransFuser and VaVAM, and fine-tune them in AlpaSim, using neural reconstructions (3DGS) of real driving logs. Driving scores increase by factors of 1.6$\times$ and 9.5$\times$, respectively. In controlled experiments OPTED matches closed-loop performance with approximately three orders of magnitude fewer simulator interactions than direct RL post-training, while staying closer to the human prior. Project page: this https URL

9. 【2609.20700】Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

链接https://arxiv.org/abs/2609.20700

作者:Lili Wang,Jing Li,Xiaowen Sun,Xiangyu Hu,Zhuangzhuang Gu,Jian Liu,Srihari Nelakuditi,Yan Tong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Episodic test-time adaptation, fixed step count, test-time adaptation resets, Episodic test-time, source weights

备注: 45 pages, 8 figures, 26 tables

点击查看摘要

Abstract:Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean $\Delta$Dice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation---the disagreement geometry between $M_0$ and the adapted mask $M_k$---predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman $\rho$ 0.50--0.60), at a quarter of gradient-norm's latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net$\to$SegFormer, Cityscapes$\to$ACDC) with coordinate, thresholds and per-bucket actions instantiated per domain.

10. 【2609.20680】owards Scaling Marine Perception with Synthetic Data

链接https://arxiv.org/abs/2609.20680

作者:Haoyu Ma,Onur Bagoren,Anja Sheppard,Elias Fandi,Ashrith Edukulla,Tanner Aslan,Natasha Sieh,Jingyu Song,Katherine A. Skinner

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Scalable machine learning, Scalable machine, challenging underwater environments, machine learning, environments is strongly

备注: Accepted at OCEANS 2026 Monterrey

点击查看摘要

Abstract:Scalable machine learning in challenging underwater environments is strongly limited by the lack of labeled real-world training data. This data is often expensive and laborious to gather, making large-scale real-world data challenging to gather and curate. However, simulated data can help close the gap, enabling many learning-based tasks for underwater perception. In this work, we extend OceanSim, an IsaacSim-based underwater perception simulator, with a Synthetic Data Generation (SDG) pipeline for training models to be used in underwater scenarios. The proposed pipeline enables users to generate large, automatically labeled, photorealistic datasets with configurable scene appearance, structure, and sensor settings. We evaluate the pipeline on a real-world sea urchin detection task and study how different forms of synthetic scene variation affect sim-to-real performance. Based on these experiments, we discuss findings on our results, main limitations of the current pipeline and identify future directions for improving underwater rendering fidelity, scene diversity, and the evaluation of sim-to-real generalization. The open-source code can be found at this https URL.

11. 【2609.20673】FunArt: Decoding Functional Structure and Articulation from Generative 3D Latents

链接https://arxiv.org/abs/2609.20673

作者:Dennis Rotondi,Abdelrhman Werby,Kai O. Arras

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:identify articulated objects, human environments, robots must identify, operate effectively, effectively in human

备注

点击查看摘要

Abstract:To operate effectively in human environments, robots must identify articulated objects, segment their movable and interactive parts, and estimate their kinematic models. Existing articulated scene representations typically recover kinematics from observed interactions, while methods operating on static scans often decouple articulation from functional interactive elements. We present FunArt, a framework that constructs articulation-aware functional 3D scene graphs from posed RGB-D observations captured in a single static configuration. FunArt reconstructs object instances, converts their fused geometry directly into the O-Voxel representation of TRELLIS.2, and exploits its frozen, sparse-compression VAE as a structural prior. A lightweight query-based decoder combines compact object-level latents with dense, surface-aligned features to jointly segment movable parts and functional interactive elements while estimating motion type, axis, origin, and range. On the Articulate3D dataset, FunArt achieves state-of-the-art performance across movable-part segmentation, articulation estimation, and functional-element segmentation, both with and without ground-truth object input. In the end-to-end setting, it outperforms the strongest baselines by 1.5 AP_{50} points for movable parts, 2.8 AP_{50} points under joint origin-and-axis constraints, and 6.7 AP_{50} points for functional elements. These results demonstrate that generative 3D latents encode actionable structural cues that can initialize robotic perception and planning before physical interaction.

12. 【2609.20669】Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

链接https://arxiv.org/abs/2609.20669

作者:Zhongbo Zhang,Zaibin Zhang,Yifan Wang,Changbo Yan,Lijun Wang,Huchuan Lu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:generating geometrically grounded, successful manipulation requires, geometrically grounded actions, Movement Trend Guidance, strong at generating

备注

点击查看摘要

Abstract:3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.

13. 【2609.20662】Earth Surface Immune System for Rapid Monitoring of Unknown Anomalies

链接https://arxiv.org/abs/2609.20662

作者:Jingtao Li,Qian Zhu,Xinyu Wang,Deren Li,Liangpei Zhang,Yanfei Zhong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:expanding human activities, remote sensing targets, limited historical data, conventional remote sensing, escalating climate change

备注: 51 pages

点击查看摘要

Abstract:Earth surface anomalies, driven by escalating climate change, and expanding human activities, are increasing in both frequency and diversity, yet their limited historical data and unpredictability make them fundamentally different from conventional remote sensing targets. Existing methods address specific anomaly categories or stop at localization, leaving a gap between detection and actionable information. Here we present ESIA, an Earth Surface Immune System whose architecture is constrained by three principles from the biological immune system, refined over millions of years against equally diverse and uncertain threats. A non-specific innate immune stage treats anomalies as unobserved changes in time-series satellite imagery, generating binary localization maps at 14.51 km2/s without assuming any anomaly category, surpassing the strongest general baseline by 37% in F1. A specific adaptive immune stage applies negative selection to filter text prompts and matches surviving prompts with localized image patches through a multi-modal foundation model, enabling open-vocabulary recognition of unknown anomaly attributes including category, affected area, and damage severity, with recognition F1 exceeding 80%. A mutation mechanism tunes minimal embeddings at test time, adapting to each scene in 3.26s using a single reference image pair. We validate ESIA on a global-scale dataset covering 19,801.60 km2 across six anomaly categories, comparing against 22 models, and further apply it to quantify degraded farmland in the Dnipro Delta following the Kakhovka Dam collapse and assess burn severity from 2025 Palisades Fire in Los Angeles. This unprecedented flexibility in handling unknown anomalies opens new avenues for real-time disaster response and environmental surveillance.

14. 【2609.20649】DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

链接https://arxiv.org/abs/2609.20649

作者:Yan Qin,Yue Chen,Wenwei Lin,Shujia Liu,Chuqiao Lyu,Kailun Su,Chenze Yu,Ping Luo,Wenbo Ding,Tianxing Chen,Renjing Xu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:manipulation requires dense, requires dense tactile, embodiment-specific sensors, Learning predictive models, requires dense

备注: Accept to IROS 2026 Workshop RoBoWoMo (Lightning Talk)

点击查看摘要

Abstract:Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict future RGB observations and bilateral tactile dynamics. Our insight is that human and robot manipulation share transferable contact dynamics when their tactile observations and action spaces are made compatible. We deploy flexible piezoresistive arrays with a shared sensing layout on both human and dexterous robot hands, and retarget human motion into the robot action space so that human interaction can supervise the same dynamics model used for real-robot prediction. DexTouch-WM couples a pretrained video expert with a lightweight tactile expert using anatomy-aware tactile tokens and aligned action conditioning. In human-to-robot scaling experiments, we keep five hours of real-robot supervision fixed while increasing human interaction from 0 to 100 hours, and observe substantial improvements in held-out robot-domain visual, geometric, and contact prediction despite disjoint human and robot task sets. Beyond prediction, we evaluate the world models as surrogate environments for policy evaluation and as generators of synthetic trajectories for real-robot policy learning, showing that scalable human interaction provides a complementary data axis for learning dexterous robot world models.

15. 【2609.20638】PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos

链接https://arxiv.org/abs/2609.20638

作者:Di Wen,Kailun Yang,Jimmy Weissert,Luc Maria Scherrer,Cedric Zöllner,Ruiping Liu,Yufan Chen,Jiale Wei,Junwei Zheng,Kunyu Peng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:assistant watching egocentric, watching egocentric video, person recovers, assistant watching, watching egocentric

备注: 9 pages, 2 figures, 4 tables. Code: [this https URL](https://github.com/Kratos-Wen/PROVIA)

点击查看摘要

Abstract:An assistant watching egocentric video should notice a mistake from past frames alone, before the next step begins, and keep working once the person recovers. A mistake changes the state of the work, so every later step has to be read against what was done rather than against the plan. The first-mistake protocol that current online methods report on cuts each recording at its first mistake, so a fixed-time rule that never looks at the video is right on every case. We evaluate on complete trials, where mistakes and recoveries arise naturally, under a validation false-alarm budget and against controls that use timing alone. PROVIA keeps two records apart: a factual state, a learned summary of the steps each actor performed, mistakes included, and the accepted progress, an exact posterior over the state of an automaton induced from correct demonstrations by Bayesian state merging and over the execution status of each actor. Procedure-state transitions occur only in the correct-status branch; the mistake and correction branches retain the source state. A sequential test turns the per-frame mistake probability into alarms. With one filter and one optimization rule, PROVIA ranks mistakes best among the evaluated controlled baselines on CaptainCook4D, IndustReal, HoloAssist and IMPACT-ego. At a validation budget of 0.1 false alarms per minute it recalls .154 against .128 on CaptainCook4D and .034 against .015 on HoloAssist, where it leads at every budget. The pipeline runs at 58-70 frames per second. The source code is available at this https URL.

16. 【2609.20633】Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

链接https://arxiv.org/abs/2609.20633

作者:Yulong Chen,Ziqian Zhang,Haoyu Zhang,Ao He,Senmao Li,Kai Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Text-guided image editing, Text-guided image, preserving unrelated source, Generative Refinement Network, preserving unrelated

备注

点击查看摘要

Abstract:Text-guided image editing must introduce the requested changes while preserving unrelated source content. Diffusion-based editors rely on spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed decoding order limits revision of earlier decisions. We introduce RefineEdit, a training-free prompt-to-prompt image editing framework built on a Generative Refinement Network. Our key idea is to couple edit localization with content generation through the global refinement of binary image codes, allowing editing evidence to be reassessed as the image evolves. RefineEdit initializes an editing branch from an intermediate source state, reusing the emerging layout. We compare the probabilities assigned by the two branches to the same source-sampled bits, using their signed differences to select editable positions and bits. Selected bits follow editing refinement, while the remaining bits copy the evolving source state. To stabilize editing across refinement steps, adaptive spatial freezing limits unnecessary mask expansion, while finite bit locking keeps recently selected bits editable. The framework requires no additional training, external masks, or attention control. Across nine editing categories of PIE-Bench, RefineEdit achieves the best background-preservation scores in PSNR, LPIPS, MSE and SSIM, together with the highest whole-image and edited-region CLIP scores among the evaluated methods.

17. 【2609.20623】PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions

链接https://arxiv.org/abs/2609.20623

作者:Rinto Yagawa,Han Cheng,Dieter Schmalstieg,Hideo Saito,Shohei Mori

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:introducing severe spatial, severe spatial redundancy, Recent single-view feed-forward, Gaussian Splatting, Recent single-view

备注

点击查看摘要

Abstract:Recent single-view feed-forward 3D Gaussian Splatting (3DGS) generation predicts a fixed number of Gaussians per camera ray, introducing severe spatial redundancy. Most existing compaction strategies target multi-view setups to exploit cross-view consistency and are incompatible with single-image models. Instead of retraining the base feed-forward network to directly output compact representations, our insight is to keep the base models frozen and apply post-hoc pruning and recurrent refinement to the generated Gaussians. Consequently, we propose a backbone-agnostic compaction pipeline for single-view feed-forward 3DGS that couples an importance-score-based pruning mechanism with a trainable, lightweight recurrent refinement module, which iteratively updates the surviving primitives to restore image quality. Our results demonstrate seamless integration with existing baselines while preserving novel-view rendering fidelity and achieving high memory reduction. Furthermore, our method supports flexible inference-time keep ratios for application needs.

18. 【2609.20615】INSPECT: Learning Robot View Selection from Assistant Use

链接https://arxiv.org/abs/2609.20615

作者:Di Wen,Kailun Yang,Wenhao Guo,Yitian Shi,Junwei Zheng,Yufan Chen,Ruiping Liu,Jiale Wei,Rania Rayyes,Kunyu Peng

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:correctly installed, egocentric assembly assistance, Robots inspecting, view, Abstract

备注: 9 pages, 3 figures, 5 tables. Code: [this https URL](https://github.com/Kratos-Wen/INSPECT)

点击查看摘要

Abstract:Robots inspecting an assembly must determine which parts are present and whether they are correctly installed. During egocentric assembly assistance, head motion and workpiece handling reveal evidence for these checks, while spoken state confirmations link observations to procedural outcomes. We introduce INSPECT, which learns robot view preferences from records of a smart-glasses assistant that answers part queries and provides next-step guidance. Presence-Invariant TwinSwap (PI-TwinSwap) calibrates object evidence through paired identity interventions. Claim-indexed supervision separates evidence requirements from camera-reproducible observation changes. Object-centered calibration adapts relative view preferences to robot poses, while clause-level screening checks predicted evidence. The robot selects views using only its current observation and known poses, without candidate images. Evaluation uses annotated assistant-video replay to simulate state feedback, without target-domain view labels for policy training. On images of physical gearbox assemblies, INSPECT achieves the highest view utility among the compared non-oracle policies and raises human-rated full verifiability from 34.8% to 41.7% compared with keeping the current view. On commercial angle-grinder recordings in IMPACT, the transferred relative-view selector increases the correct decision rate from 50.6% to 54.3% with a frozen perception head. The source code is available at this https URL.

19. 【2609.20589】RawSLAM: Online HDR Gaussian SLAM from Linear Radiance

链接https://arxiv.org/abs/2609.20589

作者:Marina Orozco González,Luis Merino

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Low Dynamic Range, tonemapped Low Dynamic, Current dense visual, High Dynamic Range, Dynamic Range

备注

点击查看摘要

Abstract:Current dense visual SLAM systems rely almost exclusively on 8-bit tonemapped Low Dynamic Range (LDR) inputs, limiting their robustness in extreme lighting where shadows and highlights trigger tracking drift and mapping collapse. Conversely, existing raw and High Dynamic Range (HDR) reconstruction pipelines operate strictly offline. They depend on Structure-from-Motion preprocessing and are not suited for large inter-frame motion. We present, to the best of our knowledge, the first online Gaussian SLAM framework that tracks and maps directly on single-exposure 16-bit linear HDR imagery. Our method rests on three core components: an architecture-agnostic HDR Gaussian Splatting module featuring an MLP-free logarithmic parameterization of Gaussian color features; a Reinhard range-compressed photometric objective; and structure-guided spatial gradient weighting. Combined, these components allow our approach to outperform a direct HDR adaptation of MonoGS in both trajectory and reconstruction accuracy, while rendering natively in linear scene radiance for post-rendering processing. The same formulation runs unchanged on standard 8-bit inputs, roughly halving the MonoGS baseline error. Furthermore, our HDR Gaussian module transfers seamlessly to SplaTAM, Gaussian SLAM, and DROID-W, eliminating all tracking failures these systems suffer on challenging illumination sequences. To enable this research, we introduce RawSLAM: a dataset of 10 real-world indoor sequences featuring 16-bit RAW imagery, aligned depth, IMU measurements, and external OptiTrack poses. Code and dataset will be made publicly available soon.

20. 【2609.20586】CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

链接https://arxiv.org/abs/2609.20586

作者:Zhikun Zhou,Kunyu Peng,Runyi Yang,Junhao Cai,Di Wen,Ruiping Liu,Danda Pani Paudel,Yi Zhou,Luc Van Gool,Kailun Yang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:relation-centric language queries, requires grounding object, cooperative referring Gaussian, embodied robots requires, understanding for embodied

备注: The established benchmark and source code will be publicly released at [this https URL](https://github.com/ruojiruoli17/CoRef-GS.git)

点击查看摘要

Abstract:Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent's observations, while spatial relations must still be interpreted from the querying robot's viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58° after coarse initialization to 0.15° after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at this https URL.

21. 【2609.20574】DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering

链接https://arxiv.org/abs/2609.20574

作者:Luca De Grandis(1),Silvia Cappelletti(1),William Raccagni(1 and 2),Marcella Cornia(1),Lorenzo Baraldi(1),Rita Cucchiara(1) ((1) University of Modena and Reggio Emilia, Modena, Italy, (2) University of Pisa, Pisa, Italy)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:costly manual effort, visual question answering, question answering remains, requires costly manual, document visual question

备注

点击查看摘要

Abstract:Answer grounding in document visual question answering remains an open challenge: most benchmarks lack grounding annotations or provide limited-quality labels, while constructing grounded datasets still requires costly manual effort. We introduce DocAttriBench (DAB), a large-scale benchmark for fine-grained, element-level source attribution in Document VQA, grounding answers to specific layout elements such as text blocks, tables, and images. To build DAB, we propose a Mask-based Perplexity-Derived Attribution method (MAPPET) that combines document layout and language modeling to identify the most informative element for each answer. MAPPET measures the increase in perplexity after masking candidate elements and attributes the answer to the element contributing most to model confidence. Applying MAPPET to multiple existing Document VQA datasets yields DAB, with 237k documents and 296k question-answer pairs with element-level grounding. We benchmark grounding-capable multimodal LLMs on DAB, evaluating answer accuracy, attribution accuracy, and overall answer quality. Results show that while larger models generally achieve higher answer accuracy, even the strongest models often fail to localize the supporting elements. DAB provides a scalable benchmark for developing grounded, verifiable, and trustworthy Document VQA models. Dataset and code are available at this https URL.

22. 【2609.20566】OmniMimic: Dynamics-completed Motion Augmentation for Multi-style Omnidirectional Quadruped Locomotion

链接https://arxiv.org/abs/2609.20566

作者:Sheng Wu,Guoqiang Zhao,Zhe Yang,Fei Teng,Zhikun Zhou,Yanlin Yang,Zheng Fang,Hong Zheng,Yaonan Wang,Kailun Yang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:provide quadruped robots, demonstrations provide quadruped, Animal demonstrations provide, distinctive gait styles, hand-crafted rewards

备注: The project page is at [this https URL](https://OmniMimic.github.io)

点击查看摘要

Abstract:Animal demonstrations provide quadruped robots with natural and distinctive gait styles that are difficult to specify through hand-crafted rewards. However, their narrow directional coverage leaves little style-consistent supervision for backward, lateral, and turning commands. We present OmniMimic, a training framework that turns directionally limited animal demonstrations into a single multi-gait policy over target per-axis velocity ranges. OmniMimic first combines temporal reversal, constrained dynamics completion, and sagittal reflection to construct robot-specific kinematic and physical supervision beyond the observed directions. It then expands commands progressively from the demonstrated velocity distribution toward the target per-axis bounds, and uses a shared actor with soft-gated, gait-specialized residual experts to balance reusable locomotion skills with gait-specific corrections. Across four gaits in simulation, OmniMimic reduces mean foot-position RMSE at forward and backward reference velocities by 12.9% and velocity-tracking RMSE on a uniform Cartesian command grid by 63.1%, compared with the matched APEX baseline. The project page is at this https URL.

23. 【2609.20562】A Dual-Stream Regulated Reconstruction and Segmentation Network with Hierarchical Artifact-Prior Modeling for Ultra-Low-Field Pediatric Neuroimaging

链接https://arxiv.org/abs/2609.20562

作者:Bahram Jafrasteh,Leo Milecki,Qingyu Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Automated quality assessment, pediatric MRI, Automated quality, weak anatomical boundaries, MRI are limited

备注

点击查看摘要

Abstract:Automated quality assessment, enhancement, and segmentation of multiple structures in $0.064\,\mathrm{T}$ ultra-low-field pediatric MRI are limited by a low signal-to-noise ratio, weak anatomical boundaries, and frequent artifacts. We present a unified framework for the LISA 2026 Challenge that performs all three tasks together within one inference pipeline. A network with two coupled streams, built on a 3D U-Net, first reconstructs an enhanced uLF volume and then combines the original and enhanced images for subcortical segmentation. To improve boundary stability, we add an auxiliary class covering brain tissue outside the target structures, derived from whole brain masks. A head conditioned on an artifact graph predicts the seven artifact ratings from reconstruction residuals and frozen segmentation features. We address the scarcity of dense annotations using diffeomorphic registration from atlas to target for label propagation and to regularize anatomical reconstruction. We report validation results across all three tasks.

24. 【2609.20509】Automated Goldsmith's Mark Retrieval in Silverware

链接https://arxiv.org/abs/2609.20509

作者:Atmik Tiwari,Vincent Christlein,Mark Fichtner,Freya Gohlke,Birgit Schübel,Theresa Witting,Heike Zech,Mathias Zinnen

类目:Computer Vision and Pattern Recognition (cs.CV); Digital Libraries (cs.DL)

关键词:goldsmith marks play, art historians, dating of artifacts, play a critical, critical role

备注: Accepted at the VISART workshop, ECCV 2026. 18 pages, 8 figures, 1 table

点击查看摘要

Abstract:For art historians, goldsmith marks play a critical role in the identification and dating of artifacts. In practice, experts must manually compare a query mark against hundreds of documented examples, a process that is both tedious and highly dependent on specialist knowledge. To address this, we present an AI-assisted retrieval pipeline that combines mark localization with metric-learning fine-tuning across three backbone architectures: an ImageNet-pretrained ResNet-50, a supervised ViT-S/16, and a self-supervised DINOv2 ViT-S/14. We conduct a systematic evaluation of cropping strategies, where we measure the impact of no cropping, manual ground-truth cropping, and learned detection-based cropping, and assess their interaction with each backbone. Our strongest configuration, DINOv2 ViT-S/14 with manual crop and metric-learning fine-tuning, achieves an mAP of 62.63% and a Top-1 accuracy of 73.74%. Our experiments show that self-supervised pretraining and mark localization are the two most impactful factors, with learned cropping recovering the majority of the gain from manual cropping without requiring ground-truth annotations at inference time. To enable reproducibility and adoption in the digital humanities, we release our manually annotated dataset and codebase, and deploy the system via a public web interface.

25. 【2609.20508】Grounded Product Understanding in Livestream Videos

链接https://arxiv.org/abs/2609.20508

作者:Xinyu Zhang,Junjie Chen,Jiawei Ge,Qianlong Li,Libin Ma,Baokun Pan,Yahui Luo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:E-commerce livestreams, product understanding, product, online consumers, important channel

备注

点击查看摘要

Abstract:E-commerce livestreams have emerged as an important channel for presenting products to online consumers, containing multiple products whose information is scattered in different moments. This poses significant challenges for downstream product understanding applications, such as product-centric livestream clipping, where models need to identify the product and its relevant segments for information gathering. However, existing benchmarks for general product understanding typically evaluate product retrieval and temporal localization in isolation, leaving the critical correspondence between product identity and temporal evidence largely unassessed. To address this limitation, we introduce GPUB, a large-scale benchmark comprising 3,000 livestream instances with quality-controlled multi-moment temporal annotations and a catalog of over 31K fashion products. GPUB supports three evaluation tasks: the main task Grounded Product Understanding (GPrU) requires jointly identifying the target product and localizing its supporting moments from a livestream video and a candidate product set; Product Retrieval and Product Moment Localization serve as two complementary subtasks. Evaluation of existing multimodal models shows that GPrU remains highly challenging, with the best-performing baseline achieving only 10.13% Pair mAP@.3. To narrow the performance gap, we further develop UniPro, a unified product understanding model that derives product-aligned and temporally structured representations from shared multimodal encoding, improving Pair mAP@.3 to 21.53% while achieving 37.23% Joint R@1@.3 on GPrU.

26. 【2609.20475】SenseFuse: Label-Free Fusion of Image and Shape Encoders for Open-Vocabulary 3D Instance Segmentation

链接https://arxiv.org/abs/2609.20475

作者:Euiseok Han,Tri Ton,Hwanhee Kim,Seungyeon Ryu,Chang D. Yoo

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:fundamental for robotics, laying the groundwork, object manipulation, Open-vocabulary scene understanding, understanding is fundamental

备注: 8 pages, 6 figures. Code: [this https URL](https://github.com/hanes1207/SenseFuse)

点击查看摘要

Abstract:Open-vocabulary scene understanding is fundamental for robotics, laying the groundwork for spatial reasoning and object manipulation. While closed-vocabulary 3D instance segmentation heavily leverages 3D shape information, state-of-the-art open-vocabulary methods remain predominantly restricted to 2D image features or image-distilled representations during mask labeling. In this paper, we propose SenseFuse, a label-free fusion method that balances 2D image and 3D shape encoders for robust open-vocabulary 3D instance segmentation, refining only the mask-labeling stage of existing pipelines. We reveal that 2D image and 3D shape encoders exhibit largely disjoint failure patterns and rarely share identical wrong labels, whereas two 2D image encoders frequently repeat the same errors. This distinct behavior makes the 2D and 3D pair inherently complementary. We introduce an adaptive mechanism that selects a scene-level fusion weight to maximize a label-free sensitivity measure, estimated directly from a single scene's unlabeled proposals in milliseconds. SenseFuse improves labeling accuracy in every evaluated setting across ScanNet200, Replica, and ScanNet++, recovering 67-100% (median 93%) of the gain achievable with an oracle weight, and it raises instance AP in 21 of 22 reported settings. Code is available at this https URL.

27. 【2609.20441】Cross-Architecture Foundation-Model Distillation for Edge Flood Segmentation

链接https://arxiv.org/abs/2609.20441

作者:Fabian Schmalstieg,Karsten Mueller,Wojciech Samek

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:provide strong flood-segmentation, size limits deployment, Geospatial foundation models, Geospatial foundation, strong flood-segmentation performance

备注: Main paper (17 pages) with supplementary material (11 pages). Submitted to IEEE JSTARS, Special Section on Generalist-Specialist Model Synergy for Remote Sensing: Theories, Methods, and Applications

点击查看摘要

Abstract:Geospatial foundation models can provide strong flood-segmentation performance, but their size limits deployment on memory-constrained edge hardware. We distill a 300-million-parameter Prithvi-EO-2.0 teacher, fine-tuned on the 252 manually labeled Sen1Floods11 training scenes, into a 0.7-million-parameter EfficientViT-B0 student. The teacher supervises additional unlabeled Sentinel-2 imagery, allowing the student training set to grow without new manual annotations. At the matched budget of 252 scenes, teacher-supervised training is competitive with direct training and improves STURM-Flood performance across tested configurations; a geometry-matched control shows that label source alone does not explain the difference. Scaling the teacher-supervised pool to 2,500 scenes narrows the remaining student--teacher gap: the float student reaches 0.787 water intersection over union on the Sen1Floods11 test split against 0.822 for the teacher, matches the teacher on STURM-Flood under our evaluation protocol, and remains below it on WorldFloods-v2. After activation replacement and quantization-aware training, the student runs as a 1.5-megabyte 8-bit integer (INT8) TensorRT engine on a Jetson Xavier NX at 5.57 milliseconds of graphics processing unit (GPU) compute per 512-by-512 image, with approximately 14 megabytes of runtime device memory. A fixed modified normalized difference water index (MNDWI) threshold is competitive with both models on the two clean external benchmarks, so we interpret those benchmarks as generalization tests rather than as evidence of learned-model superiority over a spectral rule. The results support the conclusion: foundation-model supervision can amplify a fixed manual annotation budget into a substantially larger training set and yield a compact, deployable edge model.

28. 【2609.20427】When Do Language-Grounded Explanations Help? A Graph-Bottleneck for Farm Monitoring Interpretable Sheep Facial Pain

链接https://arxiv.org/abs/2609.20427

作者:Alam Noor,Miguel Guti'errez Gait'an

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Automated pain recognition, Facial Expression Scale, make continuous welfare, continuous welfare assessment, welfare assessment practical

备注

点击查看摘要

Abstract:Automated pain recognition from facial expression could make continuous welfare assessment practical in sheep, but adoption depends on trust: a stockperson cannot act on a score that arrives without justification. We ground a model in the Sheep Pain Facial Expression Scale (SPFES) by letting each detected facial region attend over text embeddings of the clinical descriptors and then test whether the resulting explanations mean anything. They do not. Ablating an entire descriptor changes the predicted logit by about $10^{-4}$, and the most-attended cue agrees with the predicted pain level in only $32.6\%$ of regions, although the attention maps, the learned gate, and the generated text all proposed otherwise. We therefore remove the appearance bypass with a concept bottleneck whose classifier reads only SPFES concept scores, supervised by per-region state annotations that image-level pipelines discard. This costs $0.05$--$0.10$ in Cohen's $\kappa$ but yields concepts that are demonstrably learned: minority pain-indicating states are recovered at $3.5$--$8.3\times$ their base rates, and the ear and eye severity orderings emerge without severity supervision. Removing the supervision alone leaves $\kappa$ unchanged while concept accuracy falls to $0.109$, showing that architectural necessity does not imply semantic validity. We also show that pooled concept accuracy is misleading under clinical imbalance and provide a cross-validated, protocol-matched benchmark of seven methods on this dataset.

29. 【2609.20423】WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

链接https://arxiv.org/abs/2609.20423

作者:Hao Yu,Kang Liu,Linnan Zhao,Jiabo Zhan,Chong Sun,Chen Li,Jing Lyu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:requires reliable performance, converts document images, Document parsing converts, parsing converts document, acquisition conditions

备注

点击查看摘要

Abstract:Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.

30. 【2609.20414】ouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

链接https://arxiv.org/abs/2609.20414

作者:Danyan Zhou,Jinxuan Lu,Jiawei Lin,Tianxing Chen,Chuqiao Lyu,Wenbo Ding

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:dexterous robotic manipulation, enabling dexterous robotic, understanding physical interactions, robotic manipulation, provide direct contact

备注

点击查看摘要

Abstract:Tactile signals provide direct contact and force measurements that are essential for understanding physical interactions and enabling dexterous robotic manipulation. However, tactile sensing requires direct measurement at contact interfaces, making large-scale data collection reliant on intrusive, costly, and restrictive instrumentation. We present TouchSight, a monocular egocentric vision framework for dense full-hand contact force prediction that leverages 500 hours of pressure-glove recordings and extensive hand-object interaction (HOI) data. To address the appearance gap between gloved training data and bare-hand real-world scenarios, we construct TwinTouch-20H: 20 hours of paired visual data in which generative video models re-render gloved recordings as bare-hand observations against new backgrounds while preserving the original measured tactile labels. TouchSight predicts dense force from both gloved and generated bare-hand videos, outperforms prior contact prediction methods on OakInk2, qualitatively generalizes to natural bare-hand egocentric videos from unseen datasets, and improves consistently as glove supervision scales. These results demonstrate that dense tactile signals can be recovered from egocentric vision alone, without tactile instrumentation at capture time.

31. 【2609.20388】Navi-Agent: Unlocalized Monocular Navigation Agent

链接https://arxiv.org/abs/2609.20388

作者:Wenyuan Xie,Mengyang Hong,Yongzhong Wang,Yanbiao Ji,Yijin Zhou,Shaokai Wu,Shalayiding Sirejiding,Huayi Zhou,Yi-Chao Chen,Ma Ling,Yue Ding,Hongtao Lu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Continuous Environments, unknown environments, execute long-horizon instructions, zero-shot VLN-CE, requires an embodied

备注: 8 pages, 7 figures. Submitted to 2027 IEEE International Conference on Robotics Automation (ICRA)

点击查看摘要

Abstract:Vision-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to execute long-horizon instructions in unknown environments. Existing zero-shot VLN-CE systems typically maintain spatial states through geometric localization or coordinate-based representations. Recent geometry-constrained navigation removes depth and globally consistent coordinates, but maintaining persistent spatial awareness for place confirmation, progress verification, and recovery remains challenging. We present Navi-Agent, a zero-shot VLN-CE agent that constructs a coordinate-free spatial state from visual observations and executed motion histories. Navi-Agent organizes this state as a navigation topology, where nodes represent visual places and edges represent motion transitions. This representation enables observation-based approximate self-localization, task progress verification, and visual revisitation-based recovery. Navi-Agent performs closed-loop navigation by decomposing instructions into sub-goals, executing local visual navigation, and verifying visited places through the constructed spatial state. Experiments on zero-shot VLN-CE benchmark and real-world robot platforms show that Navi-Agent achieves state-of-the-art performance among geometry-constrained methods while remaining competitive with approaches relying on geometric localization.

32. 【2609.20386】Compact Vision Models for Iris Presentation Attack Detection under Presentation Attack Instrument Shift and Environmental Degradation

链接https://arxiv.org/abs/2609.20386

作者:Athanasios Angelakis,Marta Gomez-Barrero

类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)

关键词:Iris presentation attack, acquisition conditions absent, Classification Error Rate, development encounters presentation, Presentation Classification Error

备注: Accepted at BIOSIG 2026. This preprint includes minor nomenclature and editorial corrections clarifying the project-specific Patch-ABMIL and Compact-TransMIL variants

点击查看摘要

Abstract:Iris presentation attack detection (PAD) is security-critical when a subsystem that appears reliable during development encounters presentation attack instruments (PAIs) or acquisition conditions absent from validation data. We benchmark three compact scratch-trained computer-vision models, each with at most approximately 0.26 million trainable parameters, on the Notre Dame subset of LivDet-Iris 2017 under PAI-driven domain shift and environmental degradation. All models are trained without external pretraining or data augmentation and evaluated over five seeds. A validation-selected threshold is transferred unchanged to the known-attack, unknown-attack, corrupted, and pooled test partitions. From known to unknown attack presentations, Attack Presentation Classification Error Rate (APCER) increases by 17.11-30.47 percentage points and Detection Equal Error Rate (D-EER) increases by 7.38-12.73 percentage points. At the validation-selected threshold, ZACH-ViT obtains the lowest unknown-attack APCER (47.69 +/- 4.84%) and D-EER (38.87 +/- 0.93%), while Compact-TransMIL obtains the lowest Bona Fide Presentation Classification Error Rate (BPCER). ZACH-ViT also gives the lowest unknown-attack BPCER at an APCER limit of 10% (81.29 +/- 1.95%). The high absolute errors show that the comparative advantage of the best compact model does not constitute deployment readiness under unknown PAIs.

33. 【2609.20377】MM-Future: Multi-Mode Joint World-Action Modeling for Autonomous Driving

链接https://arxiv.org/abs/2609.20377

作者:Shuai Liu,Hechangle Gong,Hao Jiang,Runlin He,Junxiang Zhan,Kai Huang,Sheng Yang,Shaoqing Ren

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Autonomous driving involves, driving involves coupled, involves coupled decision-making, Autonomous driving, driving involves

备注

点击查看摘要

Abstract:Autonomous driving involves coupled decision-making and scene evolution under multi-mode uncertainty. To capture this coupling and uncertainty, we introduce MM-Future, a world-action model that generates multiple paired scene-action hypotheses and models bidirectional interaction within each pair. Each hypothesis is initialized from a structured action prior and an independent future scene source, which are then co-evolved through a modality-aware diffusion Transformer. To support efficient multi-mode rollout, MM-Future compresses multi-view video into planning-oriented representations, dubbed MM-Tokens. Finally, a future-conditioned proposal scorer ranks trajectory candidates by shared history context and their paired predicted future. On NAVSIM navtest, MM-Future achieves 94.0 PDMS and 91.5 EPDMS, while attaining a 32.3 HD-Score in zero-shot closed-loop evaluation on HUGSIM. Ablations show consistent improvements over both single-mode and action-only variants, validating the benefit of multi-mode joint world-action modeling.

34. 【2609.20348】EliGSiR: Continual RGB-D Mapping with Gaussian Splatting under Bounded Compute

链接https://arxiv.org/abs/2609.20348

作者:Björn Ellensohn,Elmar Rueckert,Christian Rauch

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Gaussian Splatting assumes, Incremental Gaussian Splatting, Gaussian Splatting, Load-adaptive Incremental Gaussian, Evidence-guided Load-adaptive Incremental

备注: 8 pages, 8 figures

点击查看摘要

Abstract:Conventional 3D Gaussian Splatting assumes a closed set of observations and long optimization schedules. Continual RGB-D mapping in contrast poses the problem that new observations arrive online, while previously reconstructed regions must be preserved. We present EliGSiR (Evidence-guided Load-adaptive Incremental Gaussian Splatting with Image Replay), a continual Gaussian mapper that controls how the available optimization budget is used as the reconstruction evolves. Map-Guided View Scheduling filters redundant incoming views and reconsiders retained views according to the current state of the map. Load-Adaptive Fidelity adjusts supervision resolution to the current mapping load instead of following a fixed resolution schedule. Targeted Geometry Growth separates depth supervision from Gaussian creation and adds geometric capacity only where repeated RGB-D observations indicate missing or misplaced structure. Together, these mechanisms adapt which views are optimized, how much image detail is used, and where the representation grows while mapping remains active. We evaluate EliGSiR on Replica, TUM RGB-D, ScanNet++, and real RGB-D sensor sequences, considering both the final reconstruction and the map available throughout acquisition. On TUM RGB-D fr3/long_office_household, EliGSiR reaches 21.52 dB with the same ground-truth mapping poses used by the controlled baselines, compared with 19.42 dB for SplaTAM. In the tracked-pose comparison, EliGSiR with live ORB-SLAM3 poses reaches 23.02 dB in 155.5 s, compared with 20.10 dB in 230.9 s for CaRtGS using its native tracker. We further evaluate reconstruction throughout acquisition and show how EliGSiR adaptive view scheduling, supervision fidelity, and geometry growth improve the use of the available mapping budget.

35. 【2609.20341】Fast Cross-Strength Multi-Contrast Brain MRI Translation using Latent Bridge Matching

链接https://arxiv.org/abs/2609.20341

作者:Siddharth Srivastava,Till Bretschneider

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Magnetic Resonance Imaging, Magnetic Resonance, Resonance Imaging, complicates downstream analysis, field strengths exhibits

备注: 10 pages, 4 figures. MRIxFields Workshop, MICCAI 2026

点击查看摘要

Abstract:Magnetic Resonance Imaging (MRI) acquired at different field strengths exhibits pronounced variation in noise, resolution, homogeneity, and contrast, which limits comparability across acquisition settings and complicates downstream analysis. We address this with a unified conditional model for controllable field-to-field synthesis, built on the framework of conditional latent bridge matching. Our single model achieves highly competitive results across the validation phase for all three tasks of the MRIxFields2026 challenge without task-specific architectures or training. We achieve fast generation with only a single inference step, producing all modality and field-strength combinations for $30$ axial slices in under $90$ seconds, as well as cross-modality-strength translation for a full volume in under $70$ seconds, on a single NVIDIA A5000 GPU. We further provide extensive ablations regarding different components of our solution. Code: this https URL

36. 【2609.20340】FreqDINO++: A Frequency-Guided Multi-Task Routing Vision Foundation Model for Universal Ultrasound Analysis

链接https://arxiv.org/abs/2609.20340

作者:Qing Xu,Yixuan Zhang,Yue Li,Xiangjian He,Qian Zhang,Mainul Haque,Rong Qu,Wenting Duan,Jieyun Bai,Zhen Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:comprehensive assessment requires, assessment requires jointly, requires jointly addressing, jointly addressing tasks, prenatal diagnosis

备注: Accepted by TBME

点击查看摘要

Abstract:Ultrasound image analysis plays a crucial role in cancer screening and prenatal diagnosis, yet comprehensive assessment requires jointly addressing tasks such as lesion segmentation and benign-malignant classification. While recent vision foundation models have shown remarkable universal representations, unlocking their potential for ultrasound is bottlenecked by the considerable domain gap from natural images. Existing methods typically fine-tune heavy vision encoders for isolated tasks, incurring substantial computational overhead while overlooking the underlying commonalities across heterogeneous tasks. In this work, we propose FreqDINO++, a frequency-guided multi-task routing vision foundation model for universal ultrasound analysis. We first introduce a Multi-task Routing Adapter (MR-Adapter) to support parameter-efficient integration of task-common and task-specific knowledge, a Frequency-aware Feature Enhancer (F$^2$-Enhancer) is then designed to capture the rich multi-scale frequency characteristics of ultrasound images, and a Task-aligned Collaborative Decoder (TC-Decoder) is devised to promote collaboration between dense and global prediction tasks through global-local token interaction. Extensive experiments on large-scale multi-task and external single-task ultrasound benchmarks demonstrate that FreqDINO++ consistently outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, while also showing promising generalization to unseen data. The code is at this https URL.

37. 【2609.20325】AgriScope: Pixel-Grounded Multimodal Understanding for Agricultural Images

链接https://arxiv.org/abs/2609.20325

作者:Abderrahmene Boudiaf,Mohamad Alanssari,Irfan Hussain,Sajid Javed

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:complex real-world conditions, Large Language Models, Multimodal Large Language, Agricultural image understanding, real-world conditions

备注

点击查看摘要

Abstract:Agricultural image understanding requires fine-grained recognition of plant diseases, pests, crop structures, and botanical species under complex real-world conditions. Despite recent advances in Multimodal Large Language Models (MLLMs), existing models remain limited to text-only outputs and lack pixel-level visual grounding capabilities. In this work, we introduce AgriScope, a unified pixel-grounded multimodal framework for agricultural image understanding. AgriScope jointly supports image-level, region-level, and pixel-level understanding within a unified framework, enabling tasks such as grounded caption generation, referring expression segmentation, and multi-turn multimodal interaction for agricultural imagery. AgriScope integrates biologically specialized semantic representations with dense spatial grounding through biological-semantic encoding, dense spatial representations, and pixel decoding. To support large-scale grounded learning, we introduce AgriGround, a large-scale pixel-grounded agricultural multimodal instruction-tuning dataset containing over 500K images and 11M instruction-following samples spanning plant disease analysis, crop and weed identification, insect pest recognition, and fine-grained botanical understanding. AgriGround is constructed through a multi-stage automatic annotation pipeline that integrates multimodal caption generation, phrase-level grounding, segmentation mask generation, and task-oriented instruction synthesis to produce densely grounded supervision. Extensive experiments across multiple agricultural vision-language tasks demonstrate the effectiveness of AgriScope in pixel-grounded multimodal understanding, establishing a strong benchmark for agricultural vision-language learning and visual grounding. The dataset and code will be made publicly available at (this https URL)

38. 【2609.20318】LLM-Guided Transformation of Non-Critical Driving Scenes into Safety-Critical Scenarios Using Augmented Reality

链接https://arxiv.org/abs/2609.20318

作者:Noura Fady,Farah Khaled,Catherine M. Elias

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Large Language Models, Autonomous Driving Systems, requires realistic safety-critical, Testing Autonomous Driving, costly and unsafe

备注

点击查看摘要

Abstract:Testing Autonomous Driving Systems (ADS) requires realistic safety-critical scenarios, but collecting such data from real-world driving is costly and unsafe. This paper presents an automated pipeline that transforms safe driving scenes into safety-critical scenarios by combining computer vision, Large Language Models (LLMs), and Augmented Reality (AR). The system detects and tracks road users, extracts safety features including distance, velocity, motion direction, and Time-to-Collision (TTC), and assesses scene criticality. Safe scenes are modified by an LLM, which generates realistic collision-inducing objects and behaviors that are integrated into the original scene using AR. The proposed pipeline was evaluated on the nuScenes dataset, achieving 97.52% safety classification accuracy and successfully generating realistic scenarios such as pedestrian crossings, rear overtaking vehicles, and sudden-stop events. The results demonstrate an effective and flexible approach for automated generation of safety-critical scenarios to support the testing and validation of autonomous driving systems.

39. 【2609.20299】Not All Layers Are Equal: Dynamic Layer Routing for Reliable CLIP OOD Detection

链接https://arxiv.org/abs/2609.20299

作者:Ignacio M. De la Jara,Cristian Rodriguez-Opazo,Damith Ranasinghe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Information aggregation, OOD detection, layer-wise information aggregation, improve OOD detection, improves OOD detection

备注

点击查看摘要

Abstract:Information aggregation across model layers are revealed to improve OOD detection. In contrast to crafting a method for layer-wise information aggregation in recent work, we investigate if layer selection is a learnable problem. In other words, we transpose the question from how to fuse layers to one asking which layers to trust for an input. Using a generalizable, weak, out of distribution context crafting approach for supervision, shown to be more effective than state of the art methods' mechanisms, we formulate learning a lightweight router to select a sparse, final-layer-anchored expert over CLIP's layer depth for OOD detection. Across three diverse benchmarks we demonstrate our learnable routing method dubbed Voyager improves OOD detection. On ImageNet-1K, Voyager achieves an average FPR@95 of 18.86, outperforming the strongest, comparable, prompt-learning method by 8.8 points. These gains persist across multiple supervision sources, including those used by existing state-of-the-art prompt-learning methods, demonstrating that, whilst our weak OOD supervision context is highly effective, the key advantage is realized from the learnable router component rather than the supervision source. Significantly, Voyager is highly practical; router learning takes approximately two minutes using less than 1 GB of memory, making it approximately 20x more efficient than current prompt-learning approaches. Anonymized Code: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.20299 [cs.CV]

(or
arXiv:2609.20299v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.20299

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
40. 【2609.20290】nyCNN: A 193K-Parameter Network for On-Device Plant Disease Detection, with a Cross-Dataset Robustness Diagnosis

链接https://arxiv.org/abs/2609.20290

作者:Ngoc-Bao Ho-Lam,Thai-Anh Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Sustainable Development Goal, Nations Sustainable Development, United Nations Sustainable, Detecting crop disease, Development Goal

备注: 13 pages, 3 figures. Accepted at ISRSD 2026

点击查看摘要

Abstract:Detecting crop disease early is central to sustainable agriculture and food security under United Nations Sustainable Development Goal 2 (Zero Hunger), and is especially urgent in resource-constrained regions where expert diagnosis is scarce but low-cost mobile devices are widespread. This paper presents TinyCNN, a lightweight convolutional neural network for on-device plant disease classification. TinyCNN uses depthwise separable convolution blocks and contains only 193,190 trainable parameters with 110.05M MACs for a 224x224 input image. On the 38-class PlantVillage benchmark, TinyCNN achieves 98.88% test accuracy and 98.03% macro-F1 while being approximately 58x smaller than ResNet18 and 11.8x smaller than a MobileNetV2 teacher, directly reducing the energy, memory, and cost footprint of inference in line with Green AI principles. The paper further analyzes vanilla knowledge distillation as a sustainable model-compression strategy; an ablation over alpha in {0.3, 0.5, 0.7} and T in {2, 4} selects alpha=0.3, T=4, producing a distilled TinyCNN with 98.81% test accuracy. Finally, cross-dataset evaluation from PlantVillage to PlantDoc reveals a substantial robustness gap under real-world conditions, which a Grad-CAM analysis attributes to off-leaf, background-driven attention consistent with shortcut learning. TinyCNN is thus an energy-efficient, deployable building block for sustainable agricultural intelligence, while field robustness remains the key barrier to durable real-world impact.

41. 【2609.20283】Queries Knew More Than We Thought: Uncovering Latent Knowledge in Segmentation Models

链接https://arxiv.org/abs/2609.20283

作者:Ignacio M. De la Jara,Cristian Rodriguez-Opazo,Damith Ranasinghe

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:deployed selection rule, Modern segmenters, model query-conditioned candidates, expensive computation, deployed selection

备注

点击查看摘要

Abstract:Modern segmenters often fail after the expensive computation has already been done: a useful mask is present among the model's query-conditioned candidates, but the deployed selection rule does not expose it. We study this output-selection bottleneck in frozen DETR-family models. A ground-truth-only oracle first shows substantial hidden headroom in already-computed mask proposals. This raises a simple question: How can we better use the masks a segmenter has already computed but does not expose? We then ask whether that headroom can be recovered without adding queries, generating new masks, rerunning the backbone, or updating weights. HYDRA is a small selector trained only on cached frozen outputs. At inference time, it scores the cached candidates against an explicit keep-baseline option and acts only when a held-out calibrated margin indicates the selected candidate is sufficiently better. Trained on training-split caches and calibrated on held-out data, HYDRA improves Mask2Former, MaskDINO, and OneFormer by up to +7.41 dataset mIoU points on ADE20k and COCO, and improves SAM 3 by +9.4 class-macro prompt-IoU points on average across eight domains while preserving useful predictions through calibration. Paired LoRA controls show that lightweight weight adaptation does not remove the bottleneck: exposed predictions are often flat or worse, while routing over the adapted candidates still recovers accuracy. Finally, we connect the effect to query specialization under bipartite matching and verify it in a controlled TinyDETR study. These results show that frozen segmenters should be evaluated not only by the masks they expose, but also by the useful candidates they suppress.

42. 【2609.20267】CleanVideo: Adaptive Concept Erasure for Text-to-Video Diffusion Models

链接https://arxiv.org/abs/2609.20267

作者:Junchi Liao,Hongji Li,Wenrui Zhou,Lijie Hu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:selectively eliminate undesired, pre-trained generative models, Concept erasure aims, eliminate undesired visual, general utility

备注

点击查看摘要

Abstract:Concept erasure aims to selectively eliminate undesired visual semantics from pre-trained generative models without compromising their general utility. Extending concept erasure from images to video is nontrivial. Target concepts emerge gradually and vary across frames and denoising steps. As a result, fixed interventions may miss the target or introduce blurring, jitter, and content distortion. We propose CleanVideo, a selective erasure framework that performs low-dimensional subspace intervention controlled by a tri-modal gating mechanism. By jointly processing spatiotemporal visual features, timestep signals, and textual semantics, CleanVideo determines where, when, and whether to intervene, steering erased content toward natural surrogate concepts when such surrogates can be clearly defined while preserving non-target content. Experiments on three video diffusion models show that CleanVideo effectively erases target concepts while maintaining visual fidelity and temporal coherence, outperforming existing baselines under frame-level and video-level evaluations and under concept-recovery attacks when the protected pipeline remains intact.

43. 【2609.20263】AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

链接https://arxiv.org/abs/2609.20263

作者:Tamoghna Chakraborty,Md Nurul Absur,Sourya Saha,Saptarshi Debroy

类目:Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)

关键词:practical detection threat, partially manipulated footages, fully fabricated clips, proliferation of generative, shifted the practical

备注: 7 Pages

点击查看摘要

Abstract:The proliferation of generative video models has shifted the practical detection threat from fully fabricated clips to partially manipulated footages. Although modern detectors achieve strong accuracy using foundation backbones of 400M+ parameters, their resource footprint precludes edge deployment. In this paper, we present a lightweight full-frame detector for partially manipulated AI-generated video, designed for deployment on edge hardware without face-detection preprocessing. The system distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student through a pipeline that combines temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter that conditions ImageNet features for artifact detection. We additionally target two failure modes specific to the partial-manipulation regime: false positives on legitimate scene cuts, addressed through within-video temporal hard negatives; and threshold-level miscalibration on the dominant pure-real class, addressed through calibration-aware sampling. Evaluation on a 55,393-sample spliced test set across fake-frame ratios from 6.2% to 31.2% demonstrates the student model closing 58% of the gap to the DINOv2-Base teacher (AUC 0.766) while running at 3.65 ms per 16-frame clip on RTX A4000 with a 150.4 MB checkpoint compatible with edge-device memory and latency budgets.

44. 【2609.20262】Generative Verification: Rethinking the Uncertainty Signal for Active Learning of Object Detection

链接https://arxiv.org/abs/2609.20262

作者:Licheng Zhang,Zheng Gong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:active object detection, object detection shares, function for active, active object, acquisition function

备注

点击查看摘要

Abstract:Nearly every acquisition function for active object detection shares one arrangement, in that the model being improved is also the model being interrogated. We depart from it. In generative verification an independent generative model re-derives the label of a detection from the pixels inside its predicted box, and the disagreement between the two becomes the acquisition signal. Two properties follow from the arrangement itself rather than from any tuning. A displaced box, a box on background and a correct box carrying the wrong label all yield a crop that fails verification, so the failure modes arrive already combined in one scalar and the hand-weighted classification and localization terms of existing criteria are no longer needed. And because the verifier never observes the detector confidence, confidently wrong detections score highest, although a self-derived signal reads them as uninteresting and they are the costliest to leave unlabeled. We build the verifier as a conditional diffusion model whose diffusion target is a label representation rather than an image. Its reverse process is stochastic, so repeated generations return a distribution whose concentration reports how firmly the evidence determines the label, where a classifier returns a single point estimate. On PASCAL VOC and MS-COCO the signal outperforms output-uncertainty, feature-geometry, perturbation and ensemble criteria, gaining about one mAP50 point per round on MS-COCO, with its largest margins in the early rounds where confident detector errors are most common.

45. 【2609.20248】Distance to Class Prototypes: Active Learning for Object Detection

链接https://arxiv.org/abs/2609.20248

作者:Licheng Zhang,Zheng Gong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:deep object detector, Deploying a deep, deep object, annotating data, setting is limited

备注

点击查看摘要

Abstract:Deploying a deep object detector in a new setting is limited less by architecture than by the cost of annotating data from that setting. Active learning lowers the cost by choosing which images to label, and the choice is only as good as the signal used to score an unlabeled image. That signal is usually the class posterior, which is cheap but poorly calibrated, or the disagreement across several models or several stochastic passes, which is better but multiplies inference over a pool far larger than the labeled set. We propose a signal richer than the posterior yet still read from one forward pass of one network. A supervised contrastive term added to the training objective shapes a per-object embedding space in which distance encodes class membership, and an unlabeled detection is scored by how far it lies from the region occupied by its predicted category, weighted by its confidence. The criterion needs no ensemble, no auxiliary predictor and no repeated inference, and its entire cost is 2.89M parameters, an increase of 8.3% over a bare detector. On PASCAL VOC and MS-COCO it beats the posterior of the same detector in every round in which a selection is made, by up to 1.08% mAP50 against run to run deviations of 0.02% to 0.18%, and it stays competitive with ensemble and Monte Carlo dropout criteria costing three to fifty forward passes per unlabeled image. Experiments use the single-stage detector under which the compared criteria report their results, so that the selection decision is isolated from the strength of the detector.

46. 【2609.20245】SAGE-Yoga: Multi-Cue Learning for Yoga Pose Classification and Joint-Level Correction

链接https://arxiv.org/abs/2609.20245

作者:Hung Le Chi,Khanh Minh Huynh,Long Nghia Tran Pham,Tan Phuc Huynh,Trong-Thuan Nguyen,Minh-Triet Tran

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Automated yoga analysis, yoga analysis requires, accurate pose classification, Automated yoga, pose classification

备注: Under Review for RIVF 2026

点击查看摘要

Abstract:Automated yoga analysis requires both accurate pose classification and interpretable feedback on pose execution. However, existing methods often rely on a single visual prediction, struggle to distinguish visually similar poses, and treat pose classification and correction as separate tasks. To address these limitations, we propose SAGE-Yoga, a unified coarse-to-fine framework for yoga pose classification and joint-level correction from a single RGB image. Inspired by how yoga instructors assess posture using multiple complementary cues, SAGE-Yoga first employs a bagging-based ensemble of complementary visual backbones to generate a ranked set of candidate pose classes. Additionally, a margin-based gating mechanism preserves confident visual predictions while invoking geometric verification only for ambiguous cases. Moreover, once the final pose class is determined, SAGE-Yoga retrieves a medoid reference pose and compares the observed joint angles with class-specific distributions to identify misaligned joints. Finally, these deviations are translated into actionable corrective feedback. Empirically, experiments on the Yoga-82 dataset show that the visual ensemble achieves 89.0% Top-1 accuracy, while the complete framework improves performance to 90.7% Top-1 accuracy and 90.1% Macro-F1. These results demonstrate that combining complementary visual evidence with selective geometric verification improves fine-grained pose classification while enabling interpretable, joint-level correction.

47. 【2609.20235】Scene-Q: Confidence-Aware Coarse-to-Fine Querying of 3D Scenes with Selective VLM Reasoning

链接https://arxiv.org/abs/2609.20235

作者:Juno Kim,Yesol Park,Hye-Jung Yoon,Byoung-Tak Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Indoor mobile robots, mobile robots require, Indoor mobile, robots require open-vocabulary, require open-vocabulary scene

备注: 8 pages, 5 figures, 3 tables. Submitted to IROS 2026

点击查看摘要

Abstract:Indoor mobile robots require open-vocabulary scene understanding that grounds natural-language queries in a consistent 3D map. Many existing systems ultimately rely on cosine-similarity retrieval with contrastive image--text encoders, which is efficient but brittle when labels are near-synonymous or multiple similar instances appear. We present Scene-Q, a confidence-aware coarse-to-fine querying framework that normalizes encoder scores with temperature scaling and selectively invokes a reasoning VLM only for low-confidence cases. High-confidence queries are answered by fast retrieval, while ambiguous ones are reranked over a small top-K candidate set using the original multi-view images and instance bounding boxes, enabling context-aware disambiguation at low cost. Scene-Q improves open-vocabulary 3D instance segmentation on ScanNet200 and natural-language 3D instance retrieval on real-world reconstructions, with the largest gains on spatial and relational queries while keeping a substantial fraction of queries on the fast path.

48. 【2609.20222】A Two-Stage Multi-Scale Attention-Based Network for Weakly Supervised Cataract Fundus Image Enhancement

链接https://arxiv.org/abs/2609.20222

作者:Xiaoyong Fang,Yue Wang,Xiangyu Li,Wanshu Fan,Dongsheng Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cataract fundus image, fundus image enhancement, images, image enhancement, fundus image

备注: Accepted by Scientific Reports

点击查看摘要

Abstract:Cataract is a major cause of vision loss and hinders further diagnosis. However, cataract fundus image enhancement often grapples with challenges such as limited paired cataract retinal images and insufficient recovery of fine details in the retinal images. To mitigate these challenges, we in this paper propose a two-stage multi-scale attention-based network (TSMSA-Net) for weakly supervised cataract fundus image enhancement. Our TSMSA-Net leverages the domain transformation to synthesis paired real-like cataract images, solving the problem of difficult acquisition of paired images. To further extract detailed information from fundus images and reduce the generation of artifacts during the enhancement process, we propose a multi-scale attention-based stage to learn more useful features for cataract image enhancement. Experimental results on Kaggle and ODIR-5K demonstrate that our TSMSA-Net outperforms current state-of-the-art cataract fundus images enhancement even without paired images and exhibits certain generalization ability. Experimental results on Kaggle and ODIR-5K datasets indicate that our TSMSA-Net outperforms the current state-of-the-art methods for cataract fundus image enhancement, even in the absence of paired images. Additionally, it demonstrates a certain level of generalization capability. The enhancement also can improve the performance of vessel segmentation and classification in cataract images.

49. 【2609.20191】VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial Robots

链接https://arxiv.org/abs/2609.20191

作者:Marco S. Tayar,Felipe Tommaselli,Gianluca Capezutto,Pedro Antonio Rabelo Saraiva,Pedro H. V. de Freitas,Lucas Kido,Guilherme Sonego,Ricardo V. Godoy,Marcelo Becker

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Systems and Control (eess.SY)

关键词:Running vision-language navigation, vision-language navigation fully, share limited compute, Running vision-language, navigation fully onboard

备注

点击查看摘要

Abstract:Running vision-language navigation fully onboard an aerial robot is hard, since grounding, planning, and control must share limited compute and a single-stage error is difficult to isolate in flight. End-to-end aerial policies fuse these stages into one network, giving up the observability and safety checks a modular stack keeps available. We propose VLN on the Fly, an onboard stack that keeps grounding, planning, and control as separate, inspectable stages. A quantized VLM grounds an instruction to a coarse image cell, depth lifts it to a 3D goal, a fast B-spline planner returns a feasible trajectory, and a pretrained reinforcement learning policy tracks it to motor commands across quadrotors. Across 15 onboard flights over three everyday referents in a controlled indoor volume, the stack reaches the target in 13 of 15 trials with 5.72 cm mean goal error and 39.3% average GPU utilization. In 6 additional cluttered-environment trials, the stack tracks collision-free trajectories under onboard perception gating.

50. 【2609.20181】MoSSGate: Memory-Modulated State-Space Gating for Skin Lesion Segmentation

链接https://arxiv.org/abs/2609.20181

作者:Anum Awan,Mahnoor Buriro,Muhammad Younas Khan,Md Imam Ahasan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:computer-aided dermatological diagnosis, limited computational budgets, reliable computer-aided dermatological, fine boundary details, jointly capture long-range

备注: 15 pages, 5 figures, accepted on International Conference on Cloud and Network Computing

点击查看摘要

Abstract:Accurate skin lesion segmentation is crucial for reliable computer-aided dermatological diagnosis, yet existing convolutional and transformer-based models often struggle to jointly capture long-range spatial dependencies and fine boundary details under limited computational budgets. This trade-off between global context modeling and boundary-aware localization frequently leads to over-segmentation, fragmented predictions, or missing thin peripheral structures. To address this challenge, we propose MoSSGate, a plug-and-play module for U-Net that integrates (i) boundary-aware spatial gating to restrict long-range propagation to informative regions, (ii) an external memory modulator that provides sample-adaptive dynamic control, and (iii) parallel 2D state-space modeling for efficient global context aggregation with linear complexity. The proposed design enables adaptive, context-aware information propagation while preserving sharp and accurate lesion boundaries. Extensive experiments on the ISIC 2017 and ISIC 2018 benchmarks demonstrate state-of-the-art accuracy with strong efficiency, achieving 86.3% and 85.9% mIoU and 92.6% and 90.6% Dice, respectively, while requiring substantially fewer FLOPs than most competing CNN-based methods. These results highlight a favorable accuracy efficiency trade-off for high-resolution medical image segmentation.

51. 【2609.20178】MTF-Net: Multi-Modal Temporal Feature Fusion Network for Pedestrian Intention Prediction

链接https://arxiv.org/abs/2609.20178

作者:Md Mahfuzur Rahman,Pengzhan Zhou,A. F. M. Abdun Noor,Md Imam Ahasan,Md Mustafizur Rahman,Fang Qu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Accurately predicting pedestrian, Accurately predicting, crucial for ensuring, ensuring safe, safe and proactive

备注: 15 pages, 5 figures, accepted on International Conference on Cloud and Network Computing

点击查看摘要

Abstract:Accurately predicting pedestrian intentions is crucial for ensuring safe and proactive interaction between autonomous vehicles and pedestrians. However, existing approaches often depend on architectures that either model temporal dependencies within individual modalities or fuse modalities only at coarse semantic levels. To address these limitations, we propose MTF-Net, a novel Multi-Modal Temporal Feature Fusion Network that jointly models kinematic, appearance, and contextual cues for pedestrian intention prediction. MTF-Net integrates four complementary modalities-bounding-box dynamics, human pose keypoints, local context, and scene-level semantics within a recurrent fusion framework enhanced by gated linear units (GLUs). These GLU-based modules adaptively regulate cross-modal information flow, enabling interpretable and efficient feature interaction across temporal scales. Through three dedicated temporal encoding branches and an attention-guided fusion head, the proposed model robustly anticipates pedestrian crossing intentions several frames before they occur. Extensive evaluations on the PIE and JAAD benchmarks demonstrate that MTF-Net surpasses recent transformer- and graph-based models, achieving up to 0.95 AUC on PIE and 0.94 AUC on JAAD, while maintaining real-time performance. The results highlight that reliable pedestrian intention prediction arises from principled multi-modal fusion rather than excessive architectural complexity.

52. 【2609.20160】Needles in a Raystack: Ultra-Sparse LiDAR Occupancy Detection for Bat Tracks

链接https://arxiv.org/abs/2609.20160

作者:Nico Klar,Pankaj Rana,Nizam Gifary,Jakob Traub,Aamir Ahmad

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Monitoring flying animals, protecting biodiversity, flying animals, animals is important, important for understanding

备注: 6 pages, 4 figures. Accepted and presented at the AI4Nature@AVSS 2026 Workshop of the 22nd International Conference on Advanced Visual and Signal-Based Systems (AVSS 2026), Lecce, Italy

点击查看摘要

Abstract:Monitoring flying animals is important for understanding and protecting biodiversity, but nocturnal species such as bats are difficult to observe in the field. Using LiDAR, bat movements at night result in ultra-sparse 3D spatio-temporal data in which standard reconstruction losses tend to predict only background and miss real flight paths. We study this problem as voxel-wise occupancy detection in sensor-centric LiDAR raystacks. A lightweight 3D U-Net is proposed that preserves temporal resolution, uses skip connections for spatial detail, and combines weighted binary cross-entropy with Dice loss to handle the strong class imbalance. In real LiDAR recordings of bats over open fields, cross-checked with acoustic monitoring, a reconstruction-based 3D convolutional autoencoder baseline fails to recover foreground trajectories. In contrast, the proposed U-Net recovers sparse foreground occupancy in diagnostic experiments and produces coherent occupancy patterns along bat flight trajectories, providing a practical basis for validation-scale experiments, later clustering of flight tracks, and future integration of bat activity information into biodiversity-aware turbine curtailment strategies.

Comments:
6 pages, 4 figures. Accepted and presented at the AI4Nature@AVSS 2026 Workshop of the 22nd International Conference on Advanced Visual and Signal-Based Systems (AVSS 2026), Lecce, Italy

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.20160 [cs.CV]

(or
arXiv:2609.20160v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.20160

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
53. 【2609.20151】Ischemic Stroke Segmentation and Net Water Uptake Quantification on Multicenter Non-Contrast CT Using Supervised Target-Domain Adaptation

链接https://arxiv.org/abs/2609.20151

作者:Linus Britt,Maximilian Nielsen,Susan Klapproth,Andre Kemmling,Michael H. Lev,Gabriel Broocks,Rene Werner,Thilo Sentker

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:non-contrast computed tomography, net water uptake, Quantitative assessment, diffusion-weighted MRI, limiting clinical applicability

备注

点击查看摘要

Abstract:Objectives: Quantitative assessment of infarct hypodensity on non-contrast computed tomography (NCCT), including net water uptake (NWU), requires manual or semi-manual lesion delineation, often guided by CT perfusion or diffusion-weighted MRI, limiting clinical applicability. Automated segmentation on NCCT could enable efficient biomarker extraction such as NWU but remains challenging across heterogeneous multicenter data. This study aimed to develop and externally test a domain-aware deep learning framework for ischemic stroke segmentation on NCCT and assess its suitability for NWU quantification. Materials Methods: In this retrospective multicenter study of 801 patients from four datasets, an nnU-Net-based model was trained on NCCT scans from the University Medical Center Hamburg-Eppendorf and the Acute Ischemic Stroke Dataset. To adapt to new domains, the model was fine-tuned on target-domain subsets from Boston (n=11) and ISLES (n=75), with evaluation on held-out cases not used for fine-tuning. Automated segmentations and NWU values were compared with expert references. Results: For lesions $\geq$ 30 mL, median Dice was 0.68 (Boston) and 0.56 (ISLES). Including smaller lesions, which predominated in ISLES, median Dice was 0.54 (interquartile range [IQR] 0.30-0.70) for acute lesion segmentation (Boston dataset) and 0.20 (IQR 0.03-0.41) for NCCT lesion segmentations when compared to post-treatment infarct (primary target of the ISLES challenge). Automated NWU mean absolute error was 1.37 percentage points (SD 1.61, Boston). Conclusion: Target-domain adaptation supported NCCT-only infarct segmentation across heterogeneous external cohorts, although performance varied across domains. The approach enabled low-error NWU quantification from baseline NCCT without advanced imaging, supporting further prospective clinical evaluation.

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2609.20151 [cs.CV]

(or
arXiv:2609.20151v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.20151

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Thilo Sentker [view email] [v1]
Thu, 17 Sep 2026 12:44:00 UTC (4,594 KB)

54. 【2609.20150】ask-Oriented Semantic Feature Transmission for Multi-Task Satellite Remote Sensing over Low-SNR Channels

链接https://arxiv.org/abs/2609.20150

作者:Shuoyuan Sun,Hongyu Wang,Mugen Peng,Wenjia Xu

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Conventional satellite remote, optimizes pixel-level fidelity, satellite remote sensing, remote sensing transmission, Conventional satellite

备注

点击查看摘要

Abstract:Conventional satellite remote sensing transmission follows a reconstruct-then-infer paradigm that optimizes pixel-level fidelity, creating an objective mismatch with downstream tasks such as classification and detection, especially at low SNR. This paper investigates a task-oriented framework that bypasses image reconstruction and directly transmits semantic features extracted by a multitask-pretrained backbone. A lightweight channel adaptation module (CAM) compresses feature dimensionality for bandwidth reduction, and a feature restorer recovers task-relevant structure after channel corruption. With the backbone frozen, the CAM and task-specific downstream heads are jointly optimized with task and feature-level supervision under random-SNR training. Under the adopted AWGN setting, experiments on scene classification and object detection show consistent gains over reconstruction-oriented JSCC baselines across different SNR conditions, with the largest improvements in the low-SNR regime.

55. 【2609.20147】Bridging Modalities on the Cortex: Surface-based MRI to PET Translation with a Diffusion Bridge

链接https://arxiv.org/abs/2609.20147

作者:Yitong Li,Alexandra Samoylova,Fabian Bongratz,Timo Grimmer,Dennis M. Hedderich,Igor Yakushev,Christian Wachinger

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Positron Emission Tomography, Fluorodeoxyglucose Positron Emission, Emission Tomography, Fluorodeoxyglucose Positron, Positron Emission

备注

点击查看摘要

Abstract:Cortical hypometabolism measured by Fluorodeoxyglucose Positron Emission Tomography (FDG-PET) is a highly sensitive biomarker for dementia diagnosis. However, high costs, radiation exposure, and limited accessibility constrain its clinical utility. While cross-modal synthesis from Magnetic Resonance Imaging (MRI) offers a promising alternative, existing volumetric generation methods do not explicitly account for the highly folded cortical geometry, where disease-related patterns predominantly reside. To address this, we introduce a novel surface-based diffusion bridge framework DB-SUiT for MRI-to-PET translation that operates natively on the cortical manifold. A conditional Spherical U-shaped vision Transformer (SUiT) is specifically designed to model the intricate cross-modal relationships while preserving surface topology. It combines spherical convolutional encoders for multi-scale surface feature extraction with bottleneck Transformers to capture long-range spatial dependencies, while incorporating demographic and subcortical conditions to refine the synthesis. Evaluated on two datasets, including subjects with different dementia types, DB-SUiT demonstrates high-fidelity synthesis that substantially outperforms other baselines. In automated dementia classification, synthesized PET surfaces improve performance over MRI by 14.2% and PET volumes by 11.3%, approaching the performance of real PET surfaces. In a blinded reader study, synthetic PET achieved 85.5% diagnostic accuracy, compared with 75.8% for MRI and 95.2% for real PET. This further demonstrates cross-cohort and cross-pathology generalization, as the model was evaluated without retraining on an external cohort that included a dementia subtype not represented during training. Our code is available at this https URL.

56. 【2609.20139】Cross-Modal Attention Acts as a Frequency Filter: Why Verbose Prompts Improve Robustness in Vision-Language Models

链接https://arxiv.org/abs/2609.20139

作者:Farooq Ahmad Wani,Maria Sofia Bucarelli,Mujtaba Hussain Mirza,Oleksandr Pryymak,Aryo Pradipta Gema,Iacopo Masi,Pasquale Minervini,Fabrizio Silvestri

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Vision-language models, Vision-language, Verbose questions, question affects VLMs, VLMs

备注

点击查看摘要

Abstract:Vision-language models (VLMs) are fragile under image corruption. We find that the wording of the question affects VLMs in two opposite ways. Verbose questions make VLMs substantially more robust---e.g., rephrasing "Is there a cat?" into "Please look carefully and answer: is there a cat?". Conversely, VLMs become more fragile under corruption when the question is semantically complex or finer-grained, e.g., "what colour is the cup left of the chair?" instead of "is there a cup?". Both effects stem from question-conditioned cross-modal attention, which induces a spectral filter over image patches: verbose questions broaden its frequency support, while fine-grained questions concentrate it onto fewer visual scales. The model's answer drifts most when this filter and the corruption sit on the same spatial frequencies. We test the filter view on Qwen3-VL and LLaVA-OneVision across GQA and CLEVR; verbose paraphrasing reduces drift variance by 70--81% on the 8B models. The practical recipe---pad the prompt---further yields measurable gains in accuracy, even under image corruption.

57. 【2609.20106】AnyviewMeter: Adapting Robotic Reward Models with Camera Geometry and Multi-View Attention

链接https://arxiv.org/abs/2609.20106

作者:Yuang Tu,Runjia Tan,Yujie Yan,Jinghan Hu,Chen Lv

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Robotic reward models, evaluate task execution, underlying task state, reward models evaluate, Robotic reward

备注: 8 pages, 3 figures, 5 tables

点击查看摘要

Abstract:Robotic reward models evaluate task execution from visual observations, but their predictions can change with camera viewpoint and occlusion even when the underlying task state is unchanged. Adapting a pretrained reward model to a local task therefore requires accounting for how that task is observed. We introduce AnyviewMeter, a geometry-conditioned adaptation framework for robotic reward models that represent task progress as a scalar reward signal. It combines low-rank fine-tuning with token-aligned Plucker rays and synchronous block attention: ray conditioning incorporates camera geometry into visual features and attention queries and keys, while block attention fuses synchronized views inside the pretrained decoder. The framework supports both single-view reward prediction and joint multi-view evaluation through parameter-efficient adaptation of a pretrained Robometer model. On PickCube, single-view adaptation improves progress prediction in every camera group and reduces mean absolute error under a changed field of view by approximately 21% relative to RGB fine-tuning. Across simulated manipulation tasks, joint multi-view prediction reduces progress error by 41-69% compared with averaging single-view RGB predictions and improves temporal ordering in approximately 88% of task-camera groups. On real tasks with fixed and wrist-mounted cameras, mean absolute error decreases by approximately 21% relative to averaged RGB fine-tuning. These results support camera geometry and joint visual evidence as useful components of task-specific robotic reward adaptation.

58. 【2609.20100】A Smaller Transformer in Your Transformer

链接https://arxiv.org/abs/2609.20100

作者:Dhananjay Tomar,Marius Aasan,Andreas Kleppe,Adín Ramírez Rivera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision Transformers settle, Vision Transformers, similar computational phases, locally similar computational, Recent findings

备注: 22 pages, 6 figures, 6 tables. Accepted at the 37th British Machine Vision Conference (BMVC 2026)

点击查看摘要

Abstract:Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.

59. 【2609.20088】G^2RA-NET: Graph-based Cross-Slice Relation Modeling with Attention Gating for Medical Image Segmentation

链接https://arxiv.org/abs/2609.20088

作者:Shengye Wang,Zonglin Wu,Liang Fan,Yule Xue,Haozhe Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:supports quantitative clinical, quantitative clinical analysis, Medical image segmentation, segmentation supports quantitative, Medical image

备注: 5 pages, 6 figures

点击查看摘要

Abstract:Medical image segmentation supports quantitative clinical analysis and computer-aided diagnosis. Recent methods for medical image segmentation have improved both local feature representation and volumetric context modeling. However, existing methods still strug- gle to efficiently model cross-slice relations in anisotropic volumet- ric images, limiting segmentation consistency and accuracy. This pa- per proposes G^2RA-Net, a medical image segmentation framework that combines graph-based cross-slice relation modeling with atten- tion gating. Graph-Based Slice Relationship Modeling (GSRM) cap- tures anatomical dependencies across consecutive slices by repre- senting each slice as a graph node and propagating semantic con- text through graph message passing. The Cross-Slice Attention Gate (CSAG) then selects relevant neighboring context and emphasizes target anatomical regions through attention-guided feature modula- tion. Experiments on brain MRI and abdominal CT datasets demon- strate that G^2RA-Net outperforms representative methods in seg- mentation accuracy and boundary quality. Ablation studies further validate the proposed design.

60. 【2609.20066】PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

链接https://arxiv.org/abs/2609.20066

作者:Zongze Wu,Baofeng Jia,Weiqi Yan,Jingyuan Zhang,Yu Zang,Xiaoyu Chen,Jing Han

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:tiny UAV detection, cameras offer high, UAV detection, tiny UAV, offer high temporal

备注: Code: [this https URL](https://github.com/wzz-z/PointEvent)

点击查看摘要

Abstract:Event cameras offer high temporal resolution and motion sensitivity for tiny UAV detection, yet distant targets generate sparse and fragmented events that are easily overwhelmed by clutter and ego-motion. Existing methods mainly rely on dense event representations or local sparse spatiotemporal modeling, resulting in redundant computation or fragmented modeling of motion continuity across distant asynchronous events. To address this limitation, we introduce serialized motion evidence accumulation, which treats motion continuity as an ordered evidence propagation process. Specifically, the same event stream is organized into locality-preserving spatiotemporal paths and chronology-preserving temporal paths through the latent complementary serializations. Based on this principle, we propose PointEvent, a lightweight event-wise state-space framework that alternates serialized scans across the complementary orders, progressively consolidating fragmented motion evidence beyond fixed local neighborhoods. A high-resolution event branch preserves fine-grained target responses, while compact context modulation suppresses interference. Experiments demonstrate that PointEvent achieves SOTA with the fewest parameters and fastest measured inference among the compared methods. Code: this https URL

61. 【2609.20064】A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

链接https://arxiv.org/abs/2609.20064

作者:Benjamin Kiessling(ALMAnaCH)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:impressive reported scores, limited practical uptake, reported scores, computational cost, dependence on large-scale

备注

点击查看摘要

Abstract:Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.

62. 【2609.20034】Astronex-World 1.0: Real-Time Interactive World Model Foundation

链接https://arxiv.org/abs/2609.20034

作者:Xin Zhou,Cong Miao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)

关键词:video world-model foundation, open controllable video, controllable video world-model, present Astronex-World, world-model foundation

备注: Technical report. 25 pages, 13 figures, 10 tables. Project page: [this https URL](https://world.astronex.com.cn) ; Code: [this https URL](https://github.com/Astronex-Robotics/Astronex-World) ; Weights: [this https URL](https://huggingface.co/Astronex-Lab/Astronex-World)

点击查看摘要

Abstract:We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.

63. 【2609.20012】GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction

链接https://arxiv.org/abs/2609.20012

作者:Enpeng Li,Yunzhou Zhang,Zhiyao Zhang,Dexuan Lyu,Chenyu Wang,Chiyuan Cui,Cheng Cheng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:paradigm for scene, scene modeling, modeling from image, GPU memory footprint, image sequences

备注: Accepted to ECCV 2026 as a Spotlight presentation

点击查看摘要

Abstract:Feed-forward 3D reconstruction provides an efficient paradigm for scene modeling from image sequences. Scaling these models to large monocular scenarios are constrained by excessive GPU memory footprint, degraded local geometry, and long-term trajectory drift. Existing chunk-based optimization strategies provide limited geometric constraints and fail to maintain global consistency over extended trajectories. We present a unified framework for stable and scalable feed-forward 3D reconstruction from long monocular sequences. Our approach builds on coarse-to-fine trajectory alignment augmented by lightweight geometric prior injection. Distilling monocular geometric cues into the feed-forward backbone via LoRA adaptation improves depth accuracy on fine structures while preserving inference efficiency. We introduce a hybrid-weight sparse ray-field optimization that leverages high-frequency geometric features to guide local point-cloud refinement and enforce consistent inter-frame ray constraints. Unlike prior chunk-based methods, this establishes strong cross-frame geometric coupling while maintaining scalability. Finally, an efficient trajectory stitching strategy with joint ray-error optimization explicitly reduces accumulated drift. Extensive experiments show that our approach achieves competitive trajectory accuracy compared with representative SLAM systems, while maintaining globally consistent 3D reconstruction in large-scale scenarios.

64. 【2609.19991】AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

链接https://arxiv.org/abs/2609.19991

作者:Longyin Zhang,Parth Sakhare Mahendra,Chengwei Wei,Ning Zhang,Lim Ming Chong,Sirui He,Ai Ti Aw

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:preserve event order, describe video content, judge audio-visual synchronization, Temporal Reasoning Assessment, preserve event

备注

点击查看摘要

Abstract:Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split's majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.

65. 【2609.19990】QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning

链接https://arxiv.org/abs/2609.19990

作者:Shengli He(1),Yongchao Liang(1),Roumeng He(2),Junjie Zeng(1),Jiyuan He(1),Can Wu(1),Li Zheng(1) ((1) Guizhou University, (2) Shanghai Ocean University)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reduce later-layer computation, high visual-token load, multimodal large language, preserve query-relevant evidence, motivates training-free pruning

备注

点击查看摘要

Abstract:The high visual-token load in multimodal large language models (MLLMs) motivates training-free pruning to reduce later-layer computation, but under a fixed budget, pruning must preserve query-relevant evidence while avoiding redundancy. Existing methods rank tokens, diversify selected subsets, or optimize coverage without using a shared per-visual query utility to weight both visual targets and candidate representatives. We introduce QCPruner, which makes both roles query-conditioned through bilateral utility weighting. Using keyword-matched query anchors, QCPruner fuses two cross-modal cues into utility and applies it to both visual targets and candidate representatives within visual-affinity-based coverage. The resulting nonnegative facility-location objective is monotone and submodular, retains the standard (1-1/e) greedy guarantee, and requires no model training or parameter updates. Across LLaVA-1.5, LLaVA-NeXT, LLaVA-Video, and Qwen2.5-VL, QCPruner achieves the highest average relative performance among evaluated complete-system pruning methods at every reported token budget. At 32 of 576 tokens on LLaVA-1.5-7B, it retains 96.1% of unpruned performance, versus 93.9% for the strongest evaluated baseline. At 256 of 1296 tokens on Qwen2.5-VL-7B, the corresponding values are 96.7% and 92.5%.

66. 【2609.19973】An Event Preserving Velocity Invariant Representation for Event Cameras

链接https://arxiv.org/abs/2609.19973

作者:Mikihiro Ikura,Luna Gava,Jiahang Wu,Chiara Bartolozzi,Arren Glover

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:cameras provide low-latency, URL novel circuitry, high temporal resolution, http URL, temporal resolution perception

备注: @inproceedings{ikura2026event, title={An Event Preserving Velocity Invariant Representation for Event Cameras}, author={Ikura, Mikihiro and Gava, Luna and Wu, Jiahang and Glover, Arren and Bartolozzi, Chiara}, year={2026}, booktitle={ECCV 2026 Workshop-Event-Based Multimodal Vision: From Imaging to Perception and Understanding} }

点击查看摘要

Abstract:Event cameras provide low-latency, high temporal resolution perception for real-time vision tasks such as this http URL novel circuitry (i.e. asynchronous, independent pixels) that enables these advantages also introduces new algorithmic challenges. Velocity-invariant representations alleviate missing observations under slow motion and motion blur under fast motion, but most discard temporal information by converting events into image-like representations. We propose Set of Centre Active Receptive Fields (SCARF), a real-time velocity-invariant representation that preserves raw events while consistently handling fast motion, stationary scenes, and independently moving objects. SCARF achieves state-of-the-art performance in both computational efficiency and representation quality.

67. 【2609.19966】Beyond the Foreground: FOV-Aware Polyp Image Synthesis via Lesion-Guided Adaptive Mucosal Context Propagation

链接https://arxiv.org/abs/2609.19966

作者:Tong Wang,Yuting He,Bin Ren,Yutong Xie,Guanyu Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scarce colonoscopy annotations, generating compatible mucosa, alleviate scarce colonoscopy, realistic synthesis requires, synthesis requires preserving

备注

点击查看摘要

Abstract:Synthetic image and mask pairs can alleviate scarce colonoscopy annotations, but realistic synthesis requires preserving the supplied lesion while generating compatible mucosa. Existing foreground-guided methods treat all non-foreground pixels as background and rely mainly on local integration. Directly applying them to colonoscopy causes two problems: non-mucosal black regions contaminate generated tissue, and local reasoning produces inconsistent mucosal texture and illumination. We propose LAMP, the first foreground-guided framework for polyp image synthesis based on lesion-guided adaptive mucosal context propagation. LAMP explicitly separates the lesion, valid mucosa, and camera exterior using a field-of-view (FOV) mask. Lesion-to-Mucosa cross-attention extracts lesion appearance conditions for valid-mucosa locations, while FOV-constrained multidirectional Vision Receptance Weighted Key Value propagates them over legal tissue support. An adaptive gate then controls their residual fusion into the diffusion U-Net. Extensive experiments on five polyp datasets demonstrate that LAMP substantially outperforms existing methods in overall generation quality and consistently improves five downstream segmentation models. Our code will be released at this https URL.

68. 【2609.19964】Enhanced Knowledge Distillation for Detection Transformer via Teacher Prediction Refinement

链接https://arxiv.org/abs/2609.19964

作者:Yitong Xing,Yuhao Cheng,Yanping Li,Yichao Yan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Detection Transformers, achieve strong performance, high computational cost, edge devices due, object detection

备注

点击查看摘要

Abstract:Detection Transformers (DETRs) achieve strong performance in object detection but remain challenging to deploy on edge devices due to their high computational cost. Existing DETR distillation methods mainly focus on aligning distillation points, while largely overlooking the quality of the teacher's supervision itself. We observe that due to stage-wise non-monotonic prediction behavior in DETRs, well-localized or correctly classified predictions from earlier stages may degrade in later ones, and some negative predictions become increasingly overconfident. As a result, relying solely on the current stage's predictions yields inaccurate and inconsistent supervision. To address this issue, we propose Teacher Prediction Refinement Distillation (TPRD), a plug-and-play module that refines teacher predictions before distillation by exploiting stage-wise prediction information. TPRD improves supervision quality through Positive Prediction Correction (PPC), which corrects degraded positive predictions by restoring more accurate ones from earlier stages, ensuring reliable localization and classification signals, and Negative Prediction Suppression (NPS) suppresses the influence of overconfident negatives, preventing them from providing misleading supervision to the student. To preserve informative dark knowledge, we further introduce Maximum Dark Knowledge Preservation (MDKP), which selectively refines target-class logits while retaining non-target relations. Extensive experiments on MS COCO and PASCAL VOC demonstrate the effectiveness and robustness of the proposed method. Our code is available at this https URL.

69. 【2609.19954】LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery

链接https://arxiv.org/abs/2609.19954

作者:Jingwei Song,Javid Hussain Jakir,Ray Zhang,Wenwei Zhang,Hao Zhou,Xiaomeng Xian,Maani Ghaffari

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:monocular laparoscopic surgery, proposes a real-time, laparoscopic surgery, work proposes, monocular laparoscopic

备注: This paper has been accepted by IEEE Transactions on Medical Robotics and Bionics (T-MRB)

点击查看摘要

Abstract:This work proposes a real-time 6 Degree-of-Freedom (6 DoF) tracking algorithm for monocular laparoscopic surgery. It provides alignment between intra-operative video and pre-operative data (e.g., CT). The 6 DoF tracking offers a solution for accurately locating the internal anatomy of the target organ despite the lack of tactile feedback and transparency. The ORB-SLAM2 framework is adopted and modified for prior-based 3D tracking with four major modifications. First, the primitive 3D shape is used for fast initialization of the ORB-SLAM2 monocular mode. Second, a pseudo-segmentation strategy is employed to separate the target organ from the background for tracking. Third, the 3D shape is incorporated as a geometric prior in its pose graph optimization. Fourth, the Multi-Scale Retinex with Chromaticity Preservation (MSRCP) algorithm is leveraged and modified for image enhancement in challenging illumination scenarios. In-vivo and ex-vivo experiments validate that LapaTrack-3D provides robust 3D tracking and effectively handles typical challenges such as poor illumination, fast motion, out-of-field-of-view scenarios, partial visibility, and ``organ-background'' relative motion. LapaTrack-3D achieves a processing rate of 13 Hz for 1280*720 pixel video.

70. 【2609.19927】DirtyMoCap: Robust Motion Capture from Unconstrained Markers

链接https://arxiv.org/abs/2609.19927

作者:Long Wang,Shuting Zhao,Shen Yan,Siyuan Yu,Xiaoben Li,Zeyu Cai,Yumeng Hou,Yuliang Xiu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:capture delivers high-fidelity, clean trajectories severely, trajectories severely limits, delivers high-fidelity human, motion capture delivers

备注: Homepage: [this https URL](https://wanglongzju.github.io/DirtyMoCap-Project-Page)

点击查看摘要

Abstract:Optical motion capture delivers high-fidelity human motion, but its reliance on strict marker layouts and clean trajectories severely limits its real-world applicability. In practice, tracking systems frequently output unconstrained markers: sparse, noisy, and unordered point clouds with unknown or varying configurations. To bridge the gap between corrupted raw markers and parametric human models, we introduce DirtyMoCap, a robust, marker-layout-free framework. Our core insight is to map unordered marker observations to a fixed set of "proxy anchors" comprising skeletal joints and body surface points, which serve as a stable intermediate representation. We first initialize and track these anchors over long sequences using a recurrent sliding-window architecture. Then, a custom differentiable Gauss-Newton solver fits the SMPL-H model to the tracked anchors to recover full-body pose, translation, and shape. By explicitly deriving geometric residuals, our solver learns adaptive observation confidence, smoothness, and prior weights end-to-end, adapting dynamically to the reliability of the input data. Extensive experiments on diverse, noisy marker configurations demonstrate that DirtyMoCap successfully generalizes across arbitrary layouts using only a single trained model. It consistently outperforms state-of-the-art configuration-specific baselines in both joint and vertex reconstruction accuracy, while our custom CUDA solver achieves up to a 100x speedup over standard PyTorch implementations. We further apply DirtyMoCap to heterogeneous raw optical MoCap recordings of traditional Chinese martial arts, yielding a Kung Fu motion dataset of temporally coherent SMPL-H reconstructions. Code and data are available at this https URL.

71. 【2609.19911】CitySTAR: Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

链接https://arxiv.org/abs/2609.19911

作者:Shuai Zhang,Hongye Hou,Qinghe Liu,Zhuoxiao Li,Dongli Wu,Jing Ou,Yuan Liu,Wufan Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:localize target entities, aims to localize, localize target, natural language, language and plays

备注

点击查看摘要

Abstract:3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.

72. 【2609.19907】GS-PI: An Optimization-Decoupled Appearance Decomposition Approach for Generating PBR Gaussian Assets

链接https://arxiv.org/abs/2609.19907

作者:Jieting Xu,Rengan Xie,Zijian Huang,Zehui Jin,Rui Wang,Yuchi Huo

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:encodes baked-in radiance, physically based rendering, tightly entangling illumination, preventing seamless integration, Gaussian Splatting

备注

点击查看摘要

Abstract:Gaussian Splatting (GS) excels at novel-view synthesis but encodes baked-in radiance, tightly entangling illumination with geometry and preventing seamless integration into physically based rendering (PBR) pipelines. Existing inverse-rendering methods attempt to disentangle materials via joint optimization, but often suffer from competing objectives that cause severe ambiguities and residual lighting artifacts. To overcome this, we present GS-PI, a novel optimization-decoupled framework that casts PBR material generation as a geometry-conditioned diffusion process on 3D point clouds. By operating directly in the 3D domain, our method inherently guarantees multi-view consistency, sidestepping the severe pixel correspondence issues that challenge 2D diffusion approaches. We introduce a multi-scale cross-view conditioning mechanism that integrates three complementary components: a global semantic prior, source-anchored photometric cues, and an absolute spatial learned view-direction conditioning signal. This design efficiently compresses complex multi-view evidence, mitigating cross-view projection misalignment and successfully preventing specular highlights from baking into intrinsic colors. By extracting a point cloud from a pre-trained Gaussian model, predicting PBR attributes via conditional diffusion, and distilling them back through differentiable rasterisation, we yield a fully relightable PBR-GS asset. GS-PI outperforms recent inverse-rendering baselines while replacing per-scene joint illumination/BRDF optimization with a learned diffusion pass followed by a short target-driven distillation, without requiring proxy meshes.

73. 【2609.19881】BinoGen: Scaling egocentric binocular data for embodied visual perception and learning

链接https://arxiv.org/abs/2609.19881

作者:Chunpeng Li,Ya-tang Li

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:temporally coherent visual, relies on temporally, temporally coherent, accumulated through continuous, continuous engagement

备注

点击查看摘要

Abstract:Embodied visual perception relies on temporally coherent visual experience accumulated through continuous engagement with the environment. However, collecting large-scale egocentric binocular observations together with dense annotations remains costly and difficult. Moreover, visual experience is shaped not only by the environment but also by the embodiment of the observer, including viewing height, field of view, binocular geometry, and motion through the scene. To address these challenges, we present BinoGen, an automated framework for generating large-scale, embodiment-aware egocentric binocular visual experiences in indoor environments. BinoGen jointly models environmental and observer variation through generative scene synthesis, probabilistic object instantiation, appearance randomization, stochastic trajectory generation, and configurable binocular camera setups. The framework produces synchronized binocular videos together with dense multimodal supervision, including depth maps, optical flow, surface normals, semantic maps, object coordinates, and camera poses. Using BinoGen, we construct a dataset comprising more than 20 million annotated images for supervised learning. We demonstrate two complementary utilities of BinoGen. First, incorporating BinoGen data consistently improves real-world visual perception, including depth estimation, object detection, and video object tracking. Second, paired human-inspired and mouse-inspired observations from the same environments enable controlled investigation of how observer embodiment affects perceptual learning. Embodiment-specific adaptation substantially improves performance, while joint training enables a single model to perform competitively across both embodiments. Together, these results demonstrate that large-scale, controllable visual experience can improve embodied perception...

74. 【2609.19876】SlugTrails: An Egocentric Benchmark for Floor Plan Localization in Large Buildings

链接https://arxiv.org/abs/2609.19876

作者:Yunqian Cheng,Roberto Manduchi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:enables infrastructure-free positioning, small residential environments, residential environments unlike, visual localization enables, localization enables infrastructure-free

备注: 8 pages, 4 figures. Preprint. Code and data: [this https URL](https://github.com/Head-inthe-Cloud/SlugTrails)

点击查看摘要

Abstract:Floor-plan-based indoor visual localization enables infrastructure-free positioning, but most methods are developed and evaluated in small residential environments unlike the large public buildings of real deployment. We introduce SlugTrails, a floor plan localization benchmark for large indoor spaces under realistic egocentric sensing: $30$ Hz Aria glasses recordings across three campus buildings and six floors ($22089$ m$^2$ of floor plan outline), CAD-derived floor plans with semantic classes and circulation space masks, and trajectories aligned into the floor plan frame using laser-surveyed anchors. One protocol covers three practical ways of gathering geometry under a limited field of view -- a single walking frame, a stationary multi-view sweep, and a walking stream with odometry -- so methods designed for different regimes are compared on the same buildings and ground truth. Evaluating five representative geometric and learned systems under their native sensing configurations, we find that stock checkpoints (official released weights) are near zero on SlugTrails (at most $0.004$ R@1m30$^{\circ}$ on walking single frames), while fine-tuning on SlugTrails improves every trainable family on all three tasks (e.g., F$^3$Loc $0.0 \rightarrow 0.141$ single-frame and $0.03 \rightarrow 0.66$ sequential), with gains compounding as observations accumulate. The same fine-tuned weights also improve cross-dataset generalization on LaMAR with no LaMAR training (sequential R@1m $0.048 \rightarrow 0.143$ for F$^3$Loc and $0.063 \rightarrow 0.127$ for UnLoc), whereas train-from-scratch on SlugTrails alone stays far below fine-tuning from stock weights -- evidence that floor plan localization is currently limited by indoor data rather than by architecture. We release the dataset, protocols, and tools at this https URL.

75. 【2609.19875】BINDER: A Latent Variable Model for Probabilistic Medical Image Registration

链接https://arxiv.org/abs/2609.19875

作者:Stefano Cerri,Amirhossein Hassankhani,Yaël Balbastre,Koen Van Leemput

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:general-purpose medical image, information registration criterion, mutual information registration, probabilistic model, model for general-purpose

备注

点击查看摘要

Abstract:We propose a new probabilistic model for general-purpose medical image registration that builds upon the mutual information registration criterion. It centers around a spatial interpolation technique that assumes latent voxel-wise correspondences between the images being registered. By exploiting these latent variables, we derive dedicated optimization and MCMC sampling techniques that only involve closed-form iterative updates. When applied to nonlinear registration, an efficient demons-like optimization algorithm is obtained that shows robust out-of-the-box performance across a variety of monomodal and multimodal registration tasks. We also demonstrate a corresponding sampler that can quantify, for the first time, uncertainty in multimodal registration scenarios with very high-dimensional 3D deformations. Our code, which we call BINDER (Bayesian INference for DEformable Registration), is freely available at this https URL.

76. 【2609.19872】PART: Learning 3D Part Assembly and Retrieval with Transformers

链接https://arxiv.org/abs/2609.19872

作者:Ruchao Bao,Wenzheng Wu,Chucheng Xiang,Zhongyuan Liu,Yuan Liu,Jinxin Dong,Ligang Liu,Ziqi Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:digital content creation, content creation, fundamental to modern, modern manufacturing, manufacturing and digital

备注: Accepted to SIGGRAPH Asia 2026 Conference Papers. 11 pages, 12 figures. Project page: [this https URL](https://iambrc.github.io/PART-project-page/)

点击查看摘要

Abstract:3D assembly is fundamental to modern manufacturing and digital content creation. In this paper, we present PART, a unified transformer-based framework for 3D part retrieval and assembly: given a target shape and a part library, PART automatically selects the appropriate parts and predicts their 6-DoF poses to reconstruct the target. While prior work has achieved impressive progress on assembling a pre-defined set of parts, this more practical retrieval-based setting remains largely unexplored. The task faces three key challenges: (i) a combinatorially explosive search space that grows exponentially with library size; (ii) variable-length outputs, as different targets require different numbers of parts; and (iii) continuous 6-DoF pose estimation for part assembly. To address these, we formulate retrieval and assembly as a set prediction problem and design a novel transformer-based framework that retrieves parts and regresses their poses with variable-length output. Additionally, we exploit the duality between part pose estimation and target segmentation through joint training and a novel segmentation-enhanced optimization module. Finally, We curate a large-scale dataset of 80K+ shapes, and the results show that PART generalizes to scene layouts, image targets, and real-world scans. Project Page: this https URL.

77. 【2609.19867】Socialized UAV Cross-Task Learning: Towards Cross-Granularity Collaboration through Hierarchical Interaction

链接https://arxiv.org/abs/2609.19867

作者:Xinjie Yao,Ruipu Zhao,Yunqi Zhu,Zhihe Fan,Zhoupeng Guo,Weihao Li,Zhen Wang,Qilong Wang,Pengfei Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:feature sharing, auxiliary supervision, task coupling, coupling, Cross-Granularity Socialized Collaboration

备注: 9 pages, 6 figures

点击查看摘要

Abstract:Joint learning across heterogeneous tasks is often treated as task coupling through feature sharing, distillation, or auxiliary supervision. However, in cross-task learning, mismatched representational and supervisory granularities make such coupling prone to interference, teacher bias, or unidirectional collapse. We argue that cross-granularity learning is fundamentally a problem of hierarchical interaction regulation rather than simple task coupling. This issue is particularly evident in UAV perception, where visual shifts and detection--segmentation objectives naturally form coarse- and fine-grained knowledge sources. To systematically study this problem, we introduce CrossUAV, a UAV benchmark for joint object detection and instance segmentation that provides a unified evaluation platform for cross-granularity task collaboration. To address these challenges, we propose Cross-Granularity Socialized Collaboration (CGSC), a progressive and adaptive framework that regulates when, where, and how tasks exchange information across network hierarchies. CGSC progressively activates cross-task interactions and adaptively adjusts the strength according to task contribution, suppressing harmful interference while exploiting complementary coarse- and fine-grained structures. Extensive experiments demonstrate consistent improvements on both tasks, validating hierarchical dynamic interaction as an effective mechanism for cross-granularity collaboration.

78. 【2609.19863】Feeling Terrain Before Crossing: World Models for Off-Road Navigation

链接https://arxiv.org/abs/2609.19863

作者:E-In Son,Dong-Wook Kim,Ji-Hoon Hwang,Kangsun Lee,Jisung Bae,Jung-Taak Kim,Seung-Woo Seo

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:candidate action sequence, action sequence produces, Navigation world models, candidate action, action sequence

备注: 8 pages, 6 figures

点击查看摘要

Abstract:Navigation world models plan by foresight, predicting the future that each candidate action sequence produces and selecting the best, rather than mapping observations to actions directly. Unlike urban settings where a predicted scene is a sufficient proxy, off-road navigation hinges on the robot--terrain interaction, so the prediction must cover not only what the camera will see but what the robot will feel. However, existing scene-focused models do not predict how much the robot will slip, tilt or shake along a planned trajectory. Proprioception captures these dynamics directly and, when used as input, improves the prediction of the physical future. We present Feel-WM, the first off-road navigation world model that conditions on proprioception and predicts what the robot will feel alongside what the camera will see. The physical future takes the form of a future proprioceptive state and a failure risk, both learned from the robot's own experience without human labels. The planner rolls out the physical future alongside the scene and weighs the predicted failure risk against goal similarity in a separable score. Experiments on real off-road data and in simulation demonstrate that Feel-WM outperforms visual-only navigation world models in open-loop planning and closed-loop rough-terrain navigation across wheeled and legged platforms. Deployed on a Husky on mountain trails, Feel-WM plans onboard, predicts rough ground ahead and steers around it, completing courses that an end-to-end policy fails.

79. 【2609.19853】PACE: Precise AI Cinematic Expression: A Typed Specification for Script-Grounded Previsualization and Geometric Conformance

链接https://arxiv.org/abs/2609.19853

作者:Bing Duan,Qiang Guo,Linpu Li,Zhijian Mao,Min Zhu,Zhirui Ren,Yiwei Yan,Xi Chu,Xiaoding Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:planning problem, camera, Automatic Drive screenplay, Precise AI Cinematic, Cinematic Expression

备注: 36 pages, 7 figures, 3 tables. Code: [this https URL](https://github.com/StudioPiLabs/pace-core)

点击查看摘要

Abstract:Between a screenplay and a film sits a planning problem that is spatial first: who stands where, and what a camera sees from where it stands. An image diffusion model asked for a shot in free text settles that plan by its own defaults. We present PACE (Precise AI Cinematic Expression), a typed representation for the plan: the screenplay evidence, the characters, props and locations it needs, where each subject stands, and what the camera does. A value is written once at the level it belongs to (script, scene, shot or panel) and inherited below it. A compiler turns the result into both the prompt sent to the diffusion model and a 3D scene built in metres, and a camera solver places the camera so that the declared framing is the framing built. Where a declared value becomes geometry, PACE measures, field by field, how far the compiled camera and the staged render sit from the declaration, rather than asking a model to judge. On the 11-scene Automatic Drive screenplay, every staged single-subject panel places its subject within 1.2% of frame width of its declared position; with two or three subjects one camera pose cannot satisfy every position, and the residual is reported rather than absorbed. On 204 external director-storyboard shots, delivered head height is 1.906 times the staged target from the director's words, 1.733 from the compiled prompt, and 0.955 with the greybox control; the condition that holds framing best draws the described action least. Declaring the pose on 30 shots raises the action drawn from 58.9% to 74.4% without moving the framing. Transitions, fitted motion and human review of the generated panels remain open. Code: this https URL

Comments:
36 pages, 7 figures, 3 tables. Code: this https URL

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

ACMclasses:
I.2.10; J.5

Cite as:
arXiv:2609.19853 [cs.CV]

(or
arXiv:2609.19853v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2609.19853

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
80. 【2609.19840】KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

链接https://arxiv.org/abs/2609.19840

作者:Hyunjung Chung,Unsang Park

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remain largely English, talking face data, datasets remain largely, standard English benchmarks, face datasets remain

备注: 22 pages, 5 figures; includes supplementary material

点击查看摘要

Abstract:High-quality 3D talking face datasets remain largely English- centric, and Korean 3D facial motion data are difficult to combine with standard English benchmarks because of differences in mesh topology, spatial scale, coordinate system, and temporal sampling. We present KoUniTalk, a lightweight articulation-centered Korean-English 3D talk- ing face benchmark that retargets VOCASET and the released Korean speech-based 3D talking face data to a shared mesh topology using de- formation transfer. Rather than proposing a new deformation-transfer algorithm or a full-head identity-preserving avatar dataset, KoUniTalk provides an identity-neutral canonical output space for controlled speech- driven facial articulation training and evaluation across English and Ko- rean. The unified template contains 1,176 vertices and focuses on the mouth and adjacent lower- and mid-face regions, reducing the output dimensionality from 15,069 and 72,147 dimensions to 3,528 dimensions, corresponding to 4.27-fold and 20.45-fold reductions compared with VO- CASET/FLAME and the original Korean mesh, respectively. To exam- ine whether retargeting preserves speech-relevant motion, we evaluate semantic mouth-landmark trajectories, including mouth opening, mouth width, aperture ratio, and mouth-opening dynamics. Since the official test set of the Korean dataset is not publicly released, we additionally define a subject-disjoint Korean benchmark split. The processed matched benchmark contains 22 speakers, 4,978 sequences, and 642,781 frames, enabling Korean-English cross-dataset evaluation of speech-driven 3D fa- cial animation models in a single compact articulation-template space. Source-reported inventory counts are listed separately from these pro- cessed counts

81. 【2609.19815】SnapPhysics: A Physics-Aware Scene Graph from a Single View for Interactive Mixed Reality Scenes

链接https://arxiv.org/abs/2609.19815

作者:Suji Kang,Seok-Young Kim,Young Bin Kim,Taewook Ha,Dieter Schmalstieg,Shohei Mori,Woontack Woo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:framework that reconstructs, training-free framework, estimates their physical, center of gravity, physical properties

备注: Accepted for publication in IEEE ISMAR, 2026

点击查看摘要

Abstract:We propose SnapPhysics, a training-free framework that reconstructs 3D objects and estimates their physical properties such as mass, friction, and center of gravity from a single image. For physically coherent interactions in mixed reality (MR), such properties are as important as geometry. Prior approaches infer them by analyzing object dynamics in video, which is computationally costly, or by querying vision-language models (VLMs) on single images, which lacks geometric grounding and inter-object relationships. We address these limitations by combining instance-level 3D reconstruction and spatial alignment with a physics-aware scene graph that encodes these relationships and per-object metric geometry as structured context for VLM-based property reasoning. Experiments on 3D-FRONT show that SnapPhysics improves scene-level F-Score by 18.6% over the best learning-based method, and on real captured scenes with ground-truth mass, it reduces the mean absolute log difference error (mALDE) by up to 20.5% and improves log-scale correlation ($r^2_{\mathrm{ls}}$) by up to 19.6% over VLM-only estimation. SnapPhysics enables physically interactive MR experiences without manual parameter tuning. Project page: this https URL.

82. 【2609.19812】Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

链接https://arxiv.org/abs/2609.19812

作者:Zhiyun Jiang,Hanyong Wang,Binbin Liang,Yu Xie,Menglong Yang,Wei Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Traditional scene understanding, scene understanding focuses, affirmative information objectively, information objectively present, Traditional scene

备注

点击查看摘要

Abstract:Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder's negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.

83. 【2609.19793】AI Smart Glasses for Wearable Intelligence: From Egocentric Sensing to Agentic Personalization

链接https://arxiv.org/abs/2609.19793

作者:Xu Yuan,Yi Wang,Zhuohang Jiang,Haohao Qu,Yujuan Ding,Shanru Lin,Guoliang Xing,Hongxia Yang,Jiannong Cao,Qing Li,Wenqi Fan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, smart glasses, reshaping smart glasses, advances in artificial, capture and display

备注

点击查看摘要

Abstract:Recent advances in artificial intelligence (AI) are reshaping smart glasses from egocentric capture and display devices into platforms for wearable intelligence. Smart glasses increasingly serve as wearable AI systems that connect first-person observation with real-time assistance under strict form-factor constraints. We frame this transition through the lens of \emph{AI smart glasses} and define them as a system-level concept in which egocentric sensing, resource-aware computing, intelligent reasoning, multimodal interaction, and real-world application constraints are co-designed for personalized assistance in the physical world. To systematically study this perspective, we organize the survey around four connected dimensions. First, we examine the hardware foundation that bounds sensing, computation, feedback delivery, and sustained deployment. Second, we study wearable intelligence, where egocentric signals are transformed into perceptual, contextual, and agentic capabilities. Third, we discuss interaction design, through which users request, receive, correct, and regulate assistance during ongoing activity. Fourth, we analyze application scenarios across healthcare, accessibility, situated learning, daily life assistance, cultural tourism, and industrial support, showing how domain requirements reshape system design and evaluation. We further identify five cross-cutting research challenges for future AI smart glasses: next-generation hardware, trustworthy egocentric intelligence, lifelong personalized memory, proactive intelligence, and embodied foundation models. By centering smart glasses as wearable-intelligence platforms, this survey provides a unified framework for organizing technologies, applications, and open challenges in this emerging area.

84. 【2609.19782】Printing the Underdetermined: Materializing Multi-solutionness in Figurative Paintings

链接https://arxiv.org/abs/2609.19782

作者:Yutao Ming,Teng Xu,Youjia Wang,Yunyang Liu,Fengmin Yang,Fuqiang Zhao,Jingyi Yu,Hua Yang,Yanjun Zhou

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:viewers infer depth, reconstruction pipelines attempt, viewers infer, infer depth, attempt to converge

备注: 10 pages, 5 figures

点击查看摘要

Abstract:Figurative paintings are often approached as if they depict a single recoverable 3D scene: viewers infer depth and occlusion, and reconstruction pipelines attempt to converge to one stable model. We instead foreground multi-solutionness, the non-uniqueness of 3D configurations compatible with a single painted image, and propose a workflow that keeps this non-uniqueness visible and material. Multi-solutionness arises from two sources: unobserved content, where backsides and occluded volumes admit multiple plausible completions, and observed cues, where perspective, shading, and occlusion still underconstrain geometry. When additional views are synthesized by a video generative model without explicit 3D constraints, small frame-level drifts become inevitable rather than exceptional. Our pipeline samples multiple camera-orbit multi-view video sequences from one painting, reconstructs each sequence with 3D Gaussian Splatting into a point-based Gaussian scene representation where density halos and ghosting expose unresolved degrees of freedom, and fabricates these representations as physical artifacts using DreamPrinting. By treating multiple compatible interpretations as explicit outputs rather than residual error, we provide a computational framework for spatial readings of figurative painting that can be inspected, compared, and discussed in both digital and physical form.

85. 【2609.19767】Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

链接https://arxiv.org/abs/2609.19767

作者:Zhiyun Jiang,Hanyong Wang,Binbin Liang,Yu Xie,Menglong Yang,Wei Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:True machine intelligence, machine intelligence requires, intelligence requires transcending, requires transcending passive, transcending passive pixel

备注

点击查看摘要

Abstract:True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbf{S}cene \textbf{N}egation \textbf{U}nderstanding under \textbf{S}afety Cognition (\textbf{SNUS}). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.

86. 【2609.19747】STAR: Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction

链接https://arxiv.org/abs/2609.19747

作者:Wontae Choi,Ki Ryum Moon,Jae Young Lee,Hyung Sup Yun,Il Yong Chun

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:noisy focal stack, ill-posed inverse problem, highly ill-posed inverse, focal stack, inverse problem

备注: 5 pages, 3 figures, 2 tables

点击查看摘要

Abstract:Light field (LF) reconstruction from limited and noisy focal stack (FS) measurements is a highly ill-posed inverse problem. Although the LF-to-FS imaging geometry is fixed for a given optical setup, LF spatial-angular structure---including within-view spatial details, cross-view angular dependencies, and disparity across views---varies across scenes. Consequently, a fixed pre-trained prior may not optimally capture the spatial-angular structure of each test LF. We propose Structure-aware Test-time Adaptation for diffusion-based light field Reconstruction (STAR), the first test-time adaptation framework for reconstructing an LF from FS. For each test LF, STAR freezes a pre-trained diffusion prior and fits three lightweight adapters to the observed FS to jointly adapt the three components of the LF's spatial-angular structure. STAR outperforms existing state-of-the-art methods in both two- and three-focal-sheet settings, with shorter inference times than those with test-time parameter updates.

87. 【2609.19745】Region-Level Policy Optimization for Fine-grained MLLM Perception

链接https://arxiv.org/abs/2609.19745

作者:Yuheng Shi,Xiaohuan Pei,Minjing Dong,Chang Xu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:language-model prefilling costs, tokens inflate vision-encoding, underlying fine-grained perception, Fine-grained visual perception, commonly improved

备注

点击查看摘要

Abstract:Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have different resolution requirements. In a controlled diagnostic, localization tolerates roughly 3 to 4 times stronger token compression than recognition, which motivates localizing from a coarse view and concentrating resolution on the selected evidence. Decoding coordinates with the MLLM can be trained end-to-end from answers, but costs a full model pass per query and depends on grounding ability. A lightweight proposal network distilled from the model's attention is fast, but inherits the noise of its attention targets. The RoI from the proposal network reaches the answer through a discrete region choice, so its faithfulness to the answer cannot supervise the network. We therefore optimize the proposal network with region-level reinforcement learning, which we call Vision-RL2. It treats coherent regions as actions, and a frozen MLLM reader scores each one by how its removal changes the answer likelihood. Complementary subtractive and additive objectives suppress distracting proposals and recover missing evidence, updating only the predictor without region annotations, response sampling, or reasoning trajectories. The refined proposal further enables a sparse encoding that magnifies evidence and excludes background tokens. Across six fine-grained benchmarks and four MLLM backbones, Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens. Code is available at this https URL .

88. 【2609.19740】Federated Learning Framework for Privacy-Preserving Kidney Stone Detection

链接https://arxiv.org/abs/2609.19740

作者:Najiyya Younas,Omar Abdulkader,Yaser Ali Shah,Muhammad Jawad Ikram,Jebran Khan,Amaad Khalil

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:pose severe threats, Recent innovations, centralized data storage, medical data security, Federated Learning

备注

点击查看摘要

Abstract:Recent innovations in deep learning have significantly enhanced the diagnosis of medical images, although they are based on the use of centralized data storage that pose severe threats to patient privacy and medical data security. To address this issue, this research proposes a Federated Learning (FL) model that is coupled with an optimized YOLOv8 network to detect the kidney stones on a computed tomography (CT) image and at the same time, protect privacy of the patients. The suggested system can help various medical organizations to jointly train a common model without exchanging the information about the patients. This is to ensure that data protection laws like GDPR and HIPAA are adhered to. The residual feature fusion and DropBlock regularization among other architectural improvements are also included in YOLOv8 to enhance detection robustness and minimize overfitting. Experimental analysis carried out on a distributed CT dataset demonstrated that the federated YOLOv8 model has a mAP at 50 of 0.733 and is able to keep the data confidential. Moreover, its lean design facilitates fast edge deployment and real-time inference across a clinical setting. Altogether, these findings indicate that Federated Learning is a safe and efficient solution to AI-assisted diagnosis in contemporary healthcare when combined with the use of sophisticated object detection models.

89. 【2609.19729】Recency Forcing: Bridging the Long-Horizon Gap in Autoregressive Video Generation

链接https://arxiv.org/abs/2609.19729

作者:Tri Cao,Hung Nguyen,Phong Nguyen,Khoi Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:memory constraints force, long horizons due, constraints force distant, force distant frames, overlooked train-inference discrepancy

备注

点击查看摘要

Abstract:Autoregressive (AR) video generation degrades over long horizons due to an overlooked train-inference discrepancy we term KV eviction mismatch: models train on short clips where all context frames reside in the KV cache, but at inference, memory constraints force distant frames to be evicted from the KV cache - removing context the model was conditioned on. Rather than simulating eviction via context truncation - which discards temporal information the model still needs and degrades motion coherence - we keep the context but while progressively reducing the influence of distant frames, making their eventual eviction negligible. To guide this design, we introduce the positional response $R( \Delta, \, t_{\text{denoise}})$, a perturbation-based sensitivity measure revealing that context influence decays steeply with temporal distance and varies systematically across denoising steps. Motivated by this analysis, we propose Recency Forcing, which applies a non-positive, timestep-dependent bias, termed Temporal Response Bias (TRB), on pre-softmax attention logits derived directly from $R$, closing the train-inference gap without modifying context length or training objectives. We further introduce Biased Attention Reparameterization (BAR), an exact reformulation that moves the bias outside the softmax, making TRB a standard FlashAttention call at zero overhead. Recency Forcing operates in both training-free mode and training-based mode. Experiments on VBench and VBench-Long demonstrate state-of-the-art long-horizon generation quality at no additional inference cost.

90. 【2609.19719】SeetaPsych v1.0: An Open-source Computer Vision Toolkit for Behavior-based Psychological Measurement

链接https://arxiv.org/abs/2609.19719

作者:Jiabei Zeng,Chiqin Li,Kaizhou Li,Fei Chang,Yong Li,Yuanhao Zhao,Dan Han,Wenqiang Yang,Xilin Chen,Shiguang Shan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Automated visual analysis, Automated visual, based psychological measurement, Automated, based psychological

备注

点击查看摘要

Abstract:Automated visual analysis opens new avenues for behavior--based psychological measurement. Nevertheless, existing technological modules are typically scattered across task specific systems with heterogeneous interfaces and disparate deployment requirements. In this work, we present SeetaPsych v1.0, an open source, unified and extensible computer vision toolkit designed to extract psychologically relevant signals from facial images and/or face based videos. The current release encompasses four major core modules aiming at behavior--based physiological perception: unified face based emotion analysis (simultaneous facial expression recognition, facial action unit detection, and valence--arousal estimation), camera based heart rate estimation, screen point--of--gaze estimation, and scene gaze following. A suite of auxiliary preprocessing modules for human centric visual analysis is also included, comprising face detection, facial landmark detection, and head detection. These functionalities are encapsulated within a modular Pipeline/Runner architecture that automatically resolves attribute dependencies, constructs computation graphs, and support intermediate result sharing among modules. SeetaPsych provides standardized Python APIs to facilitate reproducible, large scale analyses, alongside an interactive WebUI for rapid, code--free method evaluation. Overall, SeetaPsych offers an integrated and accessible visual measurement platform for research in psychology, behavioral science, human computer interaction, and related fields.

91. 【2609.19716】GAPrompt++: Multi-Granular Geometry-Aware Point Cloud Prompt for 3D Vision Model

链接https://arxiv.org/abs/2609.19716

作者:Zixiang Ai,Zhenyu Cui,Yufei Guo,Wenwen Qiang,Lei Chen,Jiwen Lu,Jiahuan Zhou

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:substantially advanced point, point cloud analysis, expensive and storage-intensive, advanced point cloud, substantially advanced

备注: Accepted by TPAMI 2026. Code at [this https URL](https://github.com/PKU-OV3-LAB/GAPromptPlus.git)

点击查看摘要

Abstract:Pre-trained 3D vision models have substantially advanced point cloud analysis, yet adapting them to downstream tasks via full fine-tuning is computationally expensive and storage-intensive. Parameter-Efficient Fine-Tuning (PEFT) offers a promising alternative by reducing both adaptation cost and storage burden. However, existing prompting-based approaches ignore the intrinsic geometric structures of point clouds, thereby limiting their adaptation capability. This limitation stems from their inability to encode both fine-grained geometric cues and coarse-grained structural semantics, as well as failing to propagate such information effectively through the model hierarchy. To address these challenges, we propose GAPrompt++, a multi-granular geometry-aware prompting method that provides richer geometric guidance for efficient 3D task adaptation. Specifically, we introduce a Point Shift Prompter that extracts multi-granular geometric features across different scales, enabling instance-specific geometric adjustments during adaptation. Next, a Keypoint Prompter adaptively generates point-level prompts to highlight local geometric saliency and fine-grained structural details. Furthermore, a Prompt Propagation mechanism injects these multi-granular geometric cues throughout the feature extraction hierarchy, strengthening the ability to capture essential geometric characteristics. Extensive experiments show that GAPrompt++ achieves state-of-the-art performance among prompting-based PEFT methods and even surpasses full fine-tuning across diverse benchmarks, while requiring less than 2\% trainable parameters. In addition, to address the saturation of existing evaluation datasets, we construct two more challenging benchmarks derived from 3D Gaussian Splatting and Multi-View Stereo reconstruction, offering diverse and realistic point cloud scenarios to promote future research.

92. 【2609.19702】Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation

链接https://arxiv.org/abs/2609.19702

作者:Daeun Kim,Junwha Hong,Changhun Oh,Yoonsung Kim,Yoonhyeong Lee,Jongse Park

类目:Computer Vision and Pattern Recognition (cs.CV); Performance (cs.PF)

关键词:Autoregressive image generation, LLM serving infrastructures, transformer-based LLM serving, paradigm for multimodal, compatibility with transformer-based

备注

点击查看摘要

Abstract:Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. However, generating thousands of visual tokens per request makes decoding increasingly bottlenecked by KV cache accesses during attention computation. Sparse attention is particularly attractive for this workload because many visual generation applications tolerate moderate quality degradation in exchange for improved performance and efficiency. While sparse attention has been extensively explored for text-based LLM inference, it remains unclear whether its sparsity assumptions generalize effectively to autoregressive image generation. We present the first systematic characterization of attention sparsity in autoregressive image generation across diverse workloads and representative open-source models. Our analysis reveals several distinguishing properties, including a pronounced prefill-decode asymmetry, strong attention concentration on prompt and local tokens, and a unique diagonal attention sparsity pattern arising from the spatial locality of visual tokens. Motivated by these observations, we propose a diagonal-aware sparse attention mechanism that selectively skips KV entries along the diagonal attention direction within a recent window. Implemented on top of a GPU-based serving system using FlexGen, FlashAttention-2, and custom kernels, our approach achieves up to 3.1x throughput and 1.19x latency improvements with less than 2% quality degradation compared to dense inference.

93. 【2609.19693】IMFD: End-to-end Multi-Face Forgery Detection through Instruction-based Large Vision-Language Models

链接https://arxiv.org/abs/2609.19693

作者:Dasom Choi,Sangjun Moon,Hyeongchan Im,Jaeeon Park,Jingun Kwon,Hidetaka Kamigaito,Taro Watanabe,Manabu Okumura

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:raised significant concerns, significant concerns due, multi-face forgery detector, multi-face forgery, social media

备注: 8 pages, 5 figures, 5 tables. Accepted to Findings of AACL-IJCNLP 2026

点击查看摘要

Abstract:The rapid increase of deepfakes has raised significant concerns due to their spread on social media. Traditional multi-face forgery detectors crop and verify each face independently, ignoring background context and inter-face relationships, which often yields suboptimal performance. To overcome these limitations, we leverage instruction-based Large Vision-Language Models (LVLMs), which can interpret entire images and follow complex textual instructions. We propose a simple yet effective single-stage multi-face forgery detector, called IMFD (Instruction-based Multi-face Forgery Detector), which is trained end-to-end to jointly localize faces and predict per-face forgery labels. Rather than treating face box prediction only as a joint objective, IMFD explicitly integrates predicted face bounding boxes into the instruction as visual cues that enhance instruction grounding and forgery detection. To support the training and evaluation of IMFD, we convert existing multi-face forgery datasets into an instruction-based format. Experimental results and analyses show that IMFD improves multi-face forgery detection by integrating face bounding boxes into the instruction, and consistently outperforms various state-of-the-art methods.

94. 【2609.19683】MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

链接https://arxiv.org/abs/2609.19683

作者:Yuan Liao,Jae-sun Seo

类目:Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV)

关键词:necessitating aggressive, memory bandwidth, devices is severely, severely bottlenecked, bottlenecked by memory

备注: Accepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)

点击查看摘要

Abstract:The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer "microscaling collapse," where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.

95. 【2609.19669】Beyond Patch Removal: Persistent Adversarial Effects in Vision-Language-Action Policies

链接https://arxiv.org/abs/2609.19669

作者:Enhao Wu,Fusen Guo,Yuxin Cao,Ziyang Lyu,Lin Li,Wei Song

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:action corruption, VLA, Adversarial patches, effects, adversarial effects

备注: 8 pages, 2 figures

点击查看摘要

Abstract:Adversarial patches to Vision-Language-Action (VLA) policies can cause both immediate action corruption and persistent state effects that remain after the patch is removed. Existing evaluations largely focus on continuous attacks and do not separate these two effects. We introduce a state-restoration protocol that removes the patch at matched action-chunk boundaries and measures subsequent recoverability under the same remaining step budget. Clean, random-patch, deviation-matched, and fixed-direction controls distinguish adversarial effects from occlusion, action-error magnitude, and directional persistence. We also evaluate a recovery adapter trained on attack-induced states under controlled intervention latency. On OpenVLA-OFT with EDPA attacks, only 36.2% of LIBERO-Long episodes remain recoverable after five chunks, compared with 89.9% and 87.0% for the deviation-matched and fixed-direction controls. Similar persistent effects are observed on autoregressive OpenVLA. The recovery adapter improves recovery from 7.7% to 47.4% at one-chunk latency, but its benefit decreases substantially with delayed intervention. These results show that adversarial effects can persist after patch removal and that timely intervention is critical for recovery.

96. 【2609.19664】VideoResearcher: Self-Improving Tool Design for Long-Video Understanding

链接https://arxiv.org/abs/2609.19664

作者:Dingqiang Ye,Dongdi Zhao,Kaishen Wang,Qingqiao Hu,Jingchen Sun,Yijun Liang,Yuqi Jia,Yiqiao Huang,Yunjie Tian,Jiaxing Zhang,Chuanyang Jin,Ke Zhang,Vishal M. Patel,Di Fu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:made substantial progress, made substantial, substantial progress, long-video understanding, Abstract

备注

点击查看摘要

Abstract:Video agents have made substantial progress in long-video understanding. Yet effective video-agent systems require costly, time-consuming manual design and trial and error. Current self-improvement methods either refine low-impact prompts, recombine predefined micro-tools, or struggle with convergence in harness optimization. To bridge this gap, we target high-impact video-tool with VideoResearcher, a training-free multi-agent framework that autonomously designs, tests, and refines tools for video understanding, like a human researcher. VideoResearcher operates through dual Solving and Evolving loops: it analyzes tool-use trajectories to identify capability gaps, coordinates specialized agents to develop and validate executable tools, and reuses evolved tools to strengthen evidence acquisition in subsequent video reasoning. Through iterative tool refinement and validation, it progressively strengthens evidence acquisition without updating model parameters. VideoResearcher achieves state-of-the-art performance among self-improving agents and approaches the human-designed upper bound, demonstrating a training-free paradigm for long-video understanding that expands agent capabilities through autonomous tool development while reducing costly manual engineering.

97. 【2609.19662】owards Active Cross-View Object Geo-Localization

链接https://arxiv.org/abs/2609.19662

作者:Shunyu Yao,Xiaohan Zhang,Zhuoran Yang,Haoqi Lai,Qi Ming,Xiaoxi Hu,Hui-Liang Shen,Si-Yuan Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Cross-view object geo-localization, Active Cross-View Object, Cross-view object, object geo-localization, fixed query image

备注

点击查看摘要

Abstract:Cross-view object geo-localization (CVOGL) typically assumes a fixed query image, overlooking the ability of mobile agents to actively acquire more informative observations. To address this limitation, we introduce Active Cross-View Object Geo-Localization (ActiveGeo), where an agent sequentially selects new viewpoints and determines when to stop, aiming to improve localization with minimal observations. We further propose ActiveMoPT, an ActiveGeo framework with three-stage training. First, Multi-View Prompt-Preserving Adaptation enables the model to aggregate multiple query views while reusing the initial prompt. Second, Trajectory-Guided Policy Initialization uses supervised agent trajectories to learn viewpoint selection and initial stopping behavior. Third, Cost-Aware Policy Refinement employs GRPO with a gain-cost reward to jointly optimize localization accuracy and observation efficiency. We also construct ActiveGeo-858, a zero-shot test set containing 858 scenes and 1,716 target annotations. Experiments show that ActiveMoPT achieves state-of-the-art performance on MoP-UAV using only 1.45 query views on average, and substantially outperforms previous CVOGL approaches under zero-shot evaluation on ActiveGeo-858.

98. 【2609.19634】Scientific Image Quality Assessment via Multi-modal Retrieval-Augmented Generation

链接https://arxiv.org/abs/2609.19634

作者:Yinuo Zhang,Bingshuo Liu,Zhiying Tu,Dianhui Chu,Qingbin Liu,Xi Chen,Jiang Bian,Xiaoyan Yu,Dianbo Sui

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:image quality assessment, scientific image quality, Retrieval-Augmented Generation, complex scientific images, quality assessment

备注

点击查看摘要

Abstract:This paper proposes a Retrieval-Augmented Generation (RAG) framework for scientific image quality assessment, designed to simultaneously address both the understanding track (SIQA-U) and the scoring track (SIQA-S) of the SIQA challenge. We construct a multimodal index that integrates textual semantics with fine-grained visual features, and develop a multi-route retrieval and fusion mechanism to provide large language models with highly relevant reference cases, thereby enhancing their capability to evaluate complex scientific images. Experimental results demonstrate that the proposed framework effectively aligns with the judgment criteria of human experts. Ultimately, our method achieves 1st place in the SIQA-U track of the SIQA challenge at the ICME 2026 Grand Challenges.

99. 【2609.19631】Instance Segmentation and Fine-grained Classification for Urban Buildings with Adaptive Region Dividing and Spatially-Supervised Contrastive Learning

链接https://arxiv.org/abs/2609.19631

作者:Weiyuan Zhang,Qi Zhang,Hui Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:digital city modeling, Accurate instance-level, essential for digital, digital city, city modeling

备注: 10 pages, 4 figures

点击查看摘要

Abstract:Accurate instance-level and functional understanding of urban buildings in large-scale point clouds is essential for digital city modeling and urban analysis. However, the extensive spatial coverage of urban scenes leads most existing methods to rely on predefined blocks for training and evaluation, although such partitions are rarely available in real-world applications and introduce additional preprocessing while fragmenting complete building structures. To address this issue, we propose an adaptive region-dividing strategy with unified scene-level evaluation. Specifically, the 3D point cloud is projected onto a bird's-eye-view (BEV) plane, where a pretrained segmentation model is used to detect building regions. The detected bounding boxes are then back-projected to the original point cloud to construct structure-aligned adaptive training blocks, enabling semantically guided dynamic partitioning without manual design. Furthermore, beyond instance-level understanding, few methods have explored fine-grained classification for urban buildings, and thus we also put forward a fine-grained classification model for urban buildings with a spatially-supervised contrastive loss. First, for each segmented building, a point transformer classifier jointly encodes its body and local context using geometric, color, and core-context information. Then, the class-balanced weighted cross-entropy is used to alleviate severe class imbalance. The proposed spatially-supervised contrastive loss further enhances inter-class discriminability by assigning greater weight to spatially proximate, same-category buildings, encouraging compact functional representations while separating easily confused categories. Extensive experiments on UrbanBIS and STPLS3D demonstrate the advantages of the proposed method in building instance segmentation and fine-grained classification compared to existing SOTA methods.

100. 【2609.19628】VGGT-GS SLAM: Uncalibrated Monocular Gaussian Splatting SLAM with Feed-Forward Priors

链接https://arxiv.org/abs/2609.19628

作者:Yuhang Han,Hao Wang,Jiaxi Cao,Xingyu Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:present VGGT-GS SLAM, Gaussian Splatting SLAM, Splatting SLAM system, present VGGT-GS, SLAM system designed

备注: 9 pages, 4 figures

点击查看摘要

Abstract:We present VGGT-GS SLAM, a monocular 3D Gaussian Splatting SLAM system designed for uncalibrated videos. Starting from feed-forward VGGT pose and depth priors, our system performs submap differentiable bundle adjustment that jointly refines camera poses and a 3D Gaussian map, while optimizing submap-shared intrinsics and radial--tangential distortion through analytic calibration Jacobians. To improve global consistency, we introduce Gaussian-native alignment (GNA) for camera-anchored scale refinement between sequential submaps and verification of loop-closure candidates. Extensive experiments on standard indoor benchmarks show consistent improvements in localization accuracy and strong rendering quality under uncalibrated settings, establishing a strong baseline for uncalibrated Gaussian SLAM.

101. 【2609.19592】Selective Cotton Boll Localization for Robotic Harvesting: Evaluation of Deep Learning Vision Models Under Field Conditions

链接https://arxiv.org/abs/2609.19592

作者:Thevathayarajh Thayananthan,Xin Zhang,Isuru Laddusinghe Badu,Jonathan Harjono,Glen C. Rains,Beiwen Li,Leonardo M. Bastos,Nuwan K. Wijewardane,Vitor S. Martins

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:study developed, segmentation, Model, models, cotton

备注: 27 Pages, 19 Figures, 15 Tables

点击查看摘要

Abstract:This study developed and evaluated a deep-learning-based perception framework for selective robotic cotton picking. The dataset contained 1,008 annotated field images collected using three cameras under varying natural lighting and weather conditions. Object-detection models from the YOLOv8 through YOLOv13 families were evaluated using their default configurations, while segmentation performance was assessed using YOLOv8-seg, YOLOv11-seg, YOLOv12-seg, the Segment Anything Model (SAM), SAMv2.1, FastSAM, and Grounded-SAM with the Recognize Anything Model (RAM). Among the detection models, GELAN-s achieved the most favorable balance between mean average precision (mAP) and inference speed, obtaining an mAP of 86.1%, precision of 81.6%, recall of 76.6%, and an F1-score of 79.0%, with an average inference time of 42.3 ms per image. Among the direct segmentation models, YOLOv12-m-seg provided the most favorable balance between AP@0.5 and FPS, achieving a segmentation AP@0.5 of 83.7% with an inference time of 20.4 ms per image. In the detection-prompted segmentation approach, bounding-box prompts generated by GELAN-s improved the localization of cotton bolls for SAM and SAMv2.1, while SAMv2.1 Tiny consistently outperformed FastSAM and Grounded-SAM with RAM. In the area-based evaluation against manually annotated segmentation masks, YOLOv12-m-seg achieved an $R^2$ value of 0.966, compared with 0.860 for GELAN-s + SAMv2.1 Tiny. Field experiments conducted using a UR5e robotic manipulator, a custom end-effector, and a ZED2i stereo camera further validated the effectiveness of the YOLOv12-m-seg model for real-time cotton boll detection, segmentation, and selective picking under varying confidence levels. These results demonstrate that YOLOv12-m-seg provides an efficient perception model for robotic cotton harvesting and has strong potential for field deployment.

102. 【2609.19555】A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

链接https://arxiv.org/abs/2609.19555

作者:Khang Nguyen Quoc,Minh-Phuoc Tran,Gia-Han Truong,Luyl-Da Quach

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Artificial intelligence, jointly interpreting visual, multi-modal models capable, textual information, classifiers to multi-modal

备注: In submission to Computers and Electronics in Agriculture Journal

点击查看摘要

Abstract:Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical deployment in precision agriculture remains limited because most existing approaches treat disease understanding as isolated prediction tasks, failing to capture the complementary relationships among symptom recognition, severity assessment, and question-driven diagnostic reasoning. In tomato pathology, accurate interpretation of diseased leaves requires more than label prediction; it demands integrating visual symptoms with semantic context to support a comprehensive and explainable understanding. Here, we present SOLAR, a multimodal generative model that understands tomato disease spanning six question-answering tasks. SOLAR learns to align visual features with task-aware language representations by Fusion Expert module based on mixture-of-expert, enabling it to generate contextually relevant answers across diverse diagnostic tasks. By formulating tomato disease analysis as a generative Visual Question Answering (VQA) task, SOLAR provides a flexible framework that supports multi-task inference within a single model while improving performance and cross-task knowledge sharing. We evaluate SOLAR on $41,677$ images, including $216,209$ Question-Answering (QA) pairs to understand tomato leaf disease under both closed and open-ended QA settings. Experimental results show that SOLAR consistently outperforms state-of-the-art vision-only, vision-language, and task-specific models across all tasks, demonstrating superior accuracy, robustness, and multimodal reasoning. These findings highlight the potential of generative multimodal modeling as an effective direction for understanding of plant disease. The code for this study is available at this https URL.

103. 【2609.19554】VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

链接https://arxiv.org/abs/2609.19554

作者:Zhongbo Zhang,Jiayi Jin,Yifan Wang,Zaibin Zhang,Haiwen Diao,Lijun Wang,Huchuan Lu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:Spatial intelligence requires, describing object locations, intelligence requires, Spatial intelligence, common spatial frame

备注

点击查看摘要

Abstract:Spatial intelligence requires more than describing object locations. Under incomplete observation, models must identify and acquire missing evidence, interpret it in a common spatial frame, and act on it. We introduce VA-Bench to evaluate the complete observe-reason-act-revise loop. General-purpose MLLMs learn procedural context from RGB-only demonstrations, actively select camera viewpoints, issue metric Cartesian commands, and revise them from execution feedback. Models receive no privileged object poses, oracle trajectories, or learned action heads. A fixed model-agnostic controller executes only model-specified targets. VA-Bench contains 14 base task families (11 single-arm and three dual-arm), seven held-out geometry/layout variants, and a long-horizon five-object composition track. We evaluate 12 primary model conditions in three independent runs over the same 20 physically verified seeds per base task, reporting terminal success, nine trajectory-level behavioral diagnostics, and subtask progress. First, the best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run. Its three-run macro-average task success is only 53.93+/-3.17%. Second, active camera control significantly improves task success over passive multi-view observation. In one matched comparison, success rises from 27.86% to 57.50%. Third, held-out geometric transfer can reduce task success by over 30 percentage points. No model completes a strict long-horizon episode, despite substantial partial progress. VA-Bench thus tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.

104. 【2609.19542】PerSeM: Persistent Semantic Memory for Long-Horizon Open-Vocabulary UAV Mapping

链接https://arxiv.org/abs/2609.19542

作者:Saurbh Singh Jamwal,Ganesh Ramakrishnan

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Open-vocabulary segmentation enables, segmentation enables rich, open-vocabulary UAV mapping, long-horizon open-vocabulary UAV, enables rich semantic

备注

点击查看摘要

Abstract:Open-vocabulary segmentation enables rich semantic perception for UAVs, but frame-wise predictions can remain temporally inconsistent across repeated observations and changing viewpoints. We present PerSeM, a training-free persistent semantic memory framework for long-horizon open-vocabulary UAV mapping. PerSeM associates frame-wise semantic observations with persistent world-space voxels and constructs a majority-based semantic memory, which is conservatively refined through history-preserving spatial refinement, trust-aware replay, and context-guided verification. Experiments on the Forest and UAVScenes benchmarks show that persistent 3D memory provides substantial gains in semantic correctness and temporal stability over frame-wise predictions. Beyond this strong persistent-memory baseline, PerSeM provides consistent additional improvements, improving both semantic accuracy and temporal stability across all five evaluated UAVScenes sequences. Analysis using regions identified independently of the final PerSeM predictions further shows that these gains are concentrated in semantically difficult and temporally unstable regions, where majority-based memory is most likely to remain uncertain. These results demonstrate that persistent 3D aggregation provides a strong foundation for long-horizon semantic mapping, while conservative refinement of uncertain memory states can provide additional improvements without retraining or additional neural-network inference.

105. 【2609.19518】AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend

链接https://arxiv.org/abs/2609.19518

作者:Hengyi Wang,Lourdes Agapito

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:single consumer-grade GPU, real-time monocular SLAM, monocular SLAM system, SLAM system capable, reconstructing kilometer-scale trajectories

备注: Project page: [this https URL](https://hengyiwang.github.io/projects/amber-slam)

点击查看摘要

Abstract:We present AMB3R-SLAM, a real-time monocular SLAM system capable of reconstructing kilometer-scale trajectories over 10k frames on a single consumer-grade GPU. Our model couples a lightweight front-end for low-latency online tracking with a hierarchical backend that progressively enforces local, mid-level, and global consistency. By avoiding bundle adjustment that relies on the static world assumption, our system naturally handles complex dynamic scenes out of the box. Furthermore, we demonstrate that our method can be extended to leverage stereo, RGB-D, and LiDAR as additional inputs. AMB3R-SLAM achieves strong camera tracking performance across 9 datasets, reducing the absolute trajectory error (ATE) of previous state-of-the-art methods on VBR and Oxford Spires by over 70%. With additional LiDAR input, our model further reduces ATE to sub-meter level on KITTI and VBR datasets.

106. 【2609.19483】SCOUT: Sim-to-Real Text-Based Person Retrieval by Embedding-Space Prediction over Frozen Video Features

链接https://arxiv.org/abs/2609.19483

作者:Abdarahmane Traoré,Andy Couturier,Éric Hervet

类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)

关键词:Text-based person retrieval, synthetic training data, Text-based person, frozen text encoder, synthetic training

备注: 16 pages, 4 figures, 3 tables. Accepted at the ECCV 2026 Workshop on AI City Challenge (Track 4). Code and annotations: [this https URL](https://github.com/abtraore/SCOUT-ECCV)

点击查看摘要

Abstract:Text-based person retrieval under a sim-to-real gap (synthetic training data, a real-image gallery) is usually tackled with costly fine-tuned cross-encoders. We ask whether a frozen-encoder system can compete. We present SCOUT, which casts cross-modal retrieval as prediction in embedding space. A trainable predictor maps the patch tokens of a frozen video encoder into the embedding space of a frozen text encoder under a bidirectional InfoNCE objective, and no encoder is fine-tuned in the base model. The video encoder is V-JEPA, the text encoder is EmbeddingGemma, and the predictor is initialized from a Qwen3.5-0.8B decoder. We make three findings. First, the best frozen text encoder is simply the one whose geometry best matches the video features. A training-free alignment score ranks three candidate text encoders in the same order as their retrieval accuracy on our held-out split (Spearman $\rho = 1.0$); a fourth, LLM-based encoder shows the rule is metric-dependent, holding for a neighborhood-overlap score ($\rho = 0.8$) but not for a linear probe ($\rho = -0.2$). Second, two precision-targeted levers, parameter-efficient ExPLoRA adaptation of the video encoder and a training-free attribute-decomposed reranker built on a vision-language model, improve the top-rank precision that otherwise limits the frozen system, adding 2.2 points of leaderboard R@1. Third, a local-versus-public calibration study explains which interventions transfer to the real domain. On AI City Challenge 2026 Track 4 the full retrieve-fuse-rerank system reaches 84.25 mAP@10 on the final leaderboard, while a single frozen model submitted alone reaches 60.63. Our trained components cost about 95 GPU-hours. CMP, the dataset authors' fine-tuned cross-encoder that trains for sixteen GPU-days, is one fusion member of the full system, not an alternative. Code and annotations: this https URL

107. 【2609.19463】ParticleSplat: Self-supervised Object-centric Latent Particle Splatting

链接https://arxiv.org/abs/2609.19463

作者:Lyuxing He,Daniel Guo,Elizabeth Terveen,Deepak Pathak,David Held,Tal Daniel

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:representing semantic entities, Gaussian Splatting, particles representing semantic, self-supervised object-centric learning, object-centric learning method

备注: Project page: [this https URL](https://lyuxinghe.github.io/ParticleSplat-website/)

点击查看摘要

Abstract:We present ParticleSplat, a self-supervised object-centric learning method that decomposes scenes into a set of latent ''particles'' representing semantic entities through feedforward 3D Gaussian Splatting. Building on the Deep Latent Particles (DLP) framework, which represents images as a set of particles with attributes such as position, scale, and visual appearance, we address a key limitation of DLP: its inherently 2D nature, which prevents explicit 3D spatial and geometric reasoning that are critical for downstream tasks such as robotic manipulation. Leveraging the structural similarity between latent particles and 3D Gaussian primitives, we introduce a 3D latent particle space trained with a novel view synthesis objective. Our model jointly encodes multiple views with camera poses into a shared 3D object-centric latent space, then transforms particles into particle-aligned 3D Gaussians whose composition reconstructs the full scene. On simulated and real-world datasets, we show that this formulation inherently learns object masks without supervision and supports controllable 3D scene editing, such as moving objects by modifying particles in the latent space. We further establish that the learned 3D representation improves downstream performance on robotic manipulation tasks.

108. 【2609.19451】Efficient Unified Multimodal Understanding (EUMU): Winning Solution for the MUMU Track at the 8th LSVOS Challenge

链接https://arxiv.org/abs/2609.19451

作者:Dayoung Kil,Seong-heum Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Unified Multimodal Understanding, Mobile Unified Multimodal, Efficient Unified Multimodal, Mobile Unified, Multimodal Understanding

备注

点击查看摘要

Abstract:The Mobile Unified Multimodal Understanding (MUMU) Challenge requires a single efficient model to jointly perform multi-concept image tagging, open-vocabulary object detection, and image captioning. We present Efficient Unified Multimodal Understanding (EUMU), the winning solution for the MUMU Track of the 8th LSVOS Challenge. EUMU builds on a shared pretrained multimodal model, using its prompt-based capabilities for detection and captioning and training lightweight heads on shared visual features to predict quality, scene, and event tags. Rather than treating the three tasks independently, EUMU applies task-aware inference refinement by reusing task outputs as cross-task cues. For detection, caption cues help recover objects missed by the initial detection. For captioning, detection cues help refine the caption to better reflect the detected objects. For tagging, image statistics refine quality predictions, while caption and detection cues refine scene and event predictions. This design unifies all three tasks within a single model while satisfying the challenge's resource constraints. EUMU contains 239.169M parameters, requires 23.947 GFLOPs, uses 4.5 GB of peak inference memory, and achieves a final challenge score of 17.3409. Code and models are available at this https URL.

109. 【2609.19445】From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning

链接https://arxiv.org/abs/2609.19445

作者:Pan Wang,Siwei Song,Hui Ji,Siqi Cao,Heng Yu,Zhijian Liu,Huanrui Yang,Yingyan Celine Lin,Beidi Chen,Mohit Bansal,Xiaoming Liu,Pengfei Zhou,Ming-Hsuan Yang,Tianlong Chen,Jingtong Hu

类目:Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:surfaced formidable bottlenecks, pivotal research frontier, bottlenecks in computation, catalyzing the rise, Efficient Multimodal Learning

备注: TMLR

点击查看摘要

Abstract:The rapid expansion of multimodal models has surfaced formidable bottlenecks in computation, memory, and deployment, catalyzing the rise of Efficient Multimodal Learning (EML) as a pivotal research frontier. Despite intensive progress, a cohesive understanding of what, how, and where efficiency is manifested across the learning stack remains fragmented. This survey systematizes the EML landscape by introducing the first structured, model-to-system taxonomy. We distill insights from over 300 seminal works into three hierarchical levels--model, algorithm, and system--addressing architectural parsimony, execution refinement, and hardware-aware orchestration, respectively. Moving beyond a purely categorical review, we offer a methodological synthesis of the vertical synergies between these layers, elucidating how cross-layer co-design contributes to the fundamental "Efficiency-Utility-Privacy" trade-off. Through an integrative case study of Multimodal Large Language Models (MLLMs), we trace the field's evolutionary trajectory from initial structural adjustments to modern full-stack resource orchestration. Furthermore, we provide a holistic discussion and application-specific optimization blueprints for diverse domains and posit a paradigm shift toward self-regulating intelligence, where efficiency is an intrinsic, emergent property of the model's fundamental design rather than a post-hoc constraint. Finally, we present open challenges and future directions that will define the trajectory of EML research. This survey establishes a structured framework for multimodal systems that are not only high-performing and generalizable but natively efficient and ready for ubiquitous deployment. A continuously updated version is available at this https URL.

110. 【2609.19444】Seeing Abnormal from Normal: Glomerular Abnormality in Representations of Normal Renal Morphology

链接https://arxiv.org/abs/2609.19444

作者:Greta Hasko,Rachit Saluja,Tianyu Shi,Leiyue Zhao,Yuechen Yang,Daniel Reisenbuechler,Tianyuan Yao,Zhenhao Guo,John Cannon,Haichun Yang,Yuankai Huo,Yuling Chi,Lorraine Gudas,Mert R. Sabuncu,Yihe Yang,Ruining Deng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:segmental glomerulosclerosis, Fine-grained evaluation, global and segmental, distinguish normal glomeruli, Fine-grained

备注

点击查看摘要

Abstract:Fine-grained evaluation of glomerular pathology must distinguish normal glomeruli from abnormalities such as global and segmental glomerulosclerosis, obsolescent, ischemic, solidified, disappearing, and atubular glomeruli. Supervised classification requires labeled examples of every category, which is impractical when subtypes are rare or absent from the training cohort. One-class anomaly detection offers an alternative by modeling normal data and scoring deviations, allowing previously unseen abnormalities to be detected. We use the frozen residual U-Net backbone of Omni-Seg, pretrained to segment structurally normal renal primitives without abnormal-subtype labels. We propose NoRDeC (Normal-Reference Detection and Characterization), a framework combining Mahalanobis normal-reference scoring with layer-wise representation analysis to determine whether and where glomerular pathology is encoded, how spatial aggregation affects detection, and whether abnormalities alter inter-layer relationships differently. Using glomerular images from two institutions, we evaluate backbone layers and aggregation strategies, compare NoRDeC with PaDiM and PatchCore, and analyze representations using centered kernel alignment (CKA). Layer 4 with Center-70 aggregation achieved a pooled AUROC of $0.926\pm0.013$. NoRDeC achieved the highest AUROC in six of seven abnormality categories and in the pooled analysis, while CKA suggested subtype-dependent changes in inter-layer relationships not captured by anomaly scores alone. The normal-reference model is fitted using only normal glomeruli; abnormality labels are used for configuration selection, evaluation, and grouping in the representation analysis. These results show that a frozen renal feature extractor can support both detection and representation-level characterization of glomerular abnormalities without using abnormal examples to fit the detector.

111. 【2609.19421】RGS: Reflection-aware Gaussian Splatting via Learning Geometry Continuity for Reflective Objects

链接https://arxiv.org/abs/2609.19421

作者:Xiaobiao Du,Yida Wang,Cheng Bi,Kun Zhan,Xin Yu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:explicit Gaussian representation, Gaussian Splatting, Reflection-aware Gaussian Splatting, Gaussian Splatting methods, Gaussian representation

备注: Project Page: [this https URL](https://xiaobiaodu.github.io/reflectivegs/) Published in ICRA2026

点击查看摘要

Abstract:Gaussian Splatting has significantly improved the quality of novel view synthesis with explicit Gaussian representation. However, we observed that existing 3D Gaussian Splatting methods (3DGS) often suffer from surface collapse issues on reflective regions, and thus produce inferior geometry and low-quality specular. In this work, we propose a physically-based deferred rendering framework, named Reflection-aware Gaussian Splatting (RGS), that can accurately model specular regions and improve novel view synthesis performance. Specifically, we found that a powerful 3D foundation model can provide a strong 3D geometric prior to foster correct geometric modeling. Based on this, we propose a cross-view shape consistency regularization to regularize the geometry surface with the large model prior and cross-view constraints. In this manner, our RGS can produce smoother geometric surfaces on reflective regions while reducing geometric hollows. To further improve rendering results on reflective regions, we present a reflection-aware densification strategy that is designed to capture specular variations across various views. With this strategy, our RGS is able to render novel views of objects in higher quality. Extensive experiments demonstrate our method consistently renders high-quality reflective objects, achieving state-of-the-art performance.

112. 【2609.19393】WZPlanner: Safe End-to-End Path Planning for Autonomous Driving in Work Zones

链接https://arxiv.org/abs/2609.19393

作者:Nishad Sahu,Changzhong Qian,Guangzhou Cai,Shounak Sural,Ragunathan(Raj)Rajkumar

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:challenging autonomous vehicle, temporary traffic controls, alter lane geometry, work zone boundaries, Work zones alter

备注

点击查看摘要

Abstract:Work zones alter lane geometry through temporary traffic controls and closures that may be absent from on-board maps, challenging autonomous vehicle (AV) perception and planning. Generalization is also limited by scarce public datasets with structured geometric supervision. We present WorkZonePlan, a dataset comprising 149K+ synthetic and 5K+ real-world multimodal samples with 3D annotations for lane boundaries, work zone boundaries, and driving trajectory options. It also provides 76 closed-loop CARLA scenarios replayed under three weather conditions, yielding 228 Bench2Drive-format evaluation routes. We introduce WAVE (Work-zone-focused AV data generation in Virtual and rEal Environments), a semi-automated pipeline for creating the dataset, and BoundaryFormer (BF), a transformer-based model that jointly predicts lane and work zone boundary polynomials and driving trajectories. BF uses slot attention for boundary prediction. Ablations show that a separate trajectory decoder using boundary slot features substantially improves trajectory prediction over a slot-attention-only approach. Building on this finding, BF++ offers Camera and Camera+LiDAR variants with metric ground-plane encoding, typed boundary/trajectory queries, long-range point anchors, image-space curve refinement, and conservative gated LiDAR fusion. On the 211 routes common to all four models at the evaluation freeze, BF++-Camera and BF++-Camera+LiDAR achieve Driving Scores of 63.0 and 64.4, respectively, compared with 59.3 for SimLingo and 26.1 for TransFuser++ (TF++). BF++ is 40 times smaller than SimLingo and more than 10 times smaller than TF++, while achieving higher Driving Scores. These results support jointly predicting lane boundaries, work zone boundaries, and driving trajectories as a promising direction toward safer AV operation in work zones. Code and dataset: this https URL.

113. 【2609.19384】Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models

链接https://arxiv.org/abs/2609.19384

作者:Badri N. Patro,Vijay S. Agneeswaran

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM)

关键词:Scaling deep learning, exponential training costs, faces critical bottlenecks, deep learning faces, learning faces critical

备注

点击查看摘要

Abstract:Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogeneous merging setting that retains both architectures while aligning parameter groups by semantic role. Our proposed Riemannian--Lorentz Parameter Fusion (RLPF) method projects aligned groups to common coordinates, lifts selected coordinates to the Lorentz hyperboloid model of hyperbolic space, computes a regularized geodesic barycenter, and decodes the result into the two branches. A learned gate then combines branch logits for each input. Component groups use fixed curvature values, with normalization parameters treated as Euclidean. In the results available in this manuscript, the fine-tuned system obtains 82.37\% on CIFAR-10, 75.04\% on Oxford-IIIT Pet, and 78.58\% top-1 accuracy on ImageNet-1K; the corresponding best-parent accuracies are 76.54\%, 71.42\%, and 76.42\%. On ImageNet-1K, the reported pre-fine-tuning initialization reaches 77.80\%. These results support further study of geometry-aware heterogeneous fusion, but not a training-free single-checkpoint merge: RLPF is a two-branch hybrid whose gate and reported final models are trained.

114. 【2609.19377】LinePilot Digitizer: Line-Plot Recovery with Manual and Automatic Calibration

链接https://arxiv.org/abs/2609.19377

作者:Fengbo Ma,Rayan Akhtar,Aakash H. Joshi,Xiaoting Li,Haijian Sun,Zhen Xiang,Xianyan Chen,Yiping Zhao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Recovering numerical series, line plots requires, plots requires accurate, requires accurate axis, reliable curve extraction

备注

点击查看摘要

Abstract:Recovering numerical series from line plots requires accurate axis calibration and reliable curve extraction. We present LinePilot Digitizer (LinePilot), which combines continuous color-based curve recovery with three calibration modes: LinePilot (standard), LinePilot (enhanced), and LinePilot (OCR). We also introduce DigitizerBench, the first dedicated benchmark for systematically evaluating digitizer performance, using an orthogonal design spanning signal, rendering, and plot-structure factors with complementary automatic and human-guided evaluations. We evaluate performance using failure-penalized capped normalized root-mean-square error (FPC-NRMSE), which assigns unit loss to missing, unusable, or catastrophically inaccurate outputs. On DigitizerBench-Full, LinePilot (OCR) achieves the lowest mean FPC-NRMSE (0.672) and highest trusted usability (38.2%) among the tested automatic pipelines. On DigitizerBench-Lite, LinePilot (enhanced) achieves the lowest mean FPC-NRMSE (0.081), 100% output success, and highest trusted usability (93.3%). The orthogonal benchmark design further enables factor analysis to identify the factors that most significantly affect digitizer performance. Together, the three calibration modes provide a practical trade-off between automation, user control, and accuracy within a shared curve-recovery workflow.

115. 【2609.19358】Open-vocabulary 3D object detection with promptable segmentation

链接https://arxiv.org/abs/2609.19358

作者:Ömer Faruk Deniz,Mustafa Taha Koçyiğit

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Three-dimensional object detection, corpora of human-annotated, autonomous driving, driving is dominated, trained on large

备注: 18 pages, 6 figures, 12 tables

点击查看摘要

Abstract:Three-dimensional object detection for autonomous driving is dominated by detectors trained on large corpora of human-annotated 3D boxes. Such a detector learns a fixed category list, and everything outside it is invisible. This paper asks whether the task can be solved training-free and open-vocabulary. A promptable segmentation model (SAM3), queried with class names as text prompts, supplies instance masks in the vehicle's six surround-view cameras, and the masks are turned into metric 3D boxes using the geometry of the scene. The core is a controlled three-stage comparison on nuScenes in which 2D detection is held fixed and only the source of 3D geometry changes. Geometry predicted from images alone reaches 0.183 mean average precision (mAP) under the official protocol; fitting boxes from raw LiDAR points inside the same masks with training-free rules reaches 0.298 mAP / 0.348 nuScenes detection score (NDS) at zero labeling cost; borrowing supervised box geometry at inference time lifts the same detections to 0.413 mAP / 0.555 NDS, which locates the pipeline's largest deficit in measurement precision rather than 2D detection, while class confusion and confidence calibration survive that substitution. Reversing the direction, a three-state camera-witness rule built from the same masks improves a supervised LiDAR-only detector from 0.596 to 0.630 mAP, roughly half the gain of fully supervised camera fusion, with no training. A coverage analysis shows that SAM3 finds 84% of in-range objects with a correctly named mask; the classes that fail in the official metric are misnamed or geometrically unforgiving, not unseen.

116. 【2609.19354】Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

链接https://arxiv.org/abs/2609.19354

作者:Henry O. Velesaca,David Freire-Obregon,Luigi Miranda,Abel Reyes-Angulo

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:Automated action quality, challenging task due, action quality assessment, Olympic sports remains, Automated action

备注

点击查看摘要

Abstract:Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub this https URL diving judge vlm

117. 【2609.19236】RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

链接https://arxiv.org/abs/2609.19236

作者:Fangjie Li,Mai Bui,Charan Mohan,Michael Miga,Matthieu Chabanas,Nicholas Kavoussi,Jie Ying Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:ureteroscopic kidney stone, kidney stone surgeries, Incomplete navigation, repeat interventions, anatomy during ureteroscopic

备注

点击查看摘要

Abstract:Objective: Incomplete navigation of anatomy during ureteroscopic kidney stone surgeries can contribute to repeat interventions. While skilled surgeons have lower reintervention rates, there are no objective metrics to quantify scope-navigation performance to evaluate when a trainee becomes skilled. This work aims to recover ureteroscope trajectories from endoscopic video and derive navigation metrics to quantify differences in skill. Methods: We propose RAUL, a reference-assisted reconstruction framework for recovering ureteroscope trajectories from ureteroscope videos only in phantoms. For each phantom, we use a slow, high-quality reference exploration video to generate a reference reconstruction. We localize subsequent exploration videos against this reference. We evaluate localization accuracy against electromagnetically tracked scope pose. We compute navigation metrics from phantom exploration trajectories to compare surgical residents across experience levels. Results: The proposed reference-assisted framework achieves a mean translation root mean square error of $0.5 \pm 0.1$ mm across 9 phantoms. Compared to standard Structure-from-Motion (SfM), the proposed pipeline increases frame-wise localization coverage from $50.5 \pm 14.9\%$ to $86.1 \pm 7.2\%$ of all video frames. The reconstructed trajectories revealed significant differences between high- and low-experience trainees in established navigation metrics. Conclusion: RAUL enables substantially more complete recovery of ureteroscope trajectories from videos compared to standard SfM pipelines, enabling trajectory-based skill assessment without additional tracking equipment. Significance: To the best of our knowledge, this is the first use of video-only recovery of ureteroscope trajectories without external tracking sensors for skill assessment, supporting scalable automated assessment of ureteroscopy navigation skill.

118. 【2609.19230】Open ultrasound foundation model for robust segmentation and clinical measurement across heterogeneous settings

链接https://arxiv.org/abs/2609.19230

作者:Chao Qin,Fahad Shahbaz Khan,Salman Khan,Sarim Ather,Siddiq Anwar,Rao Muhammad Anwer,Shadab Khan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:imaging modality worldwide, widely deployed imaging, deployed imaging modality, narrow single-task models, modality worldwide

备注: The PDF includes the Supplementary Information

点击查看摘要

Abstract:Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when device, operator, or anatomy changes. Here we present SonoCorpus, an open resource unifying 456,963 images and 1,626,085 expert masks from 53 public datasets spanning 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on it. Across fifteen evaluation datasets introducing new organs, devices, operators, and geographies, SonoBase outperforms SAM2, MedSAM2, and the concept-promptable MedSAM3 on every dataset and matches per-dataset specialist models trained on the same data; on fully external data it exceeds the accuracy these baselines achieve on their own in-distribution benchmarks. Ejection fraction derived from its segmentations falls within inter-observer variability (6.63\% error), with fewer misclassifications at the defibrillator-candidacy threshold than either promptable baseline (13\% versus 18--42\%); fetal head-circumference (1.81~mm) and gestational-age (1.2 days) errors fall below inter-observer variability. Where a baseline fails outright, one in four test cases, SonoBase recovers a usable segmentation in 81\% of them, including on handheld probes operated by minimally trained users in two low- and middle-income countries (Sierra Leone and Tanzania). Five labeled examples can help the model adapt to a new setting, and the identical training protocol transfers well to newer models such as SAM3, locating the advantage in ultrasound-specific pretraining rather than any single architecture. To ensure reproducibility and enable the community to build on SonoBase as a platform, we release all checkpoints, optimizer states, data-split indices, deduplication hashes, and starter code.

119. 【2609.19216】4D Radar Perception Algorithms for Autonomous Driving: A Review

链接https://arxiv.org/abs/2609.19216

作者:Xumin Wu,Jun Zhou,Jilin Mei,Chen Min,Yu Hu

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)

关键词:dynamic scene reconstruction, millimeter-wave radar perception, occupancy prediction, recent years, extending from signal

备注: 12 pages, 9 figures, 5 tables. Submitted to IEEE Sensors Journal

点击查看摘要

Abstract:Research on 4D millimeter-wave radar perception algorithms has flourished in recent years, extending from signal processing and object detection to semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. This review organizes the field according to the evolution of perception tasks and algorithms. It first introduces radar fundamentals, data representations, and quality-enhancement methods, and then reviews object-level perception, motion and localization, local and dense spatial perception, and dynamic scene understanding. Across these directions, we compare radar-only learning, multimodal fusion, and cross-modal supervision and knowledge distillation. Particular attention is paid to how elevation, Doppler measurements, and radar physical priors are exploited across tasks. We further summarize the task coverage, input data, annotations, and evaluation protocols of existing datasets, clarifying the empirical support for different research directions. Finally, we discuss the common challenges and future directions of 4D radar perception for autonomous driving. This review provides a task-oriented perspective on the transition from sparse object perception to dynamic spatial understanding.

120. 【2609.19148】Modality Discrepancy Transformer for Ambivalence and Hesitancy Recognition

链接https://arxiv.org/abs/2609.19148

作者:Shiyu Luo,Yu Wang,Jiawen Huang,Zhaoxiang Xiao,Chenxi Huang,Qi Zhang,Bin Liu

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:individuals express contradictory, express contradictory signals, Ambivalence and hesitancy, linguistic channels, affective states

备注: 10 pages

点击查看摘要

Abstract:Ambivalence and hesitancy (A/H) are affective states in which individuals express contradictory signals across facial, vocal, and linguistic channels. Automatically recognising A/H in clinical videos requires detecting cross-modal disagreement -- the signal that standard fusion methods suppress. Based on the conflict-aware multimodal fusion framework of Bekhouche et al., we present the Modality Discrepancy Transformer (MDT). MDT enriches the original 6-token design to a 9-token representation comprising three modality embeddings, three absolute-difference features, and three Hadamard-product discrepancy features learned through linear projections. These nine tokens undergo Transformer self-attention, with FiLM-based text-conditioned modulation and LoRA fine-tuning as core architectural components. A text-guided late fusion branch blends a text-only auxiliary head with the full multimodal output at inference. On the BAH dataset from the 3rd ABAW Challenge, MDT achieves 0.7408 Macro F1 on the labelled test split and 0.7368 on the private leaderboard, outperforming the strongest published baseline by over 10 points while training in under 20 minutes on a single GPU.

121. 【2609.19755】HyperAMS-Net: Adaptive Multi-Scale Spatial Hypergraph Network for Brain Disorder Classification

链接https://arxiv.org/abs/2609.19755

作者:Proloy Kumar Mondal,Md Kamran Hussin Chowdhury,Hoi Leong Lee

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:data remains challenging, substantial inter-subject heterogeneity, Accurate classification, neuroimaging data remains, data remains

备注: Accepted at the 17th International Workshop on Machine Learning in Medical Imaging (MLMI 2026), held in conjunction with MICCAI 2026

点击查看摘要

Abstract:Accurate classification of brain disorders from neuroimaging data remains challenging because of substantial inter-subject heterogeneity and the complex multi-scale patterns present in functional connectivity and morphological representations. To address these challenges, we propose HyperAMS-Net, a deep learning framework for brain disorder classification using neuroimaging representations derived from resting-state functional MRI or structural MRI. HyperAMS-Net integrates adaptive multi-scale convolution, hypergraph attention, spatial-channel attention, and adaptive feature fusion. Specifically, adaptive multi-scale convolution learns data-driven weights over multiple receptive fields to capture complementary patterns at different scales. Hypergraph attention models higher-order dependencies among learned feature representations through node--hyperedge--node message passing, while spatial-channel attention enhances discriminative feature learning. Adaptive feature fusion further aggregates complementary information across parallel network branches. HyperAMS-Net is evaluated on three benchmark datasets spanning distinct brain disorders: ABIDE for autism spectrum disorder, REST-meta-MDD for major depressive disorder, and ADNI for Alzheimer's disease, using 5-fold stratified cross-validation. HyperAMS-Net achieves state-of-the-art performance across all evaluated datasets, attaining the highest accuracy and AUC among the compared methods. Ablation studies further demonstrate the contribution of each proposed component, with the largest performance degradation observed when hypergraph attention is removed.

122. 【2609.19730】he segmentation ceiling: why explicit left-ventricular masks do not improve learned ejection-fraction regression

链接https://arxiv.org/abs/2609.19730

作者:Farshid Farhadi Khouzani,Paul La Plante,Bryar Mustafa Shareef,Laxmi Gewali

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:ventricular ejection fraction, left ventricular ejection, deep learning enables, learning enables automated, ejection fraction

备注: 15 pages, 4 figures. Submitted to Computers in Biology and Medicine

点击查看摘要

Abstract:Accurate estimation of left ventricular ejection fraction (EF) from echocardiography is central to cardiovascular care, and deep learning enables automated EF prediction from echocardiographic video. Because EF is clinically derived from left-ventricular (LV) volumes, a widely held intuition is that explicit LV segmentation should improve prediction. We introduce a quantitative criterion, the segmentation ceiling, that makes this testable: from EF as a normalized difference of end-diastolic and end-systolic volumes, we derive in closed form how per-frame segmentation area error propagates into EF error, and thus the accuracy a mask must reach before it can improve on direct regression. Using EchoNet-Dynamic, a UniFormer-S backbone, and the empirically measured within-patient error correlation, the criterion places the break-even near 10% per-frame area error, whereas a representative segmenter operates at roughly 14%, above the ceiling. Consistent with this, four strategies for injecting segmentation or area information (a predicted-mask channel, end-diastolic/end-systolic clip sampling, and per-bin and amplitude area-consistency objectives) fail to beat a raw-video baseline; ground-truth masks help only through label leakage. Input representation thus not being the limit, we identify generalization as the practical lever: weight averaging with strong augmentation attains a test R^2 of 0.806 (MAE 4.08) under a matched dense-clip protocol, comparable to an R(2+1)D baseline (0.811) while tightening the validation-to-test gap. Finally, a heteroscedastic beta-NLL formulation yields informative, well-calibrated per-prediction uncertainty, larger for clinically harder low-EF cases, where Monte-Carlo dropout does not. The segmentation ceiling gives a concrete design criterion for when mask-guided EF estimation is worthwhile, plus a simple, uncertainty-aware recipe for EF regression.

123. 【2609.19525】Compression Hurts, Pooling Helps: Information Loss in Rayleigh-Scale Estimation from B-Mode Ultrasound

链接https://arxiv.org/abs/2609.19525

作者:D. Hudson Smith,Ahmer Raza

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Applications (stat.AP)

关键词:potential data sources, standard clinical ultrasound, clinical ultrasound devices, Clinical B-mode images, tissue characterization

备注: 20 pages, 7 figures

点击查看摘要

Abstract:Clinical B-mode images are widely available as potential data sources for quantitative ultrasound (QUS) analysis for tissue characterization. However, standard clinical ultrasound devices apply unknown log-compression to RF envelope data before display and storage. Previous work has demonstrated estimation of the underlying RF envelope statistics in the presence of an unknown compression law. Using Fisher information analysis, we show that finite-offset log compression causes severe information loss when estimating the Rayleigh scale $\sigma$, which controls diffuse speckle. For a single image window, unknown compression raises the minimum achievable variance for unbiased estimation of $\sigma$ by a compression-independent factor of approximately $\FisherMinInflation$. When $M$ equal-sized windows share the same unknown compression settings, the excess variance decays as $1/M$; even in the most favorable regime, reducing the variance inflation factor below $1.1$ requires $\FisherBestCaseWindows$ windows. Our analysis treats the contrast parameter $a$ as unknown and the boundary offset $b$ as known; estimating $b$ experimentally shows even larger variance. We validate this theory using synthetic estimation experiments and demonstrate RF-scale recovery on real RF-envelope windows from the OASBUD dataset. Together, these results clarify the limitations of using routine B-mode images for QUS.

124. 【2609.19385】Mammography Foundation Models for Opportunistic Prediction of Major Adverse Cardiovascular Events

链接https://arxiv.org/abs/2609.19385

作者:Paula Feldman,Nusrat Binta Nizam,Sunwoo Kwak,Batuhan Karaman,Katerina Dodelzon,Mert Sabuncu

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:cardiovascular risk, Cardiovascular, remains the leading, routine care, risk

备注

点击查看摘要

Abstract:Cardiovascular disease (CVD) remains the leading cause of death among women, yet cardiovascular risk assessment often relies on clinical variables that may be missing, outdated, or unavailable in routine care. Screening mammography offers an opportunity for opportunistic cardiovascular risk stratification because it is routinely acquired and contains vascular features, including breast arterial calcifications (BAC), that are associated with cardiovascular risk and events. We evaluate whether mammography specific foundation models, originally pretrained for breast cancer-related tasks, can transfer to cardiovascular risk prediction without cardiovascular specific supervision or explicit BAC annotation. We constructed a 5-year major adverse cardiovascular event (MACE) cohort of 22,497 women linked to electronic health record outcomes, including 500 events (2.22% prevalence). The foundation models achieved AUROCs of 0.823 and 0.822 substantially exceeding an age-only model (AUROC 0.765), despite using only the screening mammogram as input, with no clinical variables. Both foundation models evaluated assigned substantially higher predicted risk to patients with radiologist-documented BAC, despite BAC never being used as a training label, and showed activation patterns consistent with vascular findings. Together, these findings suggest that mammography foundation models can recover clinically relevant cardiovascular risk information directly from mammographic pixels and suggest that screening mammography may provide an opportunistic source of cardiovascular risk information to complement conventional clinical assessment without additional imaging. Code is available in this https URL

125. 【2609.19215】Perceptual Refinement of an End-to-End Video Streaming Pipeline via Generative AI Layers

链接https://arxiv.org/abs/2609.19215

作者:Emanuele Artioli,Farzad Tashtarian,Christian Timmerer

类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Traditional codecs treat, frame alike, Traditional codecs, viewer attends, conference work ELVIS

备注: 28 pages. Submitted to ACM Transactions on Multimedia Computing, Communications and Applications (TOMM), special issue on MMSys and co-located workshops. Extended version of the NOSSDAV 2025 paper ELVIS ( [arXiv:2512.14185](https://arxiv.org/abs/2512.14185) ). Code: [this https URL](https://github.com/emanuele-artioli/presley)

点击查看摘要

Abstract:Traditional codecs treat every region of a frame alike; a generative layer can instead degrade the regions a viewer attends to least and reconstruct them at the client. We present PRESLEY, which extends the prior conference work ELVIS by replacing destructive block removal with adaptive in-place degradation under a removability mask, signaling per-block strength in a bit-packed side channel, and restoring via generative backbones conditioned on transmitted visual priors rather than unconditioned in-painting. We separate the problem into three goals: choosing which blocks to degrade, degrading them so the encoder spends fewer bits, and restoring them. Against its predecessor at matched rate, PRESLEY achieves a decisive mean -56.4% BD-rate reduction on delivered background quality across 13 rate ladders spanning multiple codecs and dataset families. Against pristine baselines, PRESLEY defines the operating regime of generative transport: delivering substantial bitrate savings (up to -29.4% BD-rate) and superior background quality (17/23 sequences) in the target bit-starved regime, while maintaining foreground fidelity bit-exact. We further map where the theoretical headroom in this class of architecture lies. Using an exact leave-one-superblock-out combinatorial oracle as an additive empirical bound, we show that existing complexity heuristics already capture 83.3% of bit-cost savings, bounding remaining cost-axis headroom at about 5% of total bitrate. We then identify and model the primary unaddressed axis -- post-restoration damage -- which disperses widely (4.9-8.4 dB). We prove that this damage is predictable before transmission (held-out rho = +0.400), establishing the feasibility of transmit-time restorability modeling and defining the roadmap for joint rate-distortion-restoration selection rules.