本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新510篇论文,其中:

  • 自然语言处理62
  • 信息检索13
  • 计算机视觉89

自然语言处理

1. 【2608.14509】Split the Labor: Separating Evidence Interpretation from Decision Aggregation

链接https://arxiv.org/abs/2608.14509

作者:Zhelun Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:reach a conclusion, source rewards capacity, Abstract, prompt, model to reach

备注: Atlassian. 22 pages, 2 figures

点击查看摘要

Abstract:Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return nothing. Once separated, the design problem becomes the interface between them. We propose a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance) and show that fixing it determines both halves. The separation also reveals a failure mode in how such systems combine, which we call count-scale drift. Thresholding a sum of unnormalized weights is exactly posterior thresholding, but at an operating point that slides with the number of sources consulted. The slide grows with reader reliability. When source reliabilities differ, the vote rule and the posterior order instances differently, and no threshold reconciles them. Pooling calibrated log-likelihood ratios addresses both problems. The fix is arithmetic rather than architectural, and applies to a class of rules beyond language models: score-summing triage engines, diagnostic panels scored by counting positives, and additive multi-signal detectors. We then instantiate the principle twice on one longitudinal corpus, once after outcomes resolve and once before. The same partition helps in both, at different granularities: over reading in the first, over learning capacity in the second. There, a small sequence encoder on an easy auxiliary objective plus a tree ensemble carrying the censored survival loss reaches 0.921 AUPRC against 0.805 for a hand-crafted baseline. We separate what transfers from what must be re-estimated per domain, and state five predictions that would falsify the framework, three negative results, and which comparisons remain confounded.

2. 【2608.14465】You Only Pass Once: Answering and Abstaining Together in a Single Forward Pass of a Frozen Language Model

链接https://arxiv.org/abs/2608.14465

作者:Ziyang Luo,Zhongyao Chu,Xinjie He,Youting Wang,Xukui Qin,Runxiong Wu,Yan-Syuan Chen

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:frozen language model, residual stream, coupled weaknesses, under-uses evidence, fails to detect

备注: 24 pages. Ziyang Luo and Zhongyao Chu contributed equally

点击查看摘要

Abstract:A frozen language model on reasoning tasks has two coupled weaknesses: it under-uses evidence its own residual stream already encodes, and it fails to detect when the input is insufficient to answer, so it confabulates. This paper consolidates two research lines that address these on the same residual stream: a conditional steering probe writes the stream at mid-stack layers and recovers reasoning accuracy from a frozen backbone, and a zero-shot sufficiency direction reads the stream and abstains when information is insufficient. Deployed in one forward pass they interfere: the steering write shifts the state the direction reads, costing up to 8 AUROC points of cross-domain transfer on small models; a separate clean pass doubles inference cost. We keep the direction fixed and train a small network to reconstruct the pre-steering residual from the steered one -- mean-squared error on (steered, clean) pairs, no sufficiency labels -- and read the direction on the reconstruction. The resulting system, YOPO (You Only Pass Once), answers, steers, and abstains in one forward pass of a frozen Qwen2.5 backbone (1.5B/3B/7B). End to end, three-way accuracy more than doubles the frozen baseline (0.375-0.798 on 1.5B alphaNLI) and one pass beats the two-pass reference at every scale (0.798/0.830/0.893 vs 0.753/0.790/0.863) and on ten backbones across six model families. We chart the capacity-transfer frontier quantifying the principle that abstention should not be trained in; a source-side audit catches our own alphaNLI construction leaking a surface artifact, so architectural claims are anchored on native-label replications (SQuAD2, RepLiQA, MuSiQue); and on the standard four-domain suite we contribute, to our knowledge, the first answer-or-abstain benchmark, where our gate tops every in-domain dataset and the label-free direction is the only gate family to survive domain transfer.

3. 【2608.14457】Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

链接https://arxiv.org/abs/2608.14457

作者:Isabel Cachola,William Walden,Reno Kriz,Mark Dredze

类目:Computation and Language (cs.CL)

关键词:general summary quality, specific desired properties, desired properties, focuses on general, summarization evaluation focuses

备注

点击查看摘要

Abstract:The majority of work on summarization evaluation focuses on general summary quality (e.g., ROUGE, BERTScore) or specific desired properties (e.g., readability, factuality). However, these metrics fail to measure the utility of a summary to an individual user. For example, a biomedical researcher learning about the latest vaccine research will have different informational needs from a family doctor. Query-focused summarization captures part of this need, but in practice, users rarely state everything relevant in a query: a single short query is likely inadequate to distinguish the needs of a researcher from those of a physician. By contrast, a reader's background or persona (their role and expertise) is comparatively stable across queries and recovers much of this missing context, which makes it a practical signal for assessing whether a summary satisfies that reader's needs. In this work, we assess how sensitive popular summarization metrics are to both informational and persona differences, and find that many popular metrics, including strong LLM-as-judge metrics, fail basic perturbation tests of informational content. We additionally conduct an expert human evaluation, measuring summary preferences based on information satisfaction given a specific person's background and use case. We find that both traditional and LLM-based metrics are insufficient measures of information satisfaction and agree poorly with human judgment.

4. 【2608.14399】Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice

链接https://arxiv.org/abs/2608.14399

作者:Syeda Anshrah Gillani,Mirza Samad Ahmed Baig

类目:Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language model, assistants which doctor, systems AI infomediaries, increasingly ask large, large language

备注: 26 pages, 9 figures, 10 tables

点击查看摘要

Abstract:Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.

5. 【2608.14397】LLMs Don't Pay for the Jump

链接https://arxiv.org/abs/2608.14397

作者:Paras Balani,Subhrakanta Panda

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Einstein equivalence principle, Large Language Models, Large Language, produced Einstein equivalence, Einstein equivalence

备注: 14 pages

点击查看摘要

Abstract:Zahavy [2026] argues that Large Language Models, despite their capabilities in induction and deduction, cannot perform the abductive "Jump" that produced Einstein's equivalence principle, and attributes this limitation to the absence of embodied simulation. Zheng-Xin [2026] and Farmer [2026] question whether embodiment is necessary for abduction, pointing to alternative routes to General Relativity and forms of abduction that require no sensorimotor grounding. Max Planck resolved the blackbody radiation problem in 1900. Planck's move to E = h{\nu} required no embodied simulation. It was motivated by a mathematical consequence of classical theory, an infinite predicted energy for a finite measured quantity, that could not be physically accepted. We show that neither induction nor deduction could have produced the postulate and argue that its adoption required a coupling between epistemic error and physical cost. We formalize this distinction through thermodynamic coupling and show that fixed-weight transformer inference lacks such coupling, regardless of model scale. This is consistent with empirical results showing that output entropy remains nearly unchanged across tasks with sharply increasing causal difficulty, even as accuracy falls from 100% to 17%. We therefore argue that the missing ingredient in machine abduction may lie deeper than embodiment: a system must have some physical mechanism through which epistemic error becomes costly enough to force revision.

6. 【2608.14377】A Survey of Large Models in Sports

链接https://arxiv.org/abs/2608.14377

作者:Yichen Xu,Jianzhe Ma,Chuhan Wang,Zhonghao Cao,Liangyu Chen,Wenxuan Wang,Qin Jin

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:witnessed growing global, growing global enthusiasm, cultural exchange, social connection, recent years

备注: 36 pages, 4 figures, 6 tables. Accepted to Findings of ACL 2026

点击查看摘要

Abstract:Sports have witnessed growing global enthusiasm in recent years, serving as a vital force for physical health, cultural exchange, social connection, and economic growth. The rapid advancement of large models, particularly (multimodal) large language models (M)LLMs, has demonstrated transformative potential to reshape sports understanding, analysis, and interaction across diverse domains. This paper presents a comprehensive survey of large models in sports, including (i) an overview of tasks and applications across different participant groups; (ii) a detailed analysis of sports-related datasets and benchmarks; and (iii) a critical discussion of current challenges and future directions. Our goal is to establish a foundation for advancing research and practical development of large-model-driven sports intelligence. An open-source GitHub repository is maintained at: this https URL.

7. 【2608.14375】Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages

链接https://arxiv.org/abs/2608.14375

作者:Chih-Hsuan Yang,Anjir Ahmed Chowdhury,Cheng-Hau Yang,Weijian Zheng,Fernando Llorente,Xiaolong Ma,Xinyang Li,Eliu A. Huerta,Ian T. Foster,Rajeev Thakur

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Multi-agent reasoning systems, Multi-agent reasoning, Diverse Hypothesis Deliberation, automated scores, scores to decide

备注: 24 pages, 9 figures. Includes an appendix and an ancillary reproducibility artifact

点击查看摘要

Abstract:Multi-agent reasoning systems often use agreement, confidence, or automated scores to decide which messages should shape a final answer. Such filtering assumes that a message likely to be correct is also worth keeping. Yet a wrong answer can contain a useful decomposition, constraint, or scientific principle. We test this distinction with Diverse Hypothesis Deliberation (DHD), a controlled measurement protocol that caches five independently generated messages and replays the same downstream solver, called the integrator, with each message available or hidden. The replay comparison measures a message's trajectory value: whether making the message available helps or harms subsequent reasoning. Across five mathematics and science benchmarks and two openly available model families, gpt-oss-120b and gemma-4-31B-it, wrong-helpful messages appear in every benchmark-model combination. Among wrong-answer messages that change final correctness, more than four in ten changes are helpful in each model. Controlled repeats show that the number of repeatable message effects is unlikely to arise from replay variation alone (p=0.0002). A focused intervention on repeatable wrong-helpful messages finds that the complete message works best, while retaining its reasoning preserves more success than retaining only its answer; the source of the complete-message advantage remains open. Within the same problem, repeated trajectory-value evidence also identifies a better keep-or-remove choice than answer correctness alone. Answer correctness is therefore informative but does not determine trajectory value. DHD measures this missing property and produces reusable labels for learning when agents should listen.

8. 【2608.14361】Local and Global Regimes of Geometric Complexity in Language Model Representations

链接https://arxiv.org/abs/2608.14361

作者:Arwa Osman,Marco Baroni,Iuri Macocco

类目:Computation and Language (cs.CL)

关键词:differences reflect properties, language models, properties of language, lexical diversity, probe the representational

备注: 12 pages, 9 figures

点击查看摘要

Abstract:Intrinsic dimensionality (ID) is widely used to probe the representational complexity of language models, but it remains unclear whether ID differences reflect properties of language itself or artefacts of how the underlying dataset was constructed. In this paper, we focus specifically on how lexical diversity, the number of unique last-token items present in a dataset, affects ID estimates of that dataset. We find a scale-dependent transition between two regimes: at low lexical diversity, conditions with fewer unique final words produce higher ID, while at high lexical diversity, this ordering reverses, and conditions with more unique words produce higher ID. We derive an exact, parameter-free formula for the point at which this reversal occurs, which matches the observed transition point at every scale tested. On the one hand, our results highlight how care must be taken when interpreting the intrinsic dimensionality of a set of representations as a straightforward cue of their complexity. On the other hand, our discovery of the two ID regimes reveals a general principle of organisation of linguistic data in LLMs that sheds new light on their inner manifold structures.

9. 【2608.14329】A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

链接https://arxiv.org/abs/2608.14329

作者:Dipankar Sarkar

类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:deliver good outcomes, deliver good, good outcomes, binary predicates, Principle-based regulation

备注: 7 pages, 3 figures. Accepted at the KDD 2026 Workshop on Secure and Trustworthy Large Language Models (SeT-LLM), poster

点击查看摘要

Abstract:Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our position is that any such judge must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. We release Principle-Bench, 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations authored under a pre-registered rubric; the first benchmark covering all four axes for principle-based regulation. We also introduce Ceca (Calibrated Exemplar-Cluster Assessment): a calibrated, auditable assessor that emits exact per-exemplar counterfactual attributions. Across keyword counting, three sentence-transformer embedders, an open-weight LLM-judge, and a calibrated cascade, no method dominates all four axes. A 120B LLM-judge, strongest on benign inputs, loses 47 accuracy points (0.74 to 0.27) on keyword-stuffed Consumer Duty inputs: "compliance theatre." A second judge from a different model family agrees only at Cohen's kappa = 0.16 on that split, localising the failure to the model rather than the corpus. Any deployment-grade LLM-judge for principle-based regulation must report per-principle adversarial deception and post-hoc calibration alongside aggregate accuracy.

10. 【2608.14320】AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

链接https://arxiv.org/abs/2608.14320

作者:Yiderigun Borjigin,Alexander Hermann,Christian Cyron,Roland Aydin

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:cognitive bias, initial reference, anchoring effect, frontier API models, API models

备注: Published as a conference paper at COLM 2026

点击查看摘要

Abstract:The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG, and (4) high task accuracy on the anchor-free control condition (Acc$_{10}$: answers within 10 points of gold) does not guarantee robustness: even frontier API models above 95% control accuracy remain susceptible to plausible anchors.

11. 【2608.14312】Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL

链接https://arxiv.org/abs/2608.14312

作者:Xiaojun Wu,Cehao Yang,Honghao Liu,Xueyuan Lin,Zhichao Shi,Hao Zhou,Xuhui Jiang,Chengjin Xu,Jia Li,Jian Guo

类目:Computation and Language (cs.CL)

关键词:Reinforcement learning, terminal agents, agents needs executable, executable training environments, reliable rewards

备注: 19 pages, 5 figures

点击查看摘要

Abstract:Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few-shot, Self-Instruct, and Evol-Instruct apply the same prompting policy to every seed, even when the current policy would benefit from a harder, easier, or simply different task. We present Envs-FORGE, a prompting policy that converts verifier rewards into per-seed environment-synthesis actions. Envs-FORGE estimates seed pass rates, scores six projection--direction actions around a target learning frontier, and solves a per-seed mixed-integer linear program (MILP) to choose the action that conditions generation. The selected action drives synchronized rewriting of the instruction, fixtures, oracle solution, tests, and Docker environment; only gold-verified bundles enter RL training. The indexed MILP form also supports optional soft skill coverage for portfolio planning. On Qwen 3.5 35B, Envs-FORGE improves Pass@1 over Base by 9.2 percentage points on tb-core (40.0% to 49.2%) and 6.4 points on tb-2.0 (23.0% to 29.4%), exceeding the strongest fixed-recipe baseline by 2.4 and 2.1 points. It reaches 77.1% on SWE-bench Verified versus 73.4% for Base, and improves tb-core by 6.8--9.2 points across the evaluated 4B--35B models. All synthesis methods export 100 verified environments and use 2.27M--2.88M synthesis tokens, placing the comparison at the same downstream training-set size and the same operational scale. The source code is available at this https URL.

12. 【2608.14286】Seeing Red, Thinking Bad: Color Bias in Vision Language Models

链接https://arxiv.org/abs/2608.14286

作者:Kohsuke Ide,Ryousuke Yamada,Yoshihiro Fukuhara,Hirokatsu Kataoka,Yutaka Satoh

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:industrial decision-making systems, visual styling, decision-making systems, support and recommendation, Vision language models

备注: 15 pages. Accepted to ICPR 2026

点击查看摘要

Abstract:Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: this https URL

Comments:
15 pages. Accepted to ICPR 2026

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

Cite as:
arXiv:2608.14286 [cs.CV]

(or
arXiv:2608.14286v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2608.14286

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Journalreference:
In: Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science. Springer, Cham, pp. 261-275

Related DOI:

https://doi.org/10.1007/978-3-032-31583-0_18

Focus to learn more

            DOI(s) linking to related resources</p>
13. 【2608.14277】SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

链接https://arxiv.org/abs/2608.14277

作者:Haonan He,Haodi Lei,Yun Luo,Haoran Zhang,Shunkai Zhang,Yizhuo Li,Shengji Tang,Zhilin Wang,Runzhe Zhan,Lei Bai,Ganqu Cui,Fangchen Yu,Yafu Li,Peng Ye,Ning Ding,Yu Cheng

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:On-policy distillation, stronger teacher models, response length explosion, introduces practical challenges, short-context student models

备注

点击查看摘要

Abstract:On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student distribution mismatch, response length explosion, and training instability. In this work, we study this setting by transferring proof-reasoning capabilities from the long-context reasoning model SU-01 to short-context student models. To handle tokenizer differences, we perform OPD in a shared text space and align only tokens that occupy identical text spans under the student and teacher tokenizers. To mitigate the problem of excessive generation length and frequent truncation, we introduce a student reference KL loss and mask the advantages of special termination tokens such as /think and |im_end|. This strategy constrains the student from drifting excessively from its initial policy, thereby mitigating the teacher-student distribution mismatch problem and fostering steady length growth. Experiments on both same-family and different-family student models, including Qwen3, Qwen3.5, Intern-S2, GLM-4.7, Gemma-4, show consistent gains in mathematical reasoning, especially natural-language math proving. Notably, Intern-S2-Preview improves by 21.2 points on ProofBench, reaching 55.2 and surpassing Gemini-2.5-Pro. It also improves on science benchmarks such as HLE and HiPhO, suggesting that OPD transfers reasoning capabilities that generalize beyond the mathematical training domain.

14. 【2608.14252】Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models

链接https://arxiv.org/abs/2608.14252

作者:Brett Reynolds

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Recent work suggests, Recent work, content or reference, work suggests, representations have content

备注: 24 pages, 1 figure, 1 table. A six-page methodological supplement, reproducible R script, and constructed data are included as ancillary files

点击查看摘要

Abstract:Recent work suggests that some large language model representations have content or reference. Grounding can secure either without supplying live routes for correction. This paper asks what follows from that gap. An output is answerable when discrepancies can affect what a target- and task-specific arrangement produces, accepts, or withdraws. The arrangement has corrective control only when live, sufficiently independent routes can detect and repair fresh discrepancies. A route profile records which routes constrain the arrangement and how they are related. Those profiles support analysis of truth-tracking: patterned support for representational success. Language models are the pressure case; text-only arrangements provide a task-relative limiting case. Text-trained models inherit patterns of testimony, coherence, and prior correction. Where target-sensitive correction survives training, these can supply derivative answerability (inherited constraint); live answerability is the relation supplied by a current route for fresh discrepancies. Fluent failures should follow when a task requires independently informative access to the facts. Self-consistency, retrieval, tools, code execution, multimodal input, and feedback should help selectively. Route-by-task interactions test the distinctions. The decomposition's empirical burden is to predict held-out route--task combinations or improve intervention choice without conceptual refitting. Surface improvement and truth-tracking improvement can come apart.

Comments:
24 pages, 1 figure, 1 table. A six-page methodological supplement, reproducible R script, and constructed data are included as ancillary files

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

Cite as:
arXiv:2608.14252 [cs.AI]

(or
arXiv:2608.14252v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2608.14252

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
15. 【2608.14229】he More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning

链接https://arxiv.org/abs/2608.14229

作者:Anna Borisiuk,Andrey Savchenko,Alexander Panchenko,Elena Tutubalina

类目:Computation and Language (cs.CL)

关键词:existing LLM unlearning, resist removal longer, apply uniform gradient, uniform gradient pressure, LLM unlearning methods

备注

点击查看摘要

Abstract:Popular facts are memorised more deeply during pretraining and resist removal longer than rare ones, yet existing LLM unlearning methods apply uniform gradient pressure regardless of training-data frequency. We propose the AdaPop (Adaptive Popularity) method, which combines local token confidence with a per-fact popularity-dependent exponent derived from an external proxy (e.g., Wikidata sitelinks, LLM-as-Judge), and automates the forget-retain balance via a dual-ascent controller that adjusts the retain penalty each epoch. Across three model families and two benchmarks, AdaPop leaks ~5x less forgotten content than competing methods under paraphrased queries and ~1.6x less under adversarial reformulations. We support our analysis with internal metrics: under our method, forget-set hidden states move further from the pre-unlearning model's states than under other methods, while retain-set representations remain close.

16. 【2608.14221】MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

链接https://arxiv.org/abs/2608.14221

作者:Lushi Pu,Weiming Zhang,Xinheng Xie,Zixuan Fu,Bingxiang He,Hengyu Zhao,Hongya Lyu,Xin Li,Jie Zhou,Yudong Wang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:machine-verifiable formal languages, translating natural-language mathematical, commonly framed, framed as translating, translating natural-language

备注: 25 pages, 6 figures, 8 tables

点击查看摘要

Abstract:Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable formal languages such as Lean 4. However, faithful formalization requires more than translation. Models must map mathematical concepts to the complex hierarchy of types and definitions in formal libraries such as Mathlib, while ensuring that generated statements preserve the meaning of the source propositions. Existing approaches struggle because they rely heavily on the model's parametric memory for library-specific knowledge, while common data construction pipelines often resort to filtering single-pass outputs and lack mechanisms for feedback-driven revision. To address these challenges, we introduce MathForm, an autoformalization framework for constructing verified training data through Mathlib knowledge retrieval and verification-guided iterative refinement. Before generation, a retrieval planner gathers relevant definitions and existing formalizations from Mathlib to guide the formalization generator. Generated statements are then revised using compiler diagnostics and semantic-consistency feedback. Using this framework, we construct FormalVerse, a Lean 4 dataset containing approximately 367K verified examples across diverse mathematical domains and sources. We then train MathForm-8B through supervised fine-tuning followed by reinforcement learning. Across six benchmarks, MathForm-8B achieves average Pass@8 rates of 88.06% under Syntax Check (SC) and 72.37% under Consistency Check (CC), outperforming multiple specialized 32B autoformalizers. On the challenging FATE-H and FATE-X subsets, it attains CC pass rates of 63% and 37%, exceeding the strongest specialized baselines in both cases.

17. 【2608.14210】How Much Do Legal RAG Systems Still Hallucinate?

链接https://arxiv.org/abs/2608.14210

作者:Souvick Das,Sallam Abualhaija,Domenico Bianculli

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:retrieval-augmented generation, legal RAG systems, major challenge, challenge for retrieval-augmented, ungrounded answers

备注

点击查看摘要

Abstract:Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR (in English) and a national civil law (in French). Using claim-level and answer-level evaluation, we report on hallucination density and severity, analyze performance across question categories and user personas, and validate our findings on an independent set of 142 legal-expert-authored questions. Our results show that hallucinations remain pervasive, ranging from less than 10% of responses for the best-performing systems to nearly half in the worst case. We further find that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.

18. 【2608.14198】MINT: A Universal Zero-Shot Predictor for Transaction Data

链接https://arxiv.org/abs/2608.14198

作者:Parameswaran Kamalaruban,Viktor Drobnyi,Maeve Madigan,Julia Rozanova,David Sutton,Stuart Burrell

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:including fraud prevention, credit risk assessment, Banks analyse sequential, Payments Foundation Models, sequential financial transaction

备注

点击查看摘要

Abstract:Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. However, these Foundation Models are not designed for flexible zero-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility. Existing LLM-based approaches to zero-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task-specific architectures that scale poorly. To address these limitations, we present the Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. We find that MINT achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions, while substantially reducing input tokens, latency, and memory consumption compared to text-serialization baselines. Through comprehensive analyses of representations, alignment strategies, training data, and history length, we establish that compact transaction embeddings are a superior approach to transaction representation than text serialization for multimodal reasoning and zero-shot prediction tasks.

19. 【2608.14191】KV Cache Compression Through the Lens of Transform Coding

链接https://arxiv.org/abs/2608.14191

作者:Hannah Laus,Claudio Mayrink Verdun,Hao Wang,Flavio du Pin Calmon,Felix Krahmer

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Signal Processing (eess.SP)

关键词:cache stores information, major memory bottleneck, long-context inference, stores information, information from past

备注

点击查看摘要

Abstract:The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data types and designing quantization schemes to minimize reconstruction error in the cache itself, without accounting for how that error propagates through attention mechanisms. We prove that, under a white-noise quantization model, the expected attention-aware distortion decomposes into additive key and value contributions that factor across tokens and channels. Building on transform coding and reverse water-filling, which are classical tools from signal processing and rate-distortion theory, we introduce Attention-Aware Transform Coding (AATC), which allocates bits over a calibration set to minimize attention-aware distortion. On Llama-3.1-8B-Instruct and Qwen-2.5-7B-Instruct, evaluated across LongBench, RULER, GSM8K, MMLU-Pro, and MATH-500, our method achieves near-lossless accuracy at approximately $5.8\times$ compression, whereas each baseline degrades in at least some settings.

20. 【2608.14150】Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

链接https://arxiv.org/abs/2608.14150

作者:Kexin Shi,Renhe Sun,Yuge Huang,Ximeng Wang,Jiayi Zhou,Jian Liu,Malu Zhang

类目:Computation and Language (cs.CL)

关键词:Speech Language Model, Conversational Speech Language, Multilingual Conversational Speech, conversational speech understanding, unsegmented multilingual conversations

备注

点击查看摘要

Abstract:The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversations: speaker diarization and recognition (Task 1) and conversational speech understanding (Task 2). Neither task provides oracle utterance boundaries or speaker labels at evaluation, and Task 2 provides no question-answer training set. For Task 1, we fine-tune VibeVoice-ASR-7B with random leading-silence cropping, consistent timestamp correction, and an exponential moving average (EMA) training strategy. For Task 2, we construct synthetic question-answer pairs through multimodal candidate generation, silent-audio filtering, and distribution-matched augmentation, and fine-tune Qwen3-Omni-30B-A3B-Instruct for tagged direct answering. On the Task 1 evaluation set, cropping reduces tcpMER from 18.30% to 17.27%, and EMA further reduces it to 16.73%. On the Task 2 evaluation set, jointly applying distribution-matched augmentation and tagged direct answering raises accuracy from 83.0% to 86.0%.

21. 【2608.14089】Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers

链接https://arxiv.org/abs/2608.14089

作者:Thiago Sandoval,Ufuk Topcu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG)

关键词:large language models, Safety classifiers deployed, deployment traffic evolves, deployer desired policy, Safety classifiers

备注: 16 pages including technical appendix, 6 figures

点击查看摘要

Abstract:Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. We present Regime-Conditional Verification (RCV), a lightweight wrapper that adapts an off-the-shelf safety classifier without retraining it. RCV estimates, from the classifier's internal representations, the probability that each prediction disagrees with the deployer's policy, and selectively corrects predictions likely to be wrong. The same correctness estimates also provide a label-free signal for detecting distribution shift, enabling a maintenance loop that updates the correctness estimation layer and resorts to classifier fine-tuning only when necessary. Across three off-the-shelf safety classifiers and two benchmark datasets, RCV improves adherence to the deployer's policy in every classifier-dataset combination, catching up to 0.81 of previously missed unsafe content without modifying the underlying classifier. In a deployment study with ten attack campaigns, each a harm category held out of RCV's training, RCV detects every campaign in a dedicated injection panel; in the maintenance census most drift episodes are repaired without updating the classifier, and the fine-tune is reserved for the residual episodes that repair does not restore.

22. 【2608.14079】he conditional superiority of fast silicon sampling

链接https://arxiv.org/abs/2608.14079

作者:Nickolas Hock Yuen Lam,Ji Xuan Voo,Xiangyu Ma

类目:Computation and Language (cs.CL); Materials Science (cond-mat.mtrl-sci)

关键词:Silicon sampling, surprisingly good population, produce surprisingly good, Silicon, surprisingly good

备注

点击查看摘要

Abstract:Silicon sampling can produce surprisingly good population estimates at times. Does doing it fast attenuate such fidelity? In this study, we extend and assess ongoing work in silicon sampling by comparing the algorithmic fidelity of "fast" and "slow" modes of silicon sampling among a nationally representative sample of Singaporean survey respondents. We find that silicon sampling with contemporary frontier models remains a method in early development to be used only with great caution. While silicon samples are able to produce moderately faithful estimates of population means, they continue to understate opinion variance and distort the latent contextual space behind human opinions. Conditional on such limitations, we find "fast" modes of silicon sampling to be relatively superior to traditional "slow" modes of silicon sampling. Fast silicon sampling is significantly more efficient in compute resources and run-time while being monotonically superior to slower modes of sampling in algorithmic fidelity.

23. 【2608.14055】HERMES: a multi-agent framework for structured knowledge extraction from ultra-long documents in geoscience

链接https://arxiv.org/abs/2608.14055

作者:Ziqi Song,Zongyuan Xiang,James G. Ogg,Bruce S. Lieberman,Gabi Ogg,Natalia López Carranza,Wen Du,Yufei Ye,Shuan Li,Zhong Peng,Shaoqi Yu,Juye Wei,Ying Zhou,Jieping Ye,Jiang Yang

类目:Computation and Language (cs.CL)

关键词:hinder computational access, remains largely trapped, complex layouts hinder, layouts hinder computational, Authoritative scientific knowledge

备注: 31-page main manuscript with 6 figures and 3 tables; supplementary information included

点击查看摘要

Abstract:Authoritative scientific knowledge in geoscience remains largely trapped in legacy monographs and historical literature, where unstructured text and complex layouts hinder computational access. We introduce HERMES, a scalable multi-agent framework that extracts structured data from ultra-long scientific documents. Using a coordinating large language model, HERMES integrates domain constraints, validation rules and evidence tracing within a unified document-level extraction process that incorporates parsed text, tables, figures and captions. Applied to the 55-volume Treatise on Invertebrate Paleontology, the system produced a structured database of 32,277 fossil taxonomic entities and 451,878 attributes, released online at this https URL. Extraction performance remained stable across fossil groups (average F1 scores of approximately 0.90 for entities and 0.91 for attributes), improving per-volume efficiency approximately sixfold relative to the tested fully manual baseline. Evaluation in palaeomagnetism and geochemistry, conducted without additional model training, demonstrated transfer across distinct geoscience domains. This work provides a practical pathway to transform historical scientific literature into FAIR-oriented structured data, offering a sustainable infrastructure for data-intensive disciplines and large-scale knowledge integration.

24. 【2608.14029】S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

链接https://arxiv.org/abs/2608.14029

作者:Xueqi Wang,Zhigang Wang,Runqing Zhang,Zhenqi Jia,Junfeng Zhao

类目:Computation and Language (cs.CL)

关键词:Conversational Speech Synthesis, acoustic conversational styles, Spoken Dialogue Systems, including Emotion Recognition, conversational styles

备注

点击查看摘要

Abstract:Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level retrieval is crucial for many dialogue-related tasks, including Emotion Recognition in Conversation, Spoken Dialogue Systems, and Conversational Speech Synthesis, where external dialogue examples can provide valuable semantic and stylistic references. However, existing retrieval methods are still largely limited to utterance-level or unimodal matching, and often fail to capture the global semantic coherence and stylistic consistency of an entire dialogue. To address this gap, we propose S2Dialog, a unified framework for dialogue-level semantic-style retrieval from multimodal dialogue banks. Specifically, S2Dialog consists of a Dialogue-level Textual Retriever and a Dialogue-level Acoustic Retriever, which encode the textual and acoustic modalities of a dialogue into dialogue-level representations, respectively. To further enhance multimodal retrieval, we introduce Dialogue-level Textual-Acoustic Contrastive Learning, which aligns semantically and stylistically similar dialogues while distinguishing unrelated ones. Extensive experiments on the multimodal dialogue dataset DailyTalk demonstrate that S2Dialog achieves outstanding retrieval performance.

25. 【2608.14003】Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model

链接https://arxiv.org/abs/2608.14003

作者:Yongmin Kim,Shota Takashiro,Yusuke Iwasawa,Takeshi Kojima,Yutaka Matsuo

类目:Computation and Language (cs.CL)

关键词:Large Reasoning Models, achieve strong performance, incur substantial computational, substantial computational costs, Large Reasoning

备注: Accepted at COLM 2026. 28 pages, 12 figures, 18 tables. Code: [this https URL](https://github.com/matsuolab/batch-wise-prune)

点击查看摘要

Abstract:Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference. In production settings, batched inference is essential for high throughput, yet the existing training-free adaptive pruning methods we evaluate severely degrade in this regime. Because a batch must share a single pruning mask, these methods aggregate activations across samples and then apply threshold-based selection; the threshold, calibrated offline on unaggregated activations, no longer matches the aggregated distribution, so the realized sparsity ratio drifts and accuracy on reasoning tasks collapses under batched inference. In this work, we propose a training-free adaptive pruning method designed specifically for batched inference in LRMs, built on two components. First, we replace threshold-based selection with periodic top-k selection over the aggregated importance scores, which is unaffected by the shift that aggregation induces in the activation distribution, and which runs selection once per update period rather than at every token, preserving the speedup. Second, based on the observation that important neurons re-fire periodically during long reasoning generation, we introduce an activation memory that accumulates importance across update phases so that recurring neurons are retained. Experiments on diverse reasoning benchmarks demonstrate that our method outperforms the previous state-of-the-art adaptive pruning method by 39.7 percentage points in average accuracy at batch size 4 with 50% target sparsity on DeepSeek-R1-Distill-Qwen-7B, and reaches 1.40x speedup over dense inference at 50% actual sparsity.

Comments:
Accepted at COLM 2026. 28 pages, 12 figures, 18 tables. Code: this https URL

Subjects:

Computation and Language (cs.CL)

Cite as:
arXiv:2608.14003 [cs.CL]

(or
arXiv:2608.14003v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2608.14003

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
26. 【2608.13966】QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

链接https://arxiv.org/abs/2608.13966

作者:Vincent Counathe,Ben Athiwaratkun,Christopher De Sa,Tianyi Zhang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML)

关键词:making quantization-aware training, preserving model quality, large language model, language model inference, model inference shifts

备注: 39 pages

点击查看摘要

Abstract:As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT computes the loss and surrogate gradients using a lossy reconstruction of latent full-precision weights, while applying updates to the latent weights themselves. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods mitigate a similar gap by minimizing loss-aware reconstruction error, but doing it once for a frozen model can take hours; repeating this process throughout QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that continuously performs lightweight, loss-aware reconstruction in the training loop to lower the loss floor and improve the resulting low-bit model. At each training step, QUASAR uses the exponential moving average of squared gradients as online saliency estimates, searches over a small set of clipping ranges, and fits affine dequantizers via saliency-weighted least squares. Our analysis shows that the loss-aware reconstruction error is the only reconstruction-dependent term in the QAT convergence bound and controls the loss of the final quantized model, establishing QUASAR's objective as a principled optimization target. QUASAR modifies only the training procedure and supports standard deployment formats, including integer quantization and NVFP4, with no inference-time changes or overhead. Across Qwen3 and Llama-3.1, QUASAR achieves the lowest held-out KL divergence among competitive QAT methods at 2, 3, and 4 bits, reducing KL by at least 10% at 3 and 4 bits and by 29% at 2 bits. At 2 bits, it improves average accuracy across eight tasks by 3.5-4.3 percentage points over strong QAT and PTQ baselines.

27. 【2608.13959】Repair, Not Improvement: Decomposing Constrained Decoding in Tool-Call Abstention

链接https://arxiv.org/abs/2608.13959

作者:Janghoon Lee(Redrob)

类目:Computation and Language (cs.CL)

关键词:Function calling, constrained generation explicitly, decoder contribution small, generation explicitly sets, format constraints

备注: 24 pages, 4 figures, 17 tables

点击查看摘要

Abstract:Function calling is what the recent accounting of constrained generation explicitly sets aside: it finds the decoder's contribution small for format constraints, then warns in its Section 7 against extrapolating where a constraint encodes a correctness requirement, and names function calling as one. Tool abstention is that case at its sharpest: an enum leaves the wording of an answer alone and narrows the set of answers there are, and declining to call anything is the first it drops. We measure the excluded case. Three conditions over one byte-identical prompt separate a grammar's two jobs: it fixes where generation stops as well as which tokens may be emitted. We evaluate open-weight models from 0.6B to 4B on matched English and Korean items, so the language comparison is made within item. Against an unconstrained decoder, prior work's contrast is negative on abstention in four of six cells with intervals excluding zero, worst -29.5 points, and positive with an interval excluding zero in none. The total is a sum with opposite signs: on the smallest model in Korean the stop token costs -20.0, the enum returns +19.5, and the two leave -0.5. What it recovers is form: of 698 abstentions repaired, 545 had no readable answer and 0 were judgements the scorer refused. On tool-needed items it is positive throughout; abstention leads because it is the preregistered measure, and the pooled number being kinder to the intervention makes moving to it worse rather than better. Both preregistered language claims fail.

28. 【2608.13947】Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion

链接https://arxiv.org/abs/2608.13947

作者:Hwan Chang,Yongil Kim,Heuiyeen Yeen,Yireun Kim,Jinsik Lee,Hwanhee Lee

类目:Computation and Language (cs.CL)

关键词:limiting models' ability, High-quality creative writing, High-quality creative, large language models, remains dominated

备注: CIKM 2026

点击查看摘要

Abstract:High-quality creative writing data for large language models (LLMs) remains dominated by story-centric data, limiting models' ability to follow the structural and functional conventions of diverse creative formats. We propose an attribute-guided genre expansion framework for scaling creative writing data beyond story generation. By separating thematic breadth from genre-form control, our framework leverages human-authored story prompts as diverse creative seeds, while utilizing manually curated genre attributes to enforce distinct structural, stylistic, and formatting conventions. We combine these to prompt strong LLMs for genre-faithful query-response pairs, which are then quality-filtered. Applying this framework, we construct the Multi-Genre Collection, a 50K-example corpus spanning 13 creative genres, including story, rap, lyrics, scripts, game design, character design, and other creative formats. Experiments across out-of-distribution writing benchmarks and held-out genre diagnostics demonstrate that models fine-tuned on our data consistently surpass not only base models and writing-specialized baselines, but also models trained on existing writing corpora. Genre-count ablations further indicate that controlled genre expansion, rather than story-centric scaling alone, is a key driver of robust creative writing capability.

29. 【2608.13926】Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

链接https://arxiv.org/abs/2608.13926

作者:Zhelun(Allen)Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB)

关键词:Large language models, mis-aggregated total yields, fluent wrong answer, made natural language, natural language interfaces

备注: 26 pages, 5 figures, 5 tables. Technical report. Describes architecture and design principles only; contains no code, schemas, datasets, or performance metrics

点击查看摘要

Abstract:Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indistinguishable at the point of use from a right one. Where the consumer cannot inspect the generated query, as in enterprise AI deployments and operational dashboards, and increasingly where the consumer is a tool-using agent rather than a person, accuracy alone is insufficient: nothing marks which answers to distrust. This is a reliability problem before it is an accuracy problem. We propose an architectural pattern for such systems, a trusted kernel with a generative shell, resting on one invariant: a component that can fabricate may influence which question the system answers, never which value it returns. A generative shell interprets underspecified input and phrases replies; a deterministic kernel matches fully specified questions against a bounded set of answerable question shapes and compiles them to queries by deterministic execution. The two meet at a confirmation the user reads before any value is computed, and requests the kernel cannot express are declined rather than approximated. We call this structural abstention, and distinguish it from the statistical abstention of selective prediction and calibrated confidence: refusal here needs no confidence estimate, because unanswerable requests are unrepresentable. We specify the pattern implementation-independently, give a five-decision recipe and work it across three domains, extend the invariant from returned values to the actions of agentic systems, and report a two-year production case study alongside two generative alternatives, a fine-tuned parser and a tool-retrieval agent. We close against enterprise and reliability benchmarks published since.

Comments:
26 pages, 5 figures, 5 tables. Technical report. Describes architecture and design principles only; contains no code, schemas, datasets, or performance metrics

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB)

ACMclasses:
H.2.3; H.5.2; I.2.7

Cite as:
arXiv:2608.13926 [cs.AI]

(or
arXiv:2608.13926v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2608.13926

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
30. 【2608.13925】CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

链接https://arxiv.org/abs/2608.13925

作者:Yuji Ren,Chenkai Xu,Zhuocheng Gong,Jianguo Li,Zhijie Deng

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Diffusion large language, accelerate language generation, Diffusion large, predicting multiple masks, single forward pass

备注

点击查看摘要

Abstract:Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass. However, existing dLLMs can suffer from unreliable predictions in early denoising stages under aggressive parallelism strategies, leading to errors that can propagate to later stages. To tackle this issue, we present Consistency Forcing (CForce) for dLLMs, a distillation method to force the mask predictions of early stages to align with those of later stages. CForce trains the model on pre-collected self-rollout trajectories, thereby improving training-inference alignment. We introduce Confidence Adaptive KL Divergence as a distillation objective to conjoin the merits of forward and reverse KL. We further provide a theoretical analysis for the consistency objective to explain why CForce can approximately minimize the prediction error of early stages. Critically, the same formulation applies to both mask-to-token decoding and edit-capable decoding; in the edit-capable case, later token-to-token refinements provide additional supervision for earlier masked-state predictions. Experiments on non-edit and edit-capable LLaDA models show improved speed-quality trade-offs, especially under high-parallelism decoding budgets. Code is available at: this https URL.

31. 【2608.13900】Agentic Transaction: Towards ACID-Compliant Agent Systems

链接https://arxiv.org/abs/2608.13900

作者:Zhaoyan Sun,Xiaoxiao Wang,Guoliang Li

类目:Databases (cs.DB); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large language model, execute long-horizon tasks, Large language, tasks through reasoning, workspace manipulation

备注

点击查看摘要

Abstract:Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation. As agents increasingly operate over persistent environments and multi-step workflows, they face challenges analogous to those addressed by transactional database systems: reliable execution, consistent outcomes, safe concurrency, and durable state management. We introduce the concept of an agentic transaction and propose an ACID-compliant agent system framework that reinterprets the classical ACID properties for agent execution through four semantic guarantees: Semantic Atomicity, Semantic Consistency, Semantic Isolation, and Semantic Durability. Together, these properties provide a principled foundation for building reliable agent systems despite model uncertainty and dynamic execution environments. To instantiate this framework, we develop an ACID-compliant data agent that realizes these guarantees through transactional exploration-execution-validation cycles, transactional skill hubs, confidence divergence-based validation, semantic dependency-aware isolation, and transaction-aware semantic state management. Experimental results on widely used benchmarks show that our system achieves a 10.6% improvement over state-of-the-art agents, including Claude Code. This work opens a broader research agenda on extending transactional principles and system architectures toward building trustworthy, scalable, and self-evolving AI agent systems.

32. 【2608.13866】Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification

链接https://arxiv.org/abs/2608.13866

作者:Benjamín Schindler,Gonzalo A. Ruz

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Large language models, Large language, correct class regions, synthetic training data, generate synthetic training

备注: 6 pages, 2 figures, to be published in IEEE LACCI 2026

点击查看摘要

Abstract:Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. We propose a geometric filtering framework that evaluates each LLM-generated sample by its Euclidean distance to real class examples in a sentence embedding space, selecting only geometrically consistent candidates. A soft weighting mechanism transforms filter scores into sample weights for classifier training. Evaluated across 13 datasets, 5 classifiers, 10 augmentation methods, and over 6,700 configurations, our method achieves +2.61 percentage points (pp) over SMOTE ($p0.0001$, Cohen's $d=0.95$, 88.9% win rate). The approach generalizes to named entity recognition (+9.26pp, 100% win rate) without filter modification, and is robust across 5 LLMs from 4 providers. A key finding is that the simplest distance-based filter consistently outperforms complex multi-criteria alternatives.

33. 【2608.13854】Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision

链接https://arxiv.org/abs/2608.13854

作者:Kouki Yuki,Jie Zeng,Kyoko Ogawa,Ryunosuke Ikeda,Yohei Kobashi,Takeshi Kojima,Ikuya Yamada,Yusuke Iwasawa,Yutaka Matsuo

类目:Computation and Language (cs.CL)

关键词:preserve executable behavior, preserve executable, executable behavior, largely focused, neural code translation

备注: 11 pages, 3 figures, 5 tables. Preprint under review

点击查看摘要

Abstract:Code translation must preserve executable behavior across many programming languages, yet neural code translation has largely focused on a few popular languages such as C++, Java, and Python. This leaves a niche, many-to-many setting where parallel supervision is sparse, producing plausible but non-executable translations. We address this setting with preference-based reinforcement learning driven by execution-based supervision. Our pipeline firstly expands verifiable seed Python programs into a multilingual pool of execution-validated codes. Using the pool, a base LLM generates translation candidates across language pairs, which we label by their execution outcomes. The resulting preferences are used to train a reward model that scores cross-language translation quality. Finally, we optimize our base LLMs with GRPO over 600 directed language pairs (25 x 24) using the reward model as a signal. To evaluate the niche translation capability, we introduce HumanEval-X++, an execution-based benchmark that extends HumanEval-X to a broad many-to-many language space. We evaluate our approach using Qwen-3.5 4B and 9B models. On HumanEval-X++ and existing benchmarks, it yields consistent gains over the untrained baselines. In particular, the 4B model achieves an average improvement of 13% across all languages on HumanEval-X++, with a gain of 21% on mid-tier languages. Our study establishes a reliable approach of data generation, training, and benchmarking, paving the way toward further bootstrapping the quality of many-to-many translation for programming languages.

34. 【2608.13840】ASSERT: A Measurement Pipeline for GenAI Audits

链接https://arxiv.org/abs/2608.13840

作者:Riccardo Fogliato,Abhinav Palia,Xiawei Wang,Emily Sheng,Chad Atalla,Jean Garcia-Gathright,Nicholas Pangakis,Sharman Tan,Dan Vann,Hannah Washington,P. Alex Dow,Heba Elfardy,Hanna Wallach,Sandeep Atluri

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)

关键词:reported rate, audited system complies, reported, rate, complies with policy

备注

点击查看摘要

Abstract:Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate to compare systems, track regressions, and gate deployment. A reported rate reflects both the system under audit and the measurement choices behind it, so a change in the rate can leave it unclear whether the system or those choices moved. We introduce ASSERT, a specification-driven measurement pipeline for GenAI audits that ties each reported rate to a written specification of the measurement choices used to produce it. ASSERT helps draft a behavioral rubric and test cases, then runs the audit against a GenAI system and returns a reported rate. In a case study on conversational deception, we observe that the reported rate moves substantially with the dialogue setup, the simulated user, the judge, and the evidence bar for non-compliance. These measurement choices substantially change the reported rate and can reorder GenAI system rankings. Because each reported rate is tied to an explicit specification, differences across audits are easier to attribute and interpret.

35. 【2608.13835】When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics

链接https://arxiv.org/abs/2608.13835

作者:Charu Karakkaparambil James

类目:Computation and Language (cs.CL)

关键词:Dynamic topic models, evolving word distributions, models capture evolving, capture evolving word, topic models capture

备注

点击查看摘要

Abstract:Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($\rho$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($\rho$=0.609), DBLP ($\rho$=0.721), and arXiv ($\rho$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.

36. 【2608.13787】From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL

链接https://arxiv.org/abs/2608.13787

作者:Wenyue Hua,Zachary Huang,Tyler Payne,Safoora Yousefi,Saleema Amershi,Asli Celikyilmaz

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

关键词:agents increasingly act, comparing offers, users' behalf, scheduling meetings, haggling over prices

备注: 25 pages, 3 figures

点击查看摘要

Abstract:AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller, a recruiter) whose goals may conflict with its principal's. Yet the dispositions that make an assistant pleasant can make it a poor delegate: a friendly, helpful frontier model may disclose its principal's private information unprompted and concede at the first sign of resistance. We present SocialRL, a general recipe that trains social reasoning directly, and apply it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace. Every domain is trained in-domain under the same recipe, and every policy is evaluated on all six. We find that (1) in-domain training reaches the frontier: on held-out scenarios the 4B matches or exceeds the GPT-5 family per domain, closing 73-122% of the baseline-to-frontier gap on the negotiation games, with 78% of buyer openings anchoring below target versus 3% untrained; (2) cross-domain transfer follows game structure: structurally paired games lift each other, a broad multi-issue donor lifts nearly all domains, and structurally isolated games transfer nothing; (3) guided by this transfer structure, two strategies, cascade RL and multi-teacher on-policy distillation (OPD), consolidate the per-domain specialists into a single unified 4B that reaches 0.627 average utility across all six environments, matching or exceeding GPT-4.1 (0.625), GPT-5.1 (0.619), and GPT-5.2 (0.613); (4) an explicit theory-of-mind scaffold helps only through training: distilling the ToM trace, rather than actions alone, lifts utility on every environment and generalizes better across them, and of the two ToM skills, only next-action prediction predicts negotiation outcomes.

37. 【2608.13786】Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

链接https://arxiv.org/abs/2608.13786

作者:Qingfang Liu,Qiao Jin,Joe D. Menke,Thorsten Kahnt,Zhiyong Lu

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language model, Large language, LLM chatbots, Large, Cochrane

备注

点击查看摘要

Abstract:Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.

38. 【2608.13760】Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

链接https://arxiv.org/abs/2608.13760

作者:Jean de Dieu Nyandwi,Leena Mathur,Yonatan Bisk,Robert Hawkins,Graham Neubig

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:reasoning-oriented training, correct answers, behaviors, Behavioral Lift, reasoning-oriented training amplify

备注: Published in COLM 2026

点击查看摘要

Abstract:Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

39. 【2608.13741】GALA: Generation-Aware Cross-Modal Alignment for Text-to-Time-Series Synthesis

链接https://arxiv.org/abs/2608.13741

作者:Haochen Zhang,Gengwei Zhang,Laura Yao,Nicholas Knoz,Tianlong Chen

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Synthesizing time series, controllable time series, time series generation, Synthesizing time, time series

备注: 21 pages, 6 figures

点击查看摘要

Abstract:Synthesizing time series from natural language is emerging as the most expressive form of controllable time series generation. However, existing text-conditioned generators either take caption embeddings frozen from off-the-shelf text encoders, or adapt the encoder end-to-end, letting the denoising loss shape the embeddings only as a by-product. In either case, the conditioning representation is never deliberately matched to the signal modality, leaving it ill-suited to guide generation. We address this by introducing GALA: Generation-Aware cross-modaL Alignment for text conditional time series generation. GALA is a two-stage approach that first contrastively couples a pretrained text encoder with a time-series foundation model into a shared embedding space with both encoders adapted to generation by an auxiliary generative loss, and then freezes the resulting caption embedding to drive a flow-matching generator. On TSFragment-600K, spanning four domains and three fragment lengths, GALA sets a new state of the art, ranking first in 30 of 36 metric columns and reaching an average rank of 1.08/1.08/1.42 at lengths 24/48/96 against 1.92/2.00/1.75 for the strongest baseline. We further find that generator-internal text encoders force a trade-off between fidelity and caption adherence, whereas conditioning on the aligned embedding breaks it: FID, CTTP, and JFTSD all improve at once. Ablating the auxiliary loss degrades FID, CTTP and JFTSD together, it indicates the generative term is a necessary component of the alignment rather than an add-on.

40. 【2608.13722】BM25-Augmented Many-Shot Translation for Low-Resource North-Eastern Indian Languages

链接https://arxiv.org/abs/2608.13722

作者:Aashish Dhawan,Christopher Driggers-Ellis,Dzmitry Kasinets,Christan Grant,Daisy Zhe Wang

类目:Computation and Language (cs.CL)

关键词:Low-Resource Indic Language, Florida Gators submission, Indic Language Translation, University of Florida, Florida Gators

备注

点击查看摘要

Abstract:This paper describes the University of Florida Gators submission to the WMT26 Low-Resource Indic Language Translation shared task. We adapt the retrieval-augmented many-shot translation pipeline from our AmericasNLP 2026 system to translate between English and eleven North-Eastern Indian languages in both directions. At inference time, BM25 retrieves the most similar parallel examples from a language-specific training bank, and Gemini 2.5 Flash translates the input conditioned on these examples. No model fine-tuning is involved. Training banks combine official WMT26 data with publicly available corpora such as Samanantar and prior WMT shared task releases. A grid search over retrieval count r and development exemplar count d across all 22 language-direction pairs selects the best configuration for each submission.

41. 【2608.13721】Capacity-Dependent Effects of Data Selection for Reasoning

链接https://arxiv.org/abs/2608.13721

作者:Cuong Dang,Hoang Anh Just,Ruoxi Jia

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:reasoning supervised fine-tuning, instruction can differ, differ substantially, supervised fine-tuning, student current distribution

备注: Accepted to COLM 2026

点击查看摘要

Abstract:In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that responses closer to the student distribution provide more effective supervision, motivating the hypothesis that high-likelihood responses may generally be preferable for fine-tuning. In this paper, we revisit this intuition and show that the value of likelihood-based data selection depends critically on model capacity and training duration. Through controlled experiments on mathematical reasoning, using students ranging from 1.5B to 8B parameters and supervision generated by stronger teacher models, we observe a clear \emph{capacity-dependent} ``{\color{SMALLCOLOR}\textbf{Fast-Fit}} / {\color{LARGECOLOR}\textbf{Slow-Gain}}'' pattern. High-likelihood data provides faster and more stable early improvements, especially for smaller models, but low-likelihood data becomes increasingly beneficial for larger models when training is allowed to continue longer. To explain this phenomenon, we analyze learning dynamics, showing that small models often fail to absorb low-likelihood supervision and instead fall into shallow or repetitive behaviors, while larger models are better able to move toward the teacher distribution under such data. We further provide a capacity-constrained theoretical view of distillation that clarifies how data difficulty, data span, and student capacity jointly govern transfer. Overall, our findings show that effective data selection for reasoning should be aware of model capacity and computing budget rather than based on a single universal preference for high-likelihood supervision.

42. 【2608.13717】StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition

链接https://arxiv.org/abs/2608.13717

作者:Zefang Liu,Chenyang Zhu,Sangwoo Cho,Xujun Peng,Shi-Xiong Zhang,Sambit Sahu

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)

关键词:domain-shifted target audio, labeled in-domain data, automatic speech recognition, Streaming automatic speech, underperforms on domain-shifted

备注

点击查看摘要

Abstract:Streaming automatic speech recognition (ASR) underperforms on domain-shifted target audio, where labeled in-domain data is costly to prepare while unlabeled audio is abundant. We present StreamHear, a semi-supervised pipeline that adapts a pretrained streaming student by fine-tuning an offline transducer teacher on the labeled training set, generating pseudo-labels on the unlabeled portion, and fine-tuning the student on the mixture. We further introduce a prior-regularized dynamic-programming realignment step that fixes chunk-level word placement using an ASR-hypothesis anchor. Across four datasets spanning financial calls, prepared read speech, and phone-quality dialogue, StreamHear consistently outperforms supervised student fine-tuning and narrows the gap to the offline teacher.

43. 【2608.13712】Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

链接https://arxiv.org/abs/2608.13712

作者:Divya Vetticaden,Arya Gupta,Julian Nyarko,Megan Ma

类目:Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Deposition training requires, legal-AI evaluations largely, evaluations largely focus, manage dynamic witness, factual accuracy

备注: Accepted at ICML 2026 AI4Law. 47 pages, 26 tables, 14 figures

点击查看摘要

Abstract:Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.

44. 【2608.13708】achMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Generation from Science Curriculum Materials

链接https://arxiv.org/abs/2608.13708

作者:Fatema Tuj Johora Faria,Mukaffi Bin Moin,M. F. Mridha,Jubayer Al Mahmud

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Automatically generating textbook-grounded, generating textbook-grounded assessment, Automatically generating, existing retrieval-augmented generation, science teachers' workload

备注

点击查看摘要

Abstract:Automatically generating textbook-grounded assessment items can reduce science teachers' workload, but existing retrieval-augmented generation (RAG) systems rely on flat retrieval, support only single-question generation, lack safeguards against weak evidence, and are ill-suited to low-resource, board-exam-structured curricula. We address these limitations with TeachMateGPT, a multi-agent system contributing four advances to curriculum-grounded science-assessment authoring. (i) COPE, a hierarchical knowledge base replacing token-window chunking with a multi-resolution index that segments documents along syllabus structure and links them at three granularities via a traversable graph-based lineage, matching evidence to each topic's instructional level. (ii) A staged, fail-closed agent pipeline replacing one-shot retrieve-then-generate: routing gates search, retrieval fuses dense and lexical evidence under a coverage gate that withholds generation on insufficient evidence, and specialist agents draft objective and constructed-response items. (iii) SAVER, a source-attributed verification protocol scoring faithfulness, relevance, and hallucination risk against retrieved evidence, applying stricter grounding checks across each creative question's four sub-parts, paired with teacher-in-the-loop evaluation rather than automatic filtering. (iv) NCTB-SciGen8, a curriculum-grounded dataset of 198 items (143 multiple-choice, 55 creative questions) spanning all 14 chapters of the NCTB Class 8 science textbook, produced by the pipeline and rated by three practicing teachers. TeachMateGPT raises faithfulness (0.68 $\rightarrow$ 0.96) and answer relevancy (0.60 $\rightarrow$ 0.89) over a vanilla RAG baseline.

45. 【2608.13706】CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

链接https://arxiv.org/abs/2608.13706

作者:Fatema Tuj Johora Faria,Mukaffi Bin Moin,Jubayer Al Mahmud,M. F. Mridha,Md. Alam Hossain

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:leaving inter-agent errors, inter-agent errors undetected, pipelines remain partial, Existing defenses, multi-agent pipelines remain

备注

点击查看摘要

Abstract:Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modality disagreement, debate verifies an aggregate report rather than individual claims, and such verification occurs only after drafting, leaving inter-agent errors undetected until the final text. To close this gap, we present CLAIR-Fin, a nine-agent framework that decomposes each question into atomic claims maintained in a typed Financial Claim Ledger. Each claim is resolved through Asymmetric Evidence Authority, which conditions evidence trust on claim type rather than treating all modalities as equally reliable; Chain-of-Custody Verification, which checks grounding at the hand-off between drafting and adversarial review rather than only at the pipeline's exit; an Adaptive Rebuttal Cycle, which routes contested claims through adversarial debate whose depth scales with what that debate finds; and a terminal entailment audit paired with a continuous Hallucination Risk Index that distinguishes claims that passed scrutiny from claims never contested. We evaluate CLAIR-Fin on BB-FinQA-X, a 500-question cross-modal financial evaluation set built from Bangladesh Bank Annual Report material, stratified by query type, format, and difficulty. Relative to a single-pass retrieval-augmented generation baseline, it raises faithfulness ($0.780 \rightarrow 0.889$) while abstaining on 5.4% of questions when evidence is insufficient rather than forcing an unsupported response, and it exceeds stronger retrieval-strategy baselines such as HyDE and Graph-RAG on faithfulness ($\leq 0.874$).

46. 【2608.13698】GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings

链接https://arxiv.org/abs/2608.13698

作者:Konstantin Dobler,Federico Scozzafava,Jonathan Janke,Mohamed Ali,Simon Lehnerer

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Relative Policy Optimization, Group Relative Policy, remain heavily English-centric, Reinforcement Learning, Policy Optimization

备注

点击查看摘要

Abstract:Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others. However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages. Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.

47. 【2608.13674】Asymmetric Discourse Homogenization and Shared Language Technology: Evidence from Reddit

链接https://arxiv.org/abs/2608.13674

作者:Fengming Liu

类目:Computers and Society (cs.CY); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:million Reddit comments, ideologically asymmetric break, million Reddit, Reddit comments, pre-existing diversification trend

备注: 36 pages, 5 figures, 6 tables. Replication package available (code and instructions). Keywords: political discourse; semantic similarity; discourse homogenization; generative AI; computational text analysis

点击查看摘要

Abstract:I document an ideologically asymmetric break in the pre-existing diversification trend of political discourse, emerging around late 2022, using 6 million Reddit comments from two cross-partisan forums, 2019-2025. Conservative users experienced an interruption of their prior diversification trajectory; progressive users showed no comparable change. The asymmetry is consistent across estimation strategies (ITS, DiD, RDiT, propensity-score matching) and temporal aggregations. A daily-frequency permutation test over 2,377 candidate cutoff dates shows the ChatGPT threshold produces an unremarkable estimate (49.8th percentile): the shift builds gradually instead of breaking at a single date. A continuous cumulative LLM index, tracking AI exposure across seven model releases, remains significant under a quadratic trend specification that eliminates the binary estimate. A stayer analysis narrows the mechanism: the homogenization effect disappears when the sample is restricted to authors active throughout the study period, and the stayer confidence interval excludes within-author effects even a tenth the size of the full-sample estimate. The mechanism is most parsimoniously ecological (community-level discursive convergence) rather than individual-level AI adoption, though the data cannot cleanly separate this account from concurrent secular change.

48. 【2608.13626】A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure

链接https://arxiv.org/abs/2608.13626

作者:Dekun Yang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:decodable or causally, causally usable, usable without supporting, action maps fitted, reusable action map

备注: 12 pages, 7 figures, 4 tables; includes supplementary results and ancillary reproducibility files

点击查看摘要

Abstract:A hidden state signal can be decodable or causally usable without supporting a reusable action map. We test whether action maps fitted without a source reach its natural post-action activation and compose. We organize the tests as an evidence lattice and validate the geometric branch on a known affine S_5 carrier: all held-source folds pass one-step, composition, inverse, decoding, and commutativity gates. Structured curvature and held-domain conjugacy raise error monotonically, but only 23/30 strongest cells flip a closure gate, bounding rather than universalizing calibration. In post-trained Qwen/Qwen3-4B, frozen final-token h28 affine maps have mean held-entity error .519, versus .398 for within-test-domain cross-fit. Seven randomized entity splits and map geometry do not support a purely entity-specific account. Earlier h4/h16 layers fit one-step transitions better, but h4 conflict-state decoding is weak and lexical controls remain unresolved. Three matched intervention datasets regenerated from one frozen checkpoint show causal effects only at h28/h36. Outcome-aware refitting improves h28 one-step error to .474 (.469 with weighting), yet no refit passes composition. Learned finite worlds likewise preserve relative algebraic signals or shared charts without held-source affine closure. Within the tested carriers, state availability, causal use, local geometry, and reusable closure are separable. The result is limited to one pretrained model, sampled final-token layers, two finite worlds, and the tested affine or diagnostic function classes.

49. 【2608.13624】Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation

链接https://arxiv.org/abs/2608.13624

作者:Zhe Liu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)

关键词:Large Audio Language, audio question answering, Audio Language Models, audio understanding tasks, Large Audio

备注

点击查看摘要

Abstract:Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation in spoken-input settings is challenging due to confounding factors, including semantic variation in spoken content and speaker-specific characteristics. Ignoring these factors can result in misleading conclusions about model bias. We propose a semantic-aware mixed-effects regression framework for fairness evaluation in LALMs that explicitly accounts for these confounders. Our approach incorporates sentence-level semantic embeddings of reference text as covariates and models speaker identity as a random effect. Notably, semantic representations are extracted from the same LALM under evaluation, enabling semantic control over variation as perceived by the model itself. Experiments on simulated data and real-world benchmarks demonstrate that the proposed approach substantially reduces spurious fairness findings and yields more robust and interpretable estimates of subgroup performance differences.

50. 【2608.13622】ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

链接https://arxiv.org/abs/2608.13622

作者:Yongqi Tong,Tan Li Hui Faith,Choy Zhen Wen Marcus,Zhou Jin,Kewei Fu,Jiang-Ming Yang,Jianshe Li,Xin Zhang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:admits multiple valid, provide progress updates, real-world interaction admits, interaction admits multiple, multiple valid behaviors

备注

点击查看摘要

Abstract:Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a \textit{reward fairness problem} and propose \textbf{ARC} (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core $\tau/\tau^2$ tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.

51. 【2608.13607】No Universal Signal Predicts Sample-Level LLM Regression under Version Updates

链接https://arxiv.org/abs/2608.13607

作者:Jia Sheng,Yiwei Lu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Frontier LLMs, LLMs are updated, updated frequently, frequently and typically, typically outperform

备注

点击查看摘要

Abstract:Frontier LLMs are updated frequently and typically outperform their predecessors in aggregate. But aggregate gains say little about individual samples: an update can still cause sample-level regression, where a response correct under the old model becomes incorrect under the new one. This paper studies how to predict such regressions from signals available at inference time. We compare single-model signals (confidence, logit margin, attention entropy) against cross-version signals (output KL divergence, likelihood drift, token-level KL, representation drift) under a unified added-value test that isolates each signal's gain over a confidence baseline. Across six benchmarks in three task families (multiple-choice question answering, or MCQ; math reasoning; code generation) and six model update pairs, we find that (1) signal effectiveness is task-dependent: confidence is strongest on MCQ and simpler math, while likelihood/KL signals give the most frequent gains on harder math and code; (2) no signal is universally best across model updates either; and (3) some cross-version signals stay informative even when confidence fails, including without labels, which supports a proof-of-concept selective fallback that routes high-risk samples back to the old model. Practitioners can use these task-level patterns to choose which regression signal to trust for a given update. Code is available at this https URL.

52. 【2608.13606】MobileMem: Learning from a Year of Mobile Experiences

链接https://arxiv.org/abs/2608.13606

作者:Xinle Deng,Yida Xue,Xiangyuan Ru,Haoming Xu,Shuofei Qiao,Mengru Wang,Yijun Chen,Buqiang Xu,Chen Jiang,Yuchen Eleanor Jiang,Lizhong Wang,Jianfeng Wang,Li Zeng,Haofen Wang,Guilin Qi,Huajun Chen,Ningyu Zhang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA); Multimedia (cs.MM)

关键词:persistent personal assistants, answer isolated questions, increasingly moving, moving beyond systems, systems that answer

备注: Technical Report; Project Page: [this http URL](http://mobilemem.openkg.cn/)

点击查看摘要

Abstract:The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. Such assistants require long-term memory to accumulate and leverage user-specific experiences over time, yet existing benchmarks remain inadequate for realistic mobile settings, where experiences are heterogeneous, multimodal, evolving, and deeply personal. We introduce MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences. MobileMem employs a knowledge-grounded synthesis pipeline to construct coherent and temporally consistent long-horizon trajectories from user-app sessions. It provides complementary text and multimodal settings covering multi-hop and temporal reasoning, knowledge updating, and implicit preference inference. Specifically, MobileMem enables agents to remember the past, understand the present, and adapt to the future. By modeling experiences rather than isolated facts, MobileMem moves memory beyond information retrieval toward experiential intelligence for continuous personal learning.

53. 【2608.13604】Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents

链接https://arxiv.org/abs/2608.13604

作者:Babak Abbaschian

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)

关键词:in-person interaction, AI-mediated channels, urgent problem, problem to solve, increasingly handled

备注: 49 pages, 2 figures, 8 tables, 94 references. Cross-disciplinary conceptual synthesis across multiple fields. Includes a source-by-source evidence matrix in Appendix A and a coding manual in Appendix B for independent application of the taxonomy

点击查看摘要

Abstract:Detection of misunderstanding is an urgent problem to solve because communication has moved away from real-time, in-person interaction and is increasingly handled by AI-mediated channels. This shift cuts communicators off from the resources repair depends on faster than new means of detection are being built. In this paper we analyse misunderstanding as a layered process in which a divergence is generated, may then be amplified, and is either detected and repaired or left to persist unnoticed. Consolidating accounts from nine fields of research that do not ordinarily cite one another, we identify eleven exact failure modes and show that each operates at a specific point in a communicative process rather than anywhere within it. Those points give eight analytical layers, derived from the literature rather than adopted from an existing model. Eight of the mechanisms primarily generate a divergence, two primarily amplify one already present, and one governs whether a divergence is detected and repaired. We model the eight layers formally, extending information and communication theory from the transmission of signals to the reconstruction of meaning, and we supply a source-by-source evidence matrix that makes every rating auditable, a coding manual, and nine analysed dialogue cases. No prior classification of misunderstanding both locates mechanisms at points in the process and types them by function.

54. 【2608.13591】Stable Miscalibration in Large Language Models: A Practical View of High-Confidence Errors

链接https://arxiv.org/abs/2608.13591

作者:Akira Okutomi

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:large language models, large language, treated as evidence, fragile internal inference, internal inference

备注: Accepted at the 2nd Workshop on Epistemic Intelligence in Machine Learning (EIML)@ICML 2026. 7 pages, 4 figures

点击查看摘要

Abstract:High-confidence errors in large language models are often treated as evidence of fragile internal inference. We study a different possibility: stable miscalibration, where a confident wrong answer remains locally stable under small perturbations. We combine two diagnostics: a label-aware output-level audit score that ranks domains by confidence variation and overconfident mistakes under a forced-answer baseline, and an internal sensitivity probe that measures hidden-state movement. On a multi-domain binary factual audit set, this audit score tracks where abstention-aware self-critique reduces decision loss, although direct labeled baselines rank the same gain more strongly. Internally, self-critical prompting consistently reduces hidden-state sensitivity across layers in three open-weight models. This supports prompt-induced local stabilization rather than a purely output-level abstention pattern, but it does not imply calibration: audit-defined overconfident errors are not clearly more locally sensitive than confidently correct answers, so some high-confidence errors may be stable and miscalibrated rather than simply fragile.

55. 【2608.13588】IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering

链接https://arxiv.org/abs/2608.13588

作者:JungMin Yun,YoungBin Kim

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:answering requires complex, overwhelms retrieval-augmented generation, retrieval-augmented generation systems, requires complex reasoning, question answering requires

备注: ACL 2026 Main Conference

点击查看摘要

Abstract:Multi-hop question answering requires complex reasoning across multiple evidence segments, which often overwhelms retrieval-augmented generation systems with lengthy and noisy contexts, thereby undermining both efficiency and accuracy. While existing prompt compression methods attempt to address this issue, they are typically designed for single-turn queries and fail to capture interdependent reasoning steps. We propose IterCOMP, a unified, training-free prompt compression framework that incorporates multi-hop reasoning within an iterative compression loop. IterCOMP decomposes documents into evidence segments, evaluates question answerability, and generates targeted follow-up questions to iteratively integrate essential evidence, producing a compact, reasoning-oriented prompt. Experiments on MusiQue, 2WikiMultiHopQA, and HotpotQA demonstrate that IterCOMP achieves substantial improvements in Exact Match and F1 scores while reducing the token budget, outperforming existing baselines and exhibiting robustness as reasoning complexity increases.

56. 【2608.13580】Jais 2: A Family of Arabic-Centric Open Large Language Models

链接https://arxiv.org/abs/2608.13580

作者:Mohamed Anwar,Abed Alhakim Freihat,George Ibrahim,Mostafa Awad,Abdelrahman Sadallah,Gurpreet Gosal,Gokulakrishnan Ramakrishnan,Sarath Chandran,Biswajit Mishra,Rituraj Joshi,Ahmed Frikha,Etienne Goffinet,Abhishek Maiti,Ali El Filali,Sarah AlBarri,Samujjwal Ghosh,Rahul Pal,Parvez Mullah,Awantika Shukla,Sajid siddiki,Samta Kamboj,Onkar Pandit,Sunil Kumar Sahu,AbdelRahman Elbadawy,Amr Mohamed,Ahmad Chamma,Evan Dufraisse,Abdelaziz Bounhar,Dani Bouch,Hadi Abdine,Guokan Shang,Fajri Koto,Yuxia Wang,Zhuohan Xie,Ali Mekky,Rania Elbadry,Sarfraz Ahmad,Momina Ahsan,Omar El Herraoui,Daniil Orel,Hasan Iqbal,Kareem Elzeky,Mervat Abassy,Kareem Elozeiri,Saadeldine Eletter,Farah Atif,Nurdaulet Mukhituly,Haonan Li,Xudong Han,Aaryamonvikram Singh,Zainul Abedien Ahmed Quraishi,Neha Sengupta,Larry Murray,Avraham Sheinin,Joel Hestness,Natalia Vassilieva,Hector Xuguang Ren,Zhengzhong Liu,Michalis Vazirgiannis,Preslav Nakov

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Arabic-centric language modeling, Arabic-centric large language, advance Arabic-centric language, large language models, language models developed

备注

点击查看摘要

Abstract:Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest open Arabic-centric LLM trained from scratch at 70B parameters, and a competitive 8B-parameter variant among the evaluated open models. A custom Arabic-centric vocabulary enables efficient training and inference. In addition, an optimized architecture and training recipe yield highly compute-efficient training. With a substantially smaller token budget than comparable models, Jais 2 achieves strong Arabic performance on the benchmarks considered in this report and competitive English results. The models obtain leading results among the evaluated open models on OALL2 and AraGen. They also perform strongly on several culturally grounded Arabic benchmarks, including poetry, religion, cuisine, and dream interpretation, as well as in general tasks such as translation and summarization. We release the models in HuggingFace under a commercially permissive license. Jais 2 70B is also released as a chat app on the Web, iOS, and Android; it runs on Cerebras hardware, delivering up to 2,000 tokens per second, and enabling high-throughput Arabic-centric chat serving in our deployment setting. By uniting scale, linguistic diversity, cultural fidelity, openness, and speed, Jais 2 provides an open-weight foundation intended to support further research and development in Arabic-centric LLMs.

57. 【2608.13578】BCMT: Blockwise Causal Memory Transformer

链接https://arxiv.org/abs/2608.13578

作者:Rachid Arezki

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:exhibits quadratic complexity, model long-range dependencies, Transformer architectures rely, mechanism exhibits quadratic, Blockwise Causal Memory

备注: 19 pages. Official implementation: [this https URL](https://github.com/rachidlabs/BCMT)

点击查看摘要

Abstract:Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length. We introduce BCMT (Blockwise Causal Memory Transformer), an architecture for long-context language modeling that decouples local token interactions from global context propagation. Dense causal self-attention is applied independently within local blocks, while each block produces an adaptive summary aggregated through an exponential causal memory. This memory is subsequently injected back into the token representations, enabling efficient propagation of long-range contextual information without relying on explicit global attention. Unlike standard Transformers and recurrent memory architectures, BCMT maintains neither dense interactions between distant tokens nor learned memory states. Its memory mechanism is fully parallelizable and remains compatible with standard implementations of dense self-attention. Experiments on language modeling with context lengths of up to 1024 tokens show that BCMT achieves validation performance comparable to that of Dense Transformers while significantly improving training throughput and reducing memory consumption. An ablation study further confirms that these improvements arise from the proposed memory mechanism. These results demonstrate that an exponential causal memory constructed from block summaries provides an effective alternative to dense global attention mechanisms for long-context language modeling.

58. 【2608.13571】Not All Tokens Are Equal: Inflation-Aware Routing for Agentic LLM Systems

链接https://arxiv.org/abs/2608.13571

作者:Heming Fu,Shan Lin,Qianqian Xie,Guojun Xiong

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:language model fails, consuming additional tokens, agentic system retries, consuming additional, fails to answer

备注

点击查看摘要

Abstract:When a language model fails to answer a query on the first attempt, an agentic system retries, consuming additional tokens each time. This retry overhead creates a gap between what a model's per-token price implies and what a full workflow actually costs. We call this gap \emph{token inflation} and define it as the ratio of true workflow cost to single-call cost. Systems like FrugalGPT route based on the latter, which can underestimate real cost by more than $2\times$ on difficult tasks. We address this with InflationAgent, a four-stage router that (1) measures token inflation systematically across model tiers and task types, finding inflation as high as $4.25\times$ for a 7B model on multi-hop question answering; (2) introduces CoT Branching Entropy (CBE), a pre-execution difficulty signal computed entirely from local inference, which predicts high inflation with AUROC 0.887; and (3) selects models by maximizing a Semantic Exchange Rate (SER) that divides expected accuracy by predicted true cost, with a fresh-escalation policy that discards failed chains before routing to a stronger model. On GSM8K under a fixed budget, InflationAgent achieves 94.7\% accuracy versus 91.0\% for FrugalGPT while using 31\% fewer tokens, and we show that forwarding a failed reasoning chain to GPT-4o reduces its accuracy by up to 34.8 percentage points, validating the fresh-escalation design.

59. 【2608.13570】hink in Latent, Explain in Language: Self-Explainable Latent Reasoning

链接https://arxiv.org/abs/2608.13570

作者:Dayuan Zhao,Shengcao Cao,Yu-Xiong Wang,Liang-Yan Gui

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:offering significant gains, Latent reasoning, compressing verbose reasoning, Latent, alternative to text-based

备注

点击查看摘要

Abstract:Latent reasoning has emerged as a powerful alternative to text-based Chain-of-Thought (CoT), offering significant gains in computational efficiency by compressing verbose reasoning into compact embeddings. However, compressing reasoning into the latent space renders the thinking opaque, hindering its interpretability. Current methods present a stark trade-off: they either function as unexplainable ''black boxes'' (e.g., Coconut), where the latent reasoning is not human-readable, or rely on separate post-hoc decoders for explainability (e.g., Heima), introducing architectural overhead and decoupling the explanation from the actual reasoning process. In this work, we present a unified framework for Self-Explainable Latent Reasoning (SELR) that trains a single model to perform efficient and inherently explainable latent reasoning. Our core contribution is a novel multi-task training objective that optimizes for two goals simultaneously: (1) an Answer Loss that optimizes the latent reasoning trajectory to produce accurate final answers, and (2) a CoT Loss that explicitly trains the same model to decode its own latent representations back into human-understandable reasoning steps. This design ensures that generated latent representations are both task-effective and semantically interpretable, eliminating the need for external decoders. We validate the effectiveness of SELR on both Large Language Models (LLMs) and Vision-Language Models (VLMs), demonstrating that SELR achieves superior token efficiency and accuracy compared to baselines, while uniquely providing self-contained explainability without auxiliary models. Project page is available at this https URL.

60. 【2608.13568】Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preliminary Study

链接https://arxiv.org/abs/2608.13568

作者:Pengcheng Xu

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Coding agents spend, Coding agents, context budget, Language Server Protocol, Semantic retrieval

备注: 13 pages, 6 figures. Code and data: [this https URL](https://github.com/Poytr1/lsp-vs-grep-token-study)

点击查看摘要

Abstract:Coding agents spend most of their context budget on retrieval. Lexical retrieval (grep) is universal, instant, and zero-setup, but noisy: it cannot tell a definition from a call from a comment. Semantic retrieval via the Language Server Protocol (LSP) is precise and typed, but needs a running, indexed server and pays a per-symbol round-trip. The claim that semantic retrieval is more token-efficient is, we find, asserted almost everywhere and measured almost nowhere: no public source isolates the LSP-vs-lexical token delta for an agent at equal task-success. This paper formalizes the question with one metric (tokens-to-success), specifies a five-arm ablation isolating semantic retrieval from confounds, maps three pre-stated failure modes onto measurable variables, and reports a preliminary study (Python and TypeScript repos; Claude Opus 4.8, Sonnet 4.6, Haiku 4.5). The answer is conditional and usually negative. On symbol-named localization the LSP costs tokens (+6% to +118%) and the agent ignores it when free. On reference-completeness it buys precision but not token savings and cannot raise the recall ceiling set by agent thoroughness; it saves tokens only for the weakest model. Tool choice is task-dependent: models default to grep on localization (0-6% semantic use) but reach for the LSP about half the time on reference tasks, unprompted. On edits scored by real test execution the gap is starkest: grep solves multi-file renames perfectly, a location-only LSP fails three-quarters of them by missing a call site, and even a complete, index-warmed, text-enriched LSP (each reference's line inline, as production LSP-MCP servers do) recovers most of the gap but cannot close it, since a rename must touch comments and strings that semantic references exclude. The implication is not LSP-always but an adaptive router keyed on task class, model capability, and lexical noise.

61. 【2608.13567】Modular Cognitive Architecture Emerges in Large Language Models

链接https://arxiv.org/abs/2608.13567

作者:Pengrui Han,Jacob Andreas,Evelina Fedorenko,Andrea Gregor de Varda

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:functional specialization, Large Language Models, exhibits a striking, striking degree, degree of functional

备注: [this https URL](https://pengrui-han.github.io/LLM_Modularity_Page/)

点击查看摘要

Abstract:The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization a fundamental principle of how intelligent systems must be built, or an evolutionary accident specific to biological brains? Here, we test whether a similar organization emerges in Large Language Models--another class of intelligent systems created through a very different optimization process. Using circuit analyses across N=46 tasks spanning four cognitive domains (language, formal reasoning, social reasoning, physical reasoning), we find that LLMs develop a modular architecture that mirrors the human brain: tasks drawing on the same network in humans recruit overlapping neurons in LLMs, whereas tasks drawing on different networks recruit distinct neurons. The convergent emergence of modularity in brains and neural networks suggests that it may be a fundamental property of intelligent systems.

62. 【2608.13831】VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

链接https://arxiv.org/abs/2608.13831

作者:Edresson Casanova,Jaehyeon Kim,Mariana Graterol Fuenmayor,Shehzeen Hussain,Viacheslav Klimkov,Valentin Mendelev,Mikyas Desta,Paarth Neekhara,Piotr Zelasko,Chen Chen,Elena Rastorgueva,Ke Hu,Ankita Pasad,Xuesong Yang,Aya Alja'fari,Rajarshi Roy,Rohan Badlani,Jason Roche,Jason Li,Zhehuai Chen

类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)

关键词:lack real-time adaptability, Spoken dialogue, language models remain, models remain limited, computer interaction

备注

点击查看摘要

Abstract:Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity synthesis must be optimized jointly. We propose VoiceChat-TTS, a low-latency, continuous, and streamable text-to-speech model for interactive agents. VoiceChat-TTS is driven directly by LLM text-token streams, supports explicit interruption via control tokens, and produces silence when no textual input is available. The model enables always-on, responsive speech generation while preserving modularity and high speech quality, and it supports mid-utterance interruptions without resetting the KV cache.

信息检索

1. 【2608.14429】PriCoRec: A Privacy-Aware Cloud-Device Collaborative Framework for Ad Recommendation under Feature Constraints

链接https://arxiv.org/abs/2608.14429

作者:Dairui Liu,Zhongyi Lu,Jitao Lu,Aghiles Salah,Mete Sertkan,Roger Zhe Li,Changhong Jin,Barry Smyth,Xingsheng Guo,Ruihai Dong

类目:Information Retrieval (cs.IR)

关键词:Privacy regulations increasingly, hindering traditional cloud-only, regulations increasingly restrict, increasingly restrict cloud, restrict cloud processing

备注: 5 pages, 1 figure. Accepted to RecSys'26

点击查看摘要

Abstract:Privacy regulations increasingly restrict cloud processing of sensitive user data (e.g., age, gender), hindering traditional cloud-only recommendation models. To mitigate this challenge, we propose a Privacy-aware Collaborative cloud-device ads Recommendation framework (PriCoRec) which personalizes recommendations while keeping sensitive features on-device. While separating recommendation into cloud-based and on-device stages enables privacy-aware deployment, naive splitting suffers from degraded shortlist quality and inefficient on-device inference due to limited private features. We therefore design a collaborative framework that comprises a cloud-based pre-ranking stage using cloud-accessible features, and an on-device ranking stage that locally incorporates highly personalized features. We introduce a diversity regularizer to pre-ranking to improve candidate quality. Moreover, to control device power consumption and computational cost, we incorporate a cloud-guided training mechanism that enhances device model performance while keeping the model lightweight. Experiments demonstrate that the proposed framework maintains strong recommendation performance while keeping sensitive features on-device.

2. 【2608.14068】MACS: A Hybrid Multi-Agent Framework for Reliable Conversational E-Commerce Recommendation

链接https://arxiv.org/abs/2608.14068

作者:Juli Huang,Hannah Clay,Sajjad Beygi,Thomas Sarda,Negin Golrezaei,Amin Saberi

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:large language models, real-world deployments operate, unsupported product claims, language models, merchant fixed catalog

备注: 9 pages, 2 figures, 8 tables. Will be presenting at Stanford TrustSafety Conference, already presented at Stanford Market AI Conference

点击查看摘要

Abstract:Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims. In this setting, the main challenge is reliability under hard constraints: the system must satisfy user requirements, remain grounded in available inventory, and preserve preferences across multiple conversational turns. We present MACS (Multi-Agent Commerce System), a hybrid multi-agent framework for reliable conversational recommendation in fixed-catalog settings. MACS uses LLMs for language-facing tasks such as interpreting user requests, eliciting preferences, and generating responses, while correctness-critical operations, including product retrieval, hard-constraint filtering, brand exclusion, and progressive relaxation, are executed deterministically by the merchant agent. A session-persistent preference layer tracks constraints across turns, enabling consistent handling of budget overwrites and exclusion reversals. On a 140-query single-turn benchmark, MACS achieves the highest pass rate (87.1%) and perfect brand compliance (1.000). On a 10-scenario multi-turn benchmark, MACS achieves the strongest macro Pass@5 (72% vs. 56% GPT+Catalog / 52% Gemini+Catalog) with zero constraint drift. The advantage is sharpest on exclusion reversal (100% vs. 20% / 0%) and constraint accumulation (100% vs. 60% / 40%). Mean judged response quality is similar across systems (0.751 vs. 0.736). These results suggest that hybrid architectures combining deterministic constraint enforcement with session-persistent preference tracking provide stronger reliability-oriented performance than catalog-bound prompt-only baselines in the fixed-catalog merchant setting.

3. 【2608.14066】nderKG

链接https://arxiv.org/abs/2608.14066

作者:Yacine Mokhtari(Lab-STICC_MOTEL, Lab-STICC, IMT Atlantique - INFO),Véra Pukhkoy,Grégory Smits(IMT Atlantique - INFO, Lab-STICC, Lab-STICC_MOTEL)

类目:Information Retrieval (cs.IR)

关键词:major economic activity, institutions allocate contracts, public institutions allocate, competitive tendering processes, Public procurement represents

备注

点击查看摘要

Abstract:Public procurement represents a major economic activity, where public institutions allocate contracts to companies through competitive tendering processes. Despite its importance, this domain remains underexplored by recommender systems, largely due to the lack of publicly available datasets capturing its complexity. In this paper, we introduce TenderKG, a large-scale knowledge graph dataset constructed from French public procurement data covering the period 2021--2023. The dataset models the procurement ecosystem through heterogeneous entities, including companies, tenders, lots, and domain-specific taxonomies of work domains, connected via rich semantic and structural relations. A key specificity of this setting is that only the awarded companies are visible, resulting in sparse explicit signals of awarded interactions. To overcome this limitation, TenderKG integrates extensive side information on the actors in the French tender market and the tenders, including textual descriptions, hierarchical classifications, and geographical features, enabling the study of knowledge-aware recommendation in a highly constrained and competitive environment. We provide detailed statistics and analyses of the dataset, highlighting its structural properties, sparsity patterns, and domain-specific characteristics. We believe TenderKG opens new research directions in bidder recommendation, knowledge graph-based recommendation, competition-aware matching, and provides a valuable benchmark for evaluating methods in real-world, high-stakes decision-making scenarios.

4. 【2608.14032】HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

链接https://arxiv.org/abs/2608.14032

作者:Yin Li,Ziyang Hu,Zhiyu Guo,Xiangyu Liu,Wenbin Li,Boo-Ho Yang,Rav Lawana,Ziyue Li,Wei Zeng,Fugee Tsung

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Existing multimodal RAG, multimodal RAG methods, text-image logic needed, multimodal RAG, Existing multimodal

备注

点击查看摘要

Abstract:Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement. We propose HAM-RAG, a Hierarchy-Aware Multimodal RAG framework for structure-faithful interleaved generation. HAM-RAG uses document hierarchy as a grounding signal across retrieval and generation, contextualizing textual and visual evidence and preserving source position and local text-image relations in the prompt. We further introduce HAM-Bench, covering Wukong, Wiki, arXiv, and Recipe across game walkthroughs, web pages, scientific papers, and step-wise recipe documents. Across multiple backbones, HAM-RAG improves the main multimodal average by 17.3% over the strongest non-hierarchical baseline. On Wukong, HAM-RAG improves Img-CBS by 24.2% over the strongest non-hierarchical baseline, demonstrating substantially better local text-image alignment. The main experiments and ablation study together demonstrate that document hierarchy is a key grounding signal for faithful image selection, placement, and local text-image alignment. These findings highlight the value of hierarchy-aware grounding for reliable multimodal assistants that generate answers faithful to the source organization, procedural structure, and local text-image evidence of structured documents, such as technical manuals, maintenance guides, and industrial SOPs. The code is available at this https URL.

5. 【2608.14021】Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders

链接https://arxiv.org/abs/2608.14021

作者:Keito Kozaki,Keigo Sakurai,Ren Togo,Takahiro Ogawa,Miki Haseyama

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:Transformer-based sequential recommenders, prediction remains unclear, Transformer-based sequential, remains unclear, sequential recommenders

备注: 11 pages, 6 figures. Accepted at the 20th ACM Conference on Recommender Systems (RecSys'26)

点击查看摘要

Abstract:Transformer-based sequential recommenders with causal self-attention often rely heavily on the most recent interaction at inference time, but how this behavior is structurally expressed in the representation used for prediction remains unclear. We combine prediction-time diagnostics with norm-based analysis of the full attention block. First, we show that SASRec-style models exhibit highly localized last-item reliance. We then find that, although self-attention aggregates contextual information, residual addition sharply shifts the full-block representation toward same-position contributions, which we term residual dominance. To probe this interpretation, we use inference-time residual scaling as a controlled diagnostic intervention. Changing the residual strength induces a monotonic trade-off between structural mixing and last-item reliance, while reducing residual strength recovers a subset of final-position misses for which representations at non-final positions already rank the ground-truth item correctly. Our results provide a structural account linking extreme last-item reliance to residual dominance at inference time. The code is publicly available.

6. 【2608.14011】EchoRec: Multi-Item Prediction-Empowered Generative Recommendation via Cycle-Consistent Preference Alignment

链接https://arxiv.org/abs/2608.14011

作者:Haokai Ma,Aoqi Hu,Yueao Xing,Ruobing Xie,Yonghui Yang,Teng Tu,Lei Meng,Tat-Seng Chua

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:shared token space, unifying preference modeling, recommendation autoregressively generates, Generative recommendation autoregressively, target item

备注: 10 pages, 9 figures, Under Review

点击查看摘要

Abstract:Generative recommendation autoregressively generates the semantic IDs of the target item, unifying preference modeling and index retrieval within the shared token space. Recent attempts have introduced Multi-Token Prediction (MTP) into this field, yet they primarily inherit its efficiency merit, leaving its potential as dense supervision unexplored. Unlocking this potential hinges on whether future behaviors qualify as informative supervision. Our analysis reveals that future behaviors carry a semantic echo of the current one far above that of random pairs, which nevertheless decays along horizons under intent transitions, making them informative yet order-dependent signals. Motivated by this, we propose EchoRec, which empowers MTP with cycle-consistent holistic preference alignment across multi-horizon for generative recommendation. It comprises two synergistic modules. Horizon-aware Preference Generation (HPG) sequentially chains lightweight auxiliary branches upon the base recommender, where each branch conditions on its predecessor to respect preference evolution. Verifiable Holistic-Preference Alignment (VHA) further consolidates them into the holistic preference and echoes it back through cycle-consistent projectors to suppress spurious alignment, with theoretical guarantees that exclude the rank-collapse form of spurious alignment under an invertible transport, enabling the holistic preference to be retained in the decoding representation. All auxiliary components serve as disposable scaffolding discarded at inference, introducing negligible online serving overhead. Extensive experiments on three datasets demonstrate the superiority of our EchoRec, together with its naturally acquired multi-item generation ability. Our code and datasets will be available upon acceptance.

7. 【2608.13990】Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy

链接https://arxiv.org/abs/2608.13990

作者:Liwei Deng,Jing Jiang,Zhiwei Li,Yang Wang,Guodong Long

类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)

关键词:short-video Recommender Systems, Recommender Systems, maximize user engagement, short-video Recommender, primarily optimized

备注: 9 Pages

点击查看摘要

Abstract:Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow-content videos that are effective at attracting immediate attention. However, growing evidence suggests that prolonged exposure to such content may negatively affect users' cognitive engagement and mental well-being, raising concerns about the long-term societal impact of the short-video platform. To tackle this challenge, this paper introduces a new metric, the \textbf{Content Depth Score (CDS)}, to quantify the content depth of short videos. CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes, using a seven-level scale grounded in established theories of cognitive psychology and learning. As an initial step toward this vision, we present \textbf{SCOPE-Bench}, the first benchmark for content-depth evaluation in short-video recommendation. Built upon a large-scale open-source short-video dataset, SCOPE-Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive-content perspective. Leveraging SCOPE-Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow-content videos. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives. Our code and datasets are available at this https URL.

8. 【2608.13956】How retriever redundancy and diversity impact RAG effectiveness

链接https://arxiv.org/abs/2608.13956

作者:Jonathan J Ross,Bevan Koopman,Anton van der Vegt,Guido Zuccon

类目:Information Retrieval (cs.IR)

关键词:retriever typically ranks, typically ranks documents, answer, retriever typically, typically ranks

备注

点击查看摘要

Abstract:In RAG, while the retriever typically ranks documents by their individual relevance to the query, the generator instead produces an answer based on the retrieved documents as a whole. This paper investigates how redundancy and diversity from the retrieved document set impact the generator in terms of answer correctness. Previous work has provided a mix of findings: some showing that redundancy improves generation by reinforcing relevant information, others that LLM-based paraphrasing of the same content may be beneficial. Many of these studies did not control for confounding factors like whether the documents contained the exact answer or not, and if parametric knowledge plays a role. We conduct a carefully controlled experiment investigating three key scenarios of retrieved document sets: 1) Duplicate (exact copies of the same document), 2) Paraphrased (LLM rephrased versions of one document) and 3) Diverse (documents from different genres each containing relevant information in different forms). We control for which documents contain the answer in exact match or rephrased form. Evaluation is done with FictionalQA, a synthetic, fictional question-answer dataset that ensures the LLM generator prior knowledge cannot answer the question; the answer must come from retrieved documents. We show that duplicate redundancy and LLM paraphrasing does not significantly improve answer correctness. However, providing diverse documents is highly beneficial, improving answer correctness by 17%-47%. We further show this improvement is driven by diverse forms of document genre (news, blogs, etc.) alone and not a consequence of more relevant answer being available to generator. Our findings help to direct more attention to how new retrieval methods might improve RAG by catering to the generator preference for diversity in retrieval results.

Subjects:

Information Retrieval (cs.IR)

Cite as:
arXiv:2608.13956 [cs.IR]

(or
arXiv:2608.13956v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2608.13956

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
9. 【2608.13879】Whose Posts Get Ranked: Identical-Text Exposure Gaps in Bluesky Custom Feeds

链接https://arxiv.org/abs/2608.13879

作者:Yipeng Wang,Mohit Singhal

类目:ocial and Information Networks (cs.SI); Information Retrieval (cs.IR)

关键词:independently operated recommendation, users deploy custom, operated recommendation algorithms, platform serves alongside, serves alongside thousands

备注: To appear at The 20th ACM Recommender Systems Conference (RecSys 2026), please cite accordingly

点击查看摘要

Abstract:Bluesky lets users deploy custom feeds, independently operated recommendation algorithms that the platform serves alongside thousands of others. This paper investigates how evenly these feeds treat posts with the same text. To measure this, we take repeated snapshots of the Top-50 lists that 1,366 public feeds return, and we group posts with identical text, from different authors, that were created before the same list response and closely matched in age. Exposure diverges widely inside these matched sets, which span 250 feeds: in 33% of sets, one copy appears on the list while another does not. Fixed-effects regressions show that this divergence is associated with the author's history on the specific feed. Authors new to a feed receive less exposure for the same text (-0.061 in reciprocal-rank weight), while authors whose posts the feed has returned before receive more. A new author with more followers than the competing author still loses 74\% of head-to-head comparisons. Media and post-type features show no detectable association after multiple-comparison correction. These results are early evidence that access to many independent feeds is not enough to give identical texts equal exposure.

10. 【2608.13874】Predicting Custom-Feed Returns for New Bluesky Posts: A Prospective Study

链接https://arxiv.org/abs/2608.13874

作者:Yipeng Wang,Mohit Singhal

类目:Information Retrieval (cs.IR); Computers and Society (cs.CY)

关键词:newly introduced items, cold-start recommendation addresses, introduced items, conventional approach, recommendation addresses

备注: To appear at The 20th ACM Recommender Systems Conference (RecSys 2026), please cite accordingly

点击查看摘要

Abstract:The conventional approach to cold-start recommendation addresses new users or newly introduced items. Bluesky custom feeds create a different setting: independently operated feeds filter content from a shared public stream. In this setting, newly published posts are the cold-start objects, while the feeds serve as candidates. We propose a cold-start routing task in which a newly ingested public post is the query and all rankable feeds in the monitored panel are ranked according to whether each will subsequently return it. We build a still-evolving collect-first, label-later benchmark dataset. The collected dataset covers a fixed panel of 5,000 monitored feeds and contains 17.804 million public posts, 1.865 million observable post--feed return records, and 625,083 valid feed polls. The labels record whether a post is observed among a feed's AppView Top-50 results in at least one poll during the 24 hours after publication. The current experiments use two disjoint 24-hour test folds, each paired with a 24-hour training window and separated by a 24-hour outcome-availability gap. Evaluation is conditional on the 602,186 test posts that have at least one positive observed label and satisfy the metric eligibility criteria; these posts account for 9.04% of all 6,661,658 test posts. Across the two folds, LambdaRank achieves the best equal-fold mean values among the evaluated models: 0.7361 for capped Recall@10, 0.6127 for NDCG@10, and 0.7749 for Hit@10.

11. 【2608.13842】he MPB Corpus: A Dataset of Melody, Rhythm, Harmony, and Melody-Harmony Relationships in Brazilian Popular Music

链接https://arxiv.org/abs/2608.13842

作者:Carlos de L. Almada,Hugo T. de Carvalho,Felipe D. Martins

类目:ound (cs.SD); Digital Libraries (cs.DL); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)

关键词:MPB Corpus, musical pieces encoded, presents the MPB, paper presents, pieces encoded

备注: 22 pages, 13 figures

点击查看摘要

Abstract:This paper presents the MPB Corpus, a collection of 500 musical pieces encoded across four musical parameters: melodic contour, melodic rhythm, harmony, and the relationship between melody and harmony. It constitutes the most comprehensive and detailed dataset to date for computational musicology of Brazilian music. To support the encoding process, we introduce specific analytical models designed to capture rhythmic and melodic information with precision, alongside tailored visualizations and metrics summarizing key musical parameters. Finally, we provide a brief qualitative exploratory analysis of the dataset, illustrating its potential to both formulate and systematically address musicological questions concerning the genre.

12. 【2608.13833】AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and Tool Coevolution

链接https://arxiv.org/abs/2608.13833

作者:Simiao Zuo,Chenhui Xu,Yimeng Jia,Qiang Lou,Jian Jiao,Denis Charles

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:multi-turn assistant interactions, Conversational advertising aims, aims to deliver, Conversational advertising, assistant interactions

备注

点击查看摘要

Abstract:Conversational advertising aims to deliver useful ads within multi-turn assistant interactions. Unlike conventional query-based advertising, where the user's intent is often expressed in a short standalone query, conversational ads must infer latent commercial intent from the current user query, the assistant response, and dialogue history while also deciding whether an ad would be helpful rather than intrusive. We propose AdsWorldEngine, an agentic framework for conversational advertising. AdsWorldEngine uses an Opportunity Gate to determine whether ads should be shown, an Orchestrator to generate commercial intents, call advertising tools, and construct a top-3 ad slate, and an Evaluator to score delivered ads for offline optimization. The central contribution is an iterative actor-tool training procedure: we first train the Orchestrator with supervised fine-tuning and agentic reinforcement learning, then use high- and low-reward rollouts to construct preference data to train tools. This creates a self-improving loop in which the system learns not only how to use advertising tools, but also how to improve them from rewarded behavior. To support subjective production decisions, we introduce label grounded judgment modeling, which trains judgment models from human labels collected under explicit guidelines. It enriches labels with thinking traces, filters inconsistent rationales through reflection, and further optimizes binary judgments with a cost sensitive GRPO variant that preserves asymmetric reward gaps. Offline, AdsWorldEngine improves diversity by 60% and relevance by 80% over the current production ad delivery system. In an online A/B test, it increases RPM by 22% and ads coverage by 74%.

13. 【2608.13786】Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

链接https://arxiv.org/abs/2608.13786

作者:Qingfang Liu,Qiao Jin,Joe D. Menke,Thorsten Kahnt,Zhiyong Lu

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Large language model, Large language, LLM chatbots, Large, Cochrane

备注

点击查看摘要

Abstract:Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.

计算机视觉

1. 【2608.14546】CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing

链接https://arxiv.org/abs/2608.14546

作者:Qinye Zhou,Jun Zheng,Yongchao Du,Yuan Wang,Zhengrui Chen,Zuan Gao,Taihang Hu,Chao Lin,Yefeng Shen,Xingjian Wang,Zhao Wang,Zhengtao Wu,Xiaoli Xu,Zhengze Xu,Hao Yan,Denghui Yang,Yuhang Yu,Huayu Zhang,Mingzhou Zhang,Mengting Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image editing models, editing, image editing, rapid advancement, increasingly urgent

备注: 13 pages, benchmark report

点击查看摘要

Abstract:With the rapid advancement of image editing models and their widespread application across various domains, there is an increasingly urgent need to deploy these model capabilities directly into real-world scenarios. However, existing benchmarks remain confined to simple single-image tasks, suffering from limited coverage dimensions and an inability to effectively differentiate performance among diverse models. Consequently, they fail to reliably evaluate model performance in complex multi-image editing, highly demanding reasoning instructions, and practical deployment settings. To address these limitations, we propose CPI-Bench, a Comprehensive, Practical andIntelligent benchmark for real-world image editing. CPI-Bench comprises three core subsets: CPI-General-Bench, which comprehensively covers diverse editing tasks and pioneers the inclusion of multi-image editing evaluation; CPI-Practical-Bench, which focuses on high-frequency real-user application scenarios; and CPI-Intelligent-Bench, which is dedicated to evaluating capabilities in highly demanding reasoning-based editing. Evaluation results of mainstream image editing models based on CPI-Bench demonstrate that CPI-Bench enhances performance differentiation among models. It provides a comprehensive and reliable quantification of gaps in general editing capabilities, practical deployment efficacy, and advanced reasoning-based editing, offering invaluable guidance for the future optimization of image editing models. Crucially, our ranking analysis reveals that CPI-Bench achieves the highest alignment with the Arena Image Edit Leaderboard, indicating it faithfully captures the preferences and perceptual judgments of human evaluators, serving as a robust proxy for real-world user experience.

2. 【2608.14543】MagnifiQ: Patch-aware Text Guided Progressive Upscaling for High-Resolution Image Restoration

链接https://arxiv.org/abs/2608.14543

作者:Mahesh Reddy,Yashesh Savani,Antoine Mercier,Hong Cai,Fatih Porikli,Guillaume Berger

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recovering fine-grained local, preserve global structural, global structural consistency, inputs is challenging, structural consistency

备注: Camera-ready version (ECCV workshop - LoViF'26)

点击查看摘要

Abstract:High-resolution image restoration from degraded inputs is challenging because it must preserve global structural consistency while recovering fine-grained local details, especially at 4K resolution where direct diffusion-based restoration is computationally expensive and prone to repeated or inconsistent textures. In this work, we introduce MagnifiQ, an image restoration framework that progressively upscales and restores images across resolutions, e.g., from 1024x1024 to 4096x4096. Our approach leverages a pre-trained text-to-image diffusion model such as SDXL and adapts it for more scalable high-resolution inference by replacing its original self-attention layers with convolutional operations whose computational cost grows linearly with image resolution. We further propose a progressive upscaling strategy that iteratively restores images over multiple resolution stages, refining each intermediate output rather than directly hallucinating the final 4K image, thereby improving global coherence and reducing high-resolution artifacts. To enhance local details while controlling content drift, MagnifiQ uses patch-specific text prompts that provide spatially localized semantic guidance during restoration. Extensive experiments on synthetic and real-world degraded images show that MagnifiQ outperforms prior diffusion-based restoration methods in perceptual quality and human preference, producing sharper textures and more coherent 4K results while offering practical speed--quality trade-offs through its scalable backbone and progressive design.

3. 【2608.14539】Decoding the Past: An Uncertainty-Aware Deep Learning Framework for Sex Attribution in Prehistoric Hand Stencils

链接https://arxiv.org/abs/2608.14539

作者:Karel Becerra,Boris Mederos,Dean Snow,Ramón A. Mollineda

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:created Upper Paleolithic, Upper Paleolithic hand, Upper Paleolithic, challenging problem due, Paleolithic hand stencils

备注

点击查看摘要

Abstract:Determining the biological sex of the individuals who created Upper Paleolithic hand stencils remains a challenging problem due to the absence of ground truth, population differences between contemporary and prehistoric groups, and the uncertainty introduced by image degradation. Traditional morphometric methods suffer from high structural overlap across sexes, poor cross-population generalizability, and subjective feature engineering. This study presents an uncertainty-aware deep learning framework for sex attribution in prehistoric hand stencils that explicitly models, propagates, and aggregates uncertainty throughout the analytical pipeline. The methodology combines dual image processing, dual contour extraction, structured silhouette augmentation, model architectural diversity, and ensemble-based decision aggregation. The pipeline generates twelve plausible silhouette realizations per stencil to capture boundary uncertainties, which are processed by two ensembles of ten deep neural networks each (EfficientNet-B3 and MobileViT-S) trained on 14,036 contemporary hand samples. Furthermore, a triangulated validation scheme integrates ensemble predictions with unsupervised 2D latent-space manifold mapping (UMAP + k-NN) and explainable AI spatial attributions (LayerCAM) to ensure anatomical consistency. On contemporary data, ensemble models achieve strong classification performance, with accuracies exceeding 88% in older age groups. When applied to prehistoric stencils, the framework produces both sex predictions and confidence measures of internal agreement, enabling the distinction between morphologically stable and ambiguous cases. Convergence across ensemble predictions, latent-space structure, and interpretability analyses shows that uncertainty can become a measurable component of archaeological inference, enabling robust and reproducible decoding of ancient rock art.

4. 【2608.14530】Marionette: Predicting World States, Rendering Geometry, Painting Appearance

链接https://arxiv.org/abs/2608.14530

作者:Zian Meng,Zhen Li,Chuanhao Li,Qiang Li,Kaipeng Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:typically autoregress visual, models typically autoregress, autoregress visual observations, generative sequence, world models typically

备注: Project page: [this https URL](https://alayalab.github.io/Marionette/)

点击查看摘要

Abstract:Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance. We instantiate this idea as Marionette, a world model for interactive games with articulated characters. First, a two-stage autoregressive dynamics model predicts an explicit and interpretable 276-dimensional 3D world state comprising multi-entity articulated skeletons, metric root trajectories, and rotations. Second, a zero-parameter graphics bridge converts the predicted state into pose-control videos, computing world-space geometry and occlusion in closed form. Third, a control-conditioned video-diffusion observation model synthesizes photorealistic RGB observations from the resulting structured controls. Our experiments establish two properties of Marionette. First, the predicted world state is directly controllable. Forcing a mismatched action stream changes root-aligned joint error by 31% across 48 held-out segments. Second, long-horizon behaviour is determined in the state, and can be repaired there. Left free, the two generated characters drift to 21.2 m apart (recorded sessions stay near 5 m) and a third of frames show ground penetration. Two rules imposed on the explicit state, a terrain collider and a separation cap, cut penetration by 66% and keep the pair engaged, with no change to the observation model. Routing appearance through the predicted state costs no fidelity we can detect, at an FVD of 831 against 799 for recorded pose.

5. 【2608.14435】Style or Signature? Artist-Disjoint Evaluation of Style Classification in Frozen Vision Embeddings

链接https://arxiv.org/abs/2608.14435

作者:Rory Ashton

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:CLIP are increasingly, high reported accuracy, high reported, CLIP, reported accuracy

备注: 15 pages, 3 figures. Accepted at the VISART VIII workshop, ECCV 2026

点击查看摘要

Abstract:Frozen image embeddings from models such as CLIP are increasingly used to classify paintings by art-historical style, with high reported accuracy. We ask whether this accuracy reflects an understanding of style or the recognition of individual artists. Standard evaluation uses random splits in which works by the same artist appear on both sides, so a classifier can succeed by recognising the painter rather than the movement. We re-evaluate style classification under an artist-disjoint protocol, holding out every artist in turn so that no work is ever classified using other works by its own painter. On a balanced dataset of 320 paintings across four twentieth-century movements, 5-NN style accuracy falls from 0.87 to 0.77 under this protocol, and the drop is sharply uneven. Impressionism and Cubism barely move, while Surrealism falls twenty points. The pattern holds across four image encoders, including a vision-only self-supervised model, which places the effect in visual structure rather than language. Where an encoder captures genuine shared form, individual artists are barely recognisable yet style is robust, while Surrealism shows the opposite. We argue that artist-disjoint evaluation is necessary to measure stylistic understanding in frozen embeddings.

6. 【2608.14430】Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

链接https://arxiv.org/abs/2608.14430

作者:Yixian Xu,Yuanrui Zhang,Shengjie Luo,Liwei Wang,Di He

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)

关键词:Reinforcement learning, align diffusion models, task-specific rewards, human preferences, preferences and task-specific

备注: 29 pages, 9 figures, 4 tables; work in progress

点击查看摘要

Abstract:Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic Itô integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe improves over prior diffusion-RL baselines.

7. 【2608.14428】GhostPoint: Self-Supervised Representation Learning by Hallucinating Occluded LiDAR Structure

链接https://arxiv.org/abs/2608.14428

作者:Mohamed Abdelsamad,Bin Yang,Michael Ulrich,Miao Zhang,Yakov Miron,Alexandru Paul Condurache,Abhinav Valada

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:LiDAR point clouds, autonomous driving, point clouds, core problem, problem in autonomous

备注: Accepted by ECCV2026

点击查看摘要

Abstract:3D object detection from LiDAR point clouds is a core problem in autonomous driving. Recent advances in self-supervised learning (SSL) enable scalable pretraining and transfers well to per-point tasks such as semantic and panoptic segmentation, but transfer to 3D detection remains weaker. We analyze recent SSL methods and find that most objectives are defined only on measured LiDAR returns from visible surfaces, leaving occluded and unobserved regions unconstrained. This visible-surface bias can be sufficient for point-wise prediction, but 3D detection requires robustness to missing structure. To address this gap, we propose GhostPoint, an SSL framework that hallucinates latent features in local neighborhoods around discovered instances, generated via a novel instance voxel dilation. In GhostPoint, an encoder processes observed returns, and an additional predictor infers neighborhood representations from observed context. In addition to standard encoder-level supervision, we introduce a predictor-level supervision scheme on sampled voxels from generated neighborhoods. Specifically, observed (visible/masked) voxels match teacher-encoder targets, while unobserved voxels match teacher-predictor hallucinations. This design encourages the learned representation to explicitly model structure beyond observed returns. Extensive evaluations on nuScenes and Waymo demonstrate that our method achieves state-of-the-art performance, consistently improving downstream 3D detection, especially under sparse scans and limited labels.

8. 【2608.14403】CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets

链接https://arxiv.org/abs/2608.14403

作者:Jihun Park,Kyoungmin Lee,Jongmin Gim,Hyeonseo Jo,Jaeyeul Kim,Han Zou,Zhenpeng Zhan,Yan Zhang,Sunghoon Im

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Subject-driven image personalization, visual content creation, modern visual content, Subject-driven image, content creation

备注: 20 pages, 8 figures, ACM SIGGRAPH Asia 2026

点击查看摘要

Abstract:Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is currently dominated by generalized methods that fine-tune a pretrained multimodal diffusion transformer (MMDiT) on hundreds of thousands to millions of paired \emph{(reference, composed-target)} examples, where each composed target is a synthesized image of the subject in a novel scene. Producing such targets demands a costly multi-stage curation pipeline---LLM-based prompt generation, T2I-based composed-target synthesis, reference-subject extraction, VLM-based quality filtering, and correspondence labeling---and tightly couples each method to a particular target synthesizer and curation choice. We introduce \emph{CRAFT} (Constrained Reward via Attention Fine-Tuning), a single-step ReFL framework that fine-tunes a pre-trained \emph{reference-aware} MMDiT via LoRA adapters using a compact reference-only data construction---$10$K reference images and subject masks, with no composed-target supervision. CRAFT realizes a \emph{Where to look} principle: attention-level rewards align noise- and phrase-token attention with the correct reference subject, and the resulting per-subject attention masks gate a pixel-level identity reward to keep image-space supervision consistent with the learned attention routing. Applied to FLUX.2-klein-9B, CRAFT achieves state-of-the-art performance on XVerseBench \rev{while using no composed-target supervision---only $10$K reference-only samples, whereas prior generalized methods require $150$K to over $2$M composed-target pairs}. The same recipe transfers to other reference-aware backbones, consistently improving performance. Project page: this https URL.

9. 【2608.14394】IRGNN: Efficient Invariant Radar Graph Neural Network for Radar Point Cloud Object Detection

链接https://arxiv.org/abs/2608.14394

作者:Xiao Guo,Wanke Xia,Lili Yang,Caicong Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Radar point clouds, Radar point, fundamental component, point clouds, Radar

备注: Accepted at ICONIP 2026

点击查看摘要

Abstract:Perception is a fundamental component of autonomous driving systems. While LiDAR-based methods have achieved remarkable progress in object detection, their reliability can degrade under adverse weather conditions. Radar point clouds provide a robust alternative due to their resilience to bad weather and low-illumination scenarios. However, radar point clouds are typically sparse, unordered, and less informative than LiDAR data, making it challenging to directly apply existing LiDAR-based perception methods. To address these challenges, we propose IRGNN, an Invariant Radar Graph Neural Network for radar point cloud object detection. IRGNN first reconstructs radar point clouds into graph representations using translation- and rotation-invariant feature designs, enabling robust modeling of sparse radar measurements. It then employs an improved message passing neural network (MPNN) with residual connections and a virtual node layer to enhance local feature propagation and global context modeling. Finally, task-specific heads are applied to the learned graph representations for object classification and bounding box prediction. Experimental results on the RadarScenes dataset show that IRGNN outperforms existing radar-based object detection methods and achieves competitive performance. In addition, IRGNN significantly reduces computational cost and memory usage during inference, demonstrating its effectiveness and practical potential for efficient radar-based perception in autonomous driving.

10. 【2608.14391】Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

链接https://arxiv.org/abs/2608.14391

作者:Shuo Liang,Yixing Ma,Pengfei Zhou,Xingyan Chen,Zihan Mei,Manting Li,Feihan Chen,Zhiwen Wang,Bin Xu,Haotian Zhang,Jiajun Song,Shiya Su,Run Liu,Zhenghang Ni,Yifa Yu,Jintao Hong,Bolong Feng,Yifei Liu,Zirui Zhang,Jingxuan Zhang,Songlin Zhao,Yifan Bai,Kang Tan,Yizhe Liu,Junhao Du,Yongtao Ge,Zhaopan Xv,Xinyuan Zhang,Mengru Ma,Chunhua Shen,Wei Wang,Yang You,Zheng Zhu,Kaipeng Zhang,Wangbo Zhao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:creating substantial risks, Recent video generators, Recent video, public emergencies, fabricate realistic depictions

备注: 63 pages, 20 figures, 32 tables

点击查看摘要

Abstract:Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation. Existing benchmarks, however, provide limited evidence on detector and generator behavior in such settings, including how detectability varies with generation conditions, how people perceive generated videos, and whether detectors remain reliable during social dissemination. To address this gap, we introduce RA-Bench, a benchmark for AI-generated video detection that uses Real videos as Anchors. RA-Bench contains 17,886 videos, comprising 1,830 real-video anchors across 10 social-risk categories and 16,056 generated clips from four open-source and five closed-source generators. Based on RA-Bench, we organize our evaluation along three dimensions. We first assess detector generalization across seven traditional detectors, ten zero-shot multimodal models under three review settings, and two MLLMs specifically fine-tuned on AI-generated video detection. Across these methods, none of the three detector families generalizes consistently across RA-Bench instances. We then examine how detectability varies with generation quality, conditioning information, and sampling seeds. These analyses show that generation properties affect detector families differently, while source-level detection patterns remain stable across seeds. Finally, we study human authenticity judgments and detector reliability during social dissemination. We find that videos that mislead people are also difficult for current detectors, and that social dissemination makes detection harder. Together, these findings show that current methods struggle to detect realistic AI-generated videos, highlighting the need for detectors robust to evolving video generators.

11. 【2608.14389】GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection

链接https://arxiv.org/abs/2608.14389

作者:Yingjie Ma,Zitong Yu,Wei Jia,Ajay Kumar,Linlin Shen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:hindering systematic evaluation, Existing palm presentation, multimodal video data, presentation attack detection, insufficient multimodal video

备注

点击查看摘要

Abstract:Existing palm presentation attack detection (PAD) datasets are often limited by static imagery, restricted acquisition conditions, or insufficient multimodal video data, hindering systematic evaluation across environments, modalities, and attack types. We present GBU-Palm, a large-scale multimodal video dataset and benchmark containing 21,326 videos from 105 subjects and 210 palms across six acquisition environments, including bona fide, Print, and Replay presentations, with 6,310 synchronized RGB-NIR samples. We construct leakage-controlled protocols that separate palm identity and attack lineage and benchmark four representative video architectures under environment-matched and held-out-environment settings. Results reveal substantial architecture-dependent degradation under environmental shift and show that RGB-NIR fusion does not consistently outperform RGB-only input. We further analyze model behavior through true accept (TA), true reject (TR), false accept (FA), and false reject (FR) decomposition, spectral masking, temporal-order intervention, and frozen-backbone NIR probing, revealing distinct failure patterns and evidence utilization across architectures. GBU-Palm provides a unified and challenging benchmark for developing and evaluating robust multimodal palm PAD methods under cross-environment conditions.

12. 【2608.14377】A Survey of Large Models in Sports

链接https://arxiv.org/abs/2608.14377

作者:Yichen Xu,Jianzhe Ma,Chuhan Wang,Zhonghao Cao,Liangyu Chen,Wenxuan Wang,Qin Jin

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:witnessed growing global, growing global enthusiasm, cultural exchange, social connection, recent years

备注: 36 pages, 4 figures, 6 tables. Accepted to Findings of ACL 2026

点击查看摘要

Abstract:Sports have witnessed growing global enthusiasm in recent years, serving as a vital force for physical health, cultural exchange, social connection, and economic growth. The rapid advancement of large models, particularly (multimodal) large language models (M)LLMs, has demonstrated transformative potential to reshape sports understanding, analysis, and interaction across diverse domains. This paper presents a comprehensive survey of large models in sports, including (i) an overview of tasks and applications across different participant groups; (ii) a detailed analysis of sports-related datasets and benchmarks; and (iii) a critical discussion of current challenges and future directions. Our goal is to establish a foundation for advancing research and practical development of large-model-driven sports intelligence. An open-source GitHub repository is maintained at: this https URL.

13. 【2608.14366】Weakly Supervised Polar Low Segmentation in Sentinel-1 SAR Imagery

链接https://arxiv.org/abs/2608.14366

作者:Andrea Federici,Jakob Grahn,Giacomo Boracchi,Filippo Maria Bianchi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Synthetic Aperture Radar, intense maritime cyclones, high latitudes, intense maritime, form rapidly

备注

点击查看摘要

Abstract:Polar lows are intense maritime cyclones that form rapidly at high latitudes. Deep learning can detect them in Synthetic Aperture Radar (SAR) imagery, but pixel-level segmentation remains an open challenge. No pixel-level masks are available for training, and a polar low's extent is inherently subjective, with diffuse boundaries that even experts delineate inconsistently. We propose Constrained Region Erasing with Soft Targets (CREST), a Weakly Supervised Semantic Segmentation (WSSS) framework that generates masks solely from image-level labels. Our approach builds on Adversarial Erasing (AER), which iteratively mines discriminative regions, erases them, and retrains a classifier to reveal complementary cues that become pseudo-labels for segmentation. However, standard AER also collects irrelevant background features, degrading pseudo-label quality. CREST addresses this with (i) a Constrained Ordinal Region Expansion (CORE) module that encodes the spatial-connectedness prior of polar lows, constraining region expansion from a high-confidence seed, and (ii) a Dynamic Bootstrapping (DB) loss that treats the mining order as a proxy for label reliability, attenuating supervision from noisier, later-mined regions. On Sentinel-1 SAR data, CREST follows the cyclone structure more closely than standard AER, and returns a multi-class rather than binary mask whose classes indicate the reliability assigned to each region. We further evaluate on BUS-UCLM breast ultrasound and PASCAL VOC person data, whose targets satisfy the same connectedness prior but come with the dense masks the SAR data lacks. On both datasets, CREST performs better than the equivalent AER pipeline under identical settings.

14. 【2608.14321】RIAGE: Risk-Controlled Pseudo-Label Admission for Annotation-Efficient Semi-Supervised Retinal OCT Classification

链接https://arxiv.org/abs/2608.14321

作者:Md Ashraful Hossen Akash,Shyla Afroge,Abdullah Al Mamun,Md. Kishor Morol,Tze Hui Liew

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:optical coherence tomography, diagnosing imaging modality, advanced retinal disease, retinal disease diagnosing, disease diagnosing imaging

备注: 38 pages, 12 figures, 13 tables. Code available at [this https URL](https://github.com/mdahakash/TRIAGE/)

点击查看摘要

Abstract:The advanced retinal disease diagnosing imaging modality, optical coherence tomography (OCT), encounters a lack of automation because of the high expenses for annotations performed by specialists. The use of SSL solves the problem of insufficient annotations using unlabeled B-scans; however, most of the current techniques for generating pseudo-labels are based on prediction confidence without considering the asymmetry between different types of errors. This paper proposes TRIAGE, a risk-controlled semi-supervised framework for OCT scans classification, which uses the concept of a patient-level conformal risk controller with an asymmetric cost matrix. TRIAGE unites three crucial modules: a hierarchical classifier that is capable of working with partially abnormal supervision of the disease subtypes, a patient-grouped conformal risk controller with primal-dual coverage control, and a context-aware Transformer teacher for cross-slice verification. On the dataset from Noor Eye Hospital (16,822 B-scans, 161 patients, and 554 volumes) with a test set of unseen patients, TRIAGE demonstrates 89.66% scan-level accuracy, 0.8805 macro-F1, 0.9641 macro-AUC, and an 8.34% under-grading rate when using only 20% of the labeled data. With only 5% of the labeled data, TRIAGE keeps 76.88% accuracy and a 0.1656 under-grading rate. Compared with the other six state-of-the-art semi-supervised methods, TRIAGE significantly outperforms them with ablation study demonstrating the contribution of each module in the overall framework performance (by 42.7% in terms of under-grading rate comparing to fixed threshold methods). TRIAGE demonstrates 98.00% accuracy for 3-class classification with 1% labeled data and 95.94% accuracy for 8-class classification with 10% labeled data on the OCT-C8 dataset.

15. 【2608.14317】Intelligent Detection of Mechanical, Electrical, and Plumbing (MEP) Metrics Based on 2D Floor Plans

链接https://arxiv.org/abs/2608.14317

作者:Tarandeep Singh Mandhiratta,ANK Zaman,Abdul-Rahman Mawlood-Yunis

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)

关键词:neural network-based model, neural network-based, floor plans, floor, network-based model

备注

点击查看摘要

Abstract:This research developed a neural network-based model to extract various information from 2D floor plans. We detect lighting symbols, identify the appropriate type of light, and extract the associated texts with lights. The study aims to enable efficient floor designing and determining the number and type of lights needed per floor, i.e., allow efficient design and estimate the power requirement of the floor plan. The model was developed using Mask RCNN as the base. The images were annotated and converted into a Coco data format for training the model. The model achieved bbox\_mAP and segm\_mAP values of 0.7596 and 0.7111, respectively. It also performed well at different IoU thresholds, i.e., with bbox\_mAP 50 and segm\_mAP 75 values of 0.9850 and 0.9219, respectively. The developed model will help various industries, such as architecture and construction, to improve design time and create efficient workflows by automatically detecting Mechanical, Electrical, and Plumbing (MEP) objects from floor plans, and it is the first step towards building tools that will help energy-efficient building design.

16. 【2608.14309】Spatial Message Passing in Language Space for Pathology Image Interpretation

链接https://arxiv.org/abs/2608.14309

作者:Jing-Cheng Yang,Hao-Jung Wang,Jinhao Du,Yang Hu,Ming-shan Tsai,Jens Rittscher,Bin Li

类目:Computer Vision and Pattern Recognition (cs.CV); Tissues and Organs (q-bio.TO)

关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Slide Images, gigapixel Whole Slide

备注: Accepted at MICCAI 2026 Workshop (Oral)

点击查看摘要

Abstract:Multimodal Large Language Models (MLLMs) can generate pathological descriptions from histological images, but gigapixel Whole Slide Images (WSIs) exceed their visual context limits. The standard tiling workaround makes WSIs tractable yet severs the tissue neighborhoods that define tumor-stroma interfaces and morphology. We introduce Spatial Language Message Passing (SLMP), a framework that performs spatial reasoning entirely in language space, human-readable by construction. SLMP represents a WSI region as a spatial text graph: tiles are nodes initialized with MLLM descriptions, and edges encode spatial adjacency. For each tile, an LLM refines its description by integrating language messages from adjacent tiles under a shared aggregation policy that, on the tile grid, acts as an adaptive local kernel operating on text rather than learned embeddings. This policy is an inspectable prompt that can be refined from model-observed tissue phenotypes via textual gradients, enabling automatic semantic optimization from local cellular context to broader tissue morphology without fine-tuning MLLM weights. On representative HER2 and CAMELYON16 regions, SLMP improves tile-level tumor description accuracy in settings spanning general-purpose and pathology-specialized backbones, with gains of +3.3 to +19.6 percentage points. Random-neighbor ablations confirm that these gains stem from spatial context rather than additional text alone, and inspecting the optimized policies reveals interpretable, tissue-specific decision rules. Besides, without any weight updates or fine-tuning the backbone MLLM, SLMP substantially improves general-purpose MLLMs and narrows its gap to pathology-specialized counterparts, offering a transparent and flexible mechanism for incorporating spatial reasoning into MLLM-based pathology analysis.

17. 【2608.14293】Conditional Neural Optimal Transport for Predicting Cellular Phenotypes from Molecular Structure

链接https://arxiv.org/abs/2608.14293

作者:Gauthier Avité,Maxime Sanchez-Renauld,Nicolas Bourriez,Auguste Genovesio

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:High-content microscopy enables, characterization experimentally infeasible, microscopy enables systematic, enables systematic profiling, High-content microscopy

备注: Accepted at the BioImage Computing (BIC) Workshop, ECCV 2026

点击查看摘要

Abstract:High-content microscopy enables systematic profiling of cellular responses to chemical perturbations, but the scale of the chemical space makes exhaustive phenotypic characterization experimentally infeasible. This motivates computational models that can predict image-derived phenotypes without acquiring the corresponding treated cells. We formulate molecule-induced phenotype prediction as an inductive conditional transport problem in image representation space. Given a negative-control phenotype and the structure of a molecule, we aim to predict the phenotype induced by the corresponding molecule. We first evaluate classical optimal transport baselines and show that static couplings do not yield useful predictions on large-scale phenotypic image datasets. We then introduce a molecule-conditioned Neural Optimal Transport (NOT) model with a Monge-Gap regularization training objective that learns to transport negative-control unperturbed phenotypes toward perturbed phenotypes using molecular structure as conditioning information. NOT recovers molecule-specific phenotypic effects while reducing microscopy-associated technical variation, thereby facilitating comparisons across experimental batches. On unseen active molecules, the model outperforms baseline approaches, demonstrating that chemically conditioned transport can generalize beyond the molecules observed during training. We identified the molecular encoder as the main limitation to this generalization, while transport in a compressed representation space improves performance and scalability. These results establish NOT as a promising framework for predicting cellular phenotypes from molecular structure and negative-control phenotypes, while highlighting the development of more informative molecular representations as a key direction for improving out-of-distribution performance.

18. 【2608.14287】Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels

链接https://arxiv.org/abs/2608.14287

作者:Vadym Vilhurin,Volodymyr Sydorskyi,Andrii Shevtsov

类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Passive acoustic sensing, unmanned aerial vehicles, detecting small unmanned, small unmanned aerial, acoustic sensing offers

备注

点击查看摘要

Abstract:Passive acoustic sensing offers a critical, cost-efficient, and, crucially, passive alternative for detecting small unmanned aerial vehicles. However, the practical deployment of acoustic systems is discouraged by extreme environmental noise and sensor-induced domain shift caused by heterogeneous hardware. This paper addresses these challenges by introducing a robust framework optimized for real-world battlefield conditions. We propose the integration of Per-Channel Energy Normalization (PCEN) and attention-based pooling to enhance feature extraction under low signal-to-noise ratio scenarios. We further propose a domain-aware training strategy that leverages auxiliary classes and multi-microphone data to mitigate cross-domain performance degradation. Evaluated on a unique dataset of combat-zone recordings from the Ukrainian frontlines, our approach significantly outperforms existing baselines, increasing the F1 score from 55.4% to 78.6%. This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026.

19. 【2608.14286】Seeing Red, Thinking Bad: Color Bias in Vision Language Models

链接https://arxiv.org/abs/2608.14286

作者:Kohsuke Ide,Ryousuke Yamada,Yoshihiro Fukuhara,Hirokatsu Kataoka,Yutaka Satoh

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:industrial decision-making systems, visual styling, decision-making systems, support and recommendation, Vision language models

备注: 15 pages. Accepted to ICPR 2026

点击查看摘要

Abstract:Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: this https URL

Comments:
15 pages. Accepted to ICPR 2026

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

Cite as:
arXiv:2608.14286 [cs.CV]

(or
arXiv:2608.14286v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2608.14286

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Journalreference:
In: Pattern Recognition. ICPR 2026. Lecture Notes in Computer Science. Springer, Cham, pp. 261-275

Related DOI:

https://doi.org/10.1007/978-3-032-31583-0_18

Focus to learn more

            DOI(s) linking to related resources</p>
20. 【2608.14284】PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment

链接https://arxiv.org/abs/2608.14284

作者:Yuyang Liu,Yanqing Shen,Ruike Chen,Jifan Zhao,Yuxuan Tian,Yichi Zhang,Tianfeng Long,Zixuan Yin,Yipu Wang,Ziheng Qin,Wenxing Tan,Yang Shi,Mingyu Cao,Runze Xiao,Ziqi Wang,Zhixin Yin,Shiwei Chu,Yi-Fan Zhang,Yao Mu,Yuheng Ji,Yihao Wang,Jun Yan,Zhongyuan Wang,Pengwei Wang,Xiaolong Zheng

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:binary success rates, rule-based process scores, robotic evaluation matters, matters for understanding, binary success

备注: Project page: [this https URL](https://prm-as-a-judge.github.io)

点击查看摘要

Abstract:Fine-grained robotic evaluation matters for understanding embodied models, going beyond binary success rates and rule-based process scores. We present PRM-as-a-Judge 1.5, a toolkit for robot process assessment that turns rollout videos into dense progress curves and derives multiple fine metrics. PRM-as-a-Judge 1.5 introduces three metrics, building on version 1.0, that characterize failure-side progress, post-drawdown recovery, and success-side execution quality, helping users understand embodied model capability. Based on the rollout videos from benchmarks, we perform a comprehensive assessment of the embodied models, providing some fine-grained metric results and key findings. We further introduce RoboPulse++ to evaluate the reliability of process reward models (PRM), providing evaluators with a more accurate testing platform. Moreover, we release a user-friendly assessment suite, including the benchmark, metric implementation, and visualization tools, to support reproducible manipulation process evaluation. We call on the community to rethink how robots are evaluated and establish transparent, procedural, and reproducible assessment as a foundation for the next generation of embodied intelligence.

21. 【2608.14282】MAGneT-3D: Monocular and Domain-Generalizable Temporal 3D Detection

链接https://arxiv.org/abs/2608.14282

作者:Mohamed Kotb,Johannes Meier,Christoph Reich,Oussema Dhaouadi,Luis Denninger,Daniel Cremers

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:aims to detect, monocular video, Monocular temporal, detection aims, Monocular

备注: To appear at ECCVW 2026 (DriveX workshop; Oral paper). Johannes Meier and Mohamed Kotb - both authors contributed equally. Project page: [this https URL](https://mo-sameh.github.io/MAGneT-3D-Project-Page/)

点击查看摘要

Abstract:Monocular temporal 3D detection aims to detect objects in 3D, given a monocular video. Query-based 3D detectors unify detection and cross-view association, but their learnable queries fit the spatial distribution of the training data (e.g., field-of-view). We show that this issue is especially severe when these models are applied to monocular video, hindering generalization to unseen datasets and environments. To address this limitation, we introduce MAGneT-3D, the first method for domain-generalized monocular temporal 3D object detection. Instead of relying on static learnable queries, we propose a Domain-Robust Anchor Generator (DRAG) approach that adaptively derives 3D proposals during inference. To further enable domain generalization, we propose a Temporal Refinement and Identity Merging (TRIM) strategy, reducing dependence on specific 3D proposals. To enable comprehensive domain-generalization evaluation, we establish a cross-dataset benchmark spanning nuScenes, Waymo, Lyft, and ONCE. Under zero-shot domain shifts, MAGneT-3D outperforms all baselines, improving NDS from 12.1% to 18.6% while also increasing in-domain accuracy.

22. 【2608.14281】Learning to Forecast Crop Growth from Earth Observation Data

链接https://arxiv.org/abs/2608.14281

作者:Dominik Senti,Mehmet Ozgur Turkoglu,Michele Volpi,Helge Aasen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:improving the productivity, farming systems, important for improving, operational management, management of farming

备注

点击查看摘要

Abstract:Forecasting crop growth across agricultural landscapes is important for improving the productivity, resilience, and operational management of farming systems. In this work, we investigate whether Earth observation time series and meteorological drivers can be used to predict future canopy development at country scale. We focus on winter wheat and formulate crop growth prediction as forecasting future leaf area index (LAI) trajectories beyond the last available Sentinel-2 observation. We evaluate this task on a multi-year dataset which spans the entire country of Switzerland, containing over 20 million pixel-level Sentinel-2-derived LAI time series paired with meteorological variables. Because cloud cover and revisit gaps leave LAI supervision sparse, models fit the few valid (cloud-free) LAI observations yet oscillate implausibly between them, producing trajectories no real canopy could follow. We introduce a lightweight unimodal shape regulariser which improves trajectory plausibility with negligible loss in accuracy. We compare deep learning sequence-to-sequence (Seq2Seq) models with classic machine learning baselines and show that Seq2Seq models generalise well across years, achieving $\mathrm{R}^2$ above 0.8 and consistently outperforming conventional approaches. Together, these results demonstrate that remote sensing and weather-driven sequence modelling can learn crop growth dynamics at landscape scale. S

23. 【2608.14266】Accelerating Large-scale Bundle Adjustment for LiDAR Mapping via Parallel Computing

链接https://arxiv.org/abs/2608.14266

作者:Yixi Cai,Rundong Li,Yuhan Xie,Qingwen Zhang,Patric Jensfelt,Fu Zhang

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:construct globally consistent, LiDAR bundle adjustment, globally consistent point, point cloud maps, consistent point cloud

备注: Accepted by IEEE International Conference on Automation Science and Engineering (CASE), 2026

点击查看摘要

Abstract:LiDAR bundle adjustment is widely utilized in mapping to construct globally consistent point cloud maps. In this paper, we propose the first fully parallel computing framework to accelerate LiDAR bundle adjustment for large-scale mapping, incorporating three key techniques. First, we design an adaptive, asynchronous data loading strategy to efficiently process large-scale point cloud datasets on memory-constrained GPUs. Secondly, we present a novel bottom-up voxelization method for extracting planar features, enabling fully parallelized pre-processing. Thirdly, we build upon a majorization-minimization formulation to accelerate compute-intensive tasks in the optimization via parallel computation, including the computation of residuals, Jacobian and Hessian matrices, and a parallel increment solver. To support our design, we provide both theoretical and experimental analysis of the time complexity of our approach. Extensive benchmarking on large-scale public datasets across various computational platforms validates the robustness and adaptability of our approach, achieving up to a tenfold improvement in computational efficiency while preserving mapping accuracy comparable to state-of-the-art methods. To benefit future research, the implementation code is available on GitHub.

24. 【2608.14262】On the Robustness of Temporal Vision-Language Models for Surgical Endoscopy Videos

链接https://arxiv.org/abs/2608.14262

作者:Darakshan Rashid,Raza Imam,Ufaq Khan,Muhammad Bilal,Shazad Ashraf,Dwarikanath Mahapatra,Mohammad Yaqub,Muhammad Haris Khan,Imran Razzak,Brejesh Lall,Lena Maier-Hein,Yutong Xie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:remains insufficiently characterized, clinically realistic acquisition, realistic acquisition artifacts, endoscopy remains insufficiently, offer a reusable

备注: Accepted to MICCAI 2026

点击查看摘要

Abstract:Temporal vision-language models (TVLMs) offer a reusable, prompt-based interface for surgical video understanding, yet, their robustness under clinically realistic acquisition artifacts in endoscopy remains insufficiently characterized. In practice, degradations such as defocus, haze, motion blur, noise, cautery smoke, and packet loss introduce structured distribution shifts which may compromise video-text alignment. We study the robustness of temporal VLMs under such shifts caused by corruptions in clip frames. We introduce Endo-C6, a compact corruption benchmark of six endoscopy-realistic perturbations evaluated at a fixed high severity, and apply it to public Gastrointestinal (GI) endoscopy and laparoscopic cholecystectomy videos. Under a standardized prompt protocol, we benchmark 3 recent surgical TVLM baselines and analyze robustness in both mean and worst-case settings, spanning 294 dataset-level evaluations. Finally, we present RobustEndoCLIP, obtained by few-shot parameter-efficient tuning with VeRA, outperforming existing TVLM baselines. Our findings show that off-the-shelf TVLMs can exhibit severe worst-case collapse under endoscopy-specific corruptions, whereas lightweight few-shot adaptation can substantially improve corrupted performance and robustness without changing the prompt-based interface. We expect Endo-C6 to support standardized robustness reporting and promote more reliable clinical vision-language systems.

25. 【2608.14243】Zero-Shot Skeleton-Based Action Anticipation

链接https://arxiv.org/abs/2608.14243

作者:Hongsong Wang,Pengbo Yan,Yang Zhang,Qiuxia Lai

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:recognize ongoing human, Action anticipation, aims to recognize, enabling robots, Skeleton-Based Action Anticipation

备注

点击查看摘要

Abstract:Action anticipation (AA) aims to recognize ongoing human or humanoids actions from partial observations, enabling robots to predict intentions before the actions are completed. Although skeleton-based AA offers efficiency advantages, existing approaches assume that all action classes are seen during training, which limits their deployment in real-world scenarios where novel actions inevitably arise. To address this gap, we study the new task of Zero-Shot Skeleton-Based Action Anticipation (ZS-SkAA). This task requires recognizing unseen action classes using only limited early-stage skeleton sequences, combining the challenges of partial observations, temporal dynamics, and zero-shot generalization. To establish foundational research for ZS-SkAA, we introduce:(1) A baseline model comprising a spatio-temporal feature extractor and a mutual information estimation and maximization module. This baseline model explicitly aligns partial visual features with semantic class embeddings across modalities by estimating and maximizing their mutual information, enhancing generalization to unseen classes.(2) A benchmark protocol using the NTU RGB+D dataset, which is adapted for rigorous ZS-SkAA evaluation. Experiments demonstrate the effectiveness of our model as a strong baseline for ZS-SkAA, achieving high zero-shot accuracy on NTU RGB+D. This work establishes ZS-SkAA as a vital research direction for real-world systems requiring generalization to novel actions.

26. 【2608.14235】AppleScab-LT: A Longitudinal Real-Field Apple Scab Dataset for Temporal Disease Progression Analysis

链接https://arxiv.org/abs/2608.14235

作者:Aamir Hilal,Shabir Ahmad Sofi,Neeraj Goel

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:disease, systems is constrained, constrained by limited, limited longitudinal datasets, longitudinal datasets capturing

备注

点击查看摘要

Abstract:The development of reliable plant disease monitoring systems is constrained by limited longitudinal datasets capturing disease progression under natural field conditions. Although existing plant disease datasets have advanced image-based recognition, most consist of static images acquired at a single time point, limiting analysis of temporal disease evolution and severity progression. To address this gap, this study presents AppleScab-LT, a longitudinal real-field dataset developed to monitor apple scab progression through repeated observations of individually tracked infected leaves. Guided by a research-question-driven framework, the dataset was systematically developed, validated, and characterized for reliable longitudinal disease analysis. AppleScab-LT was constructed through systematic orchard monitoring under natural environmental conditions, incorporating longitudinal leaf tracking, expert-guided disease verification, polygon-based annotation, leaf isolation, disease severity quantification, and temporal sequence construction. A comprehensive quality assurance framework, including standardized annotation protocols, expert validation, automated integrity checks, sequence-level verification, and temporal consistency analysis, was applied throughout curation. The dataset contains 21 longitudinal leaf sequences, 2,101 high-resolution images, and 264 progressive temporal samples from repeated monitoring of same infected leaves. It captures variability in severity accumulation, progression rates, monitoring duration, and inter-leaf progression. Quantitative disease descriptors based on pixel severity, color-intensity severity, and normalized relative severity provide standardized measurements for temporal disease analysis. AppleScab-LT provides a reliable resource for temporal disease intelligence, disease progression modelling, precision agriculture, and future crop health monitoring

27. 【2608.14226】RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models

链接https://arxiv.org/abs/2608.14226

作者:Ritika Allada,Pinar Yanardag

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent advances, revolutionized the field, image generation, Recent, semantics

备注: ECCV 2026

点击查看摘要

Abstract:Recent advances in text-to-image (T2I) models have revolutionized the field of image generation and editing. However, identifying semantics that a T2I model can successfully edit in an image continues to be a challenging task. Most existing approaches require users to manually specify semantics to modify a particular image, a time-consuming process that often involves extensive trial and error. In this paper, we present RankT2I, a novel, training-free, and model-agnostic framework that automates the discovery of editable semantics in diffusion and FLUX-based models. Given a visual domain, we first utilize a multimodal vision-language model to gather a broad set of candidate semantics. We then frame semantic discovery as a set selection problem and use a submodular objective to identify semantics that are relevant, editable, and diverse. Our method helps users efficiently identify a wide range of semantics for text-to-image editing models across several domains while outperforming existing methods.

28. 【2608.14207】MMUSV-Sim: A Perception-Oriented Simulation and Data-Generation Platform for Multi-USV Cooperative Perception

链接https://arxiv.org/abs/2608.14207

作者:Ziao Li,Jianxiong Ye,Biao Tang,Leping Zhang,Kun Zuo,Siyu Huang,Chenqiang Gao

类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)

关键词:unmanned surface vehicles, extend maritime target, maritime target sensing, combines complementary observations, multiple unmanned surface

备注

点击查看摘要

Abstract:Cooperative perception among multiple unmanned surface vehicles (USVs) combines complementary observations to extend maritime target sensing beyond the view range and field of a single platform. Developing such systems at scale calls for a unified workflow for configurable multi-USV scenarios, multimodal acquisition, and shared annotations. We present MMUSV-Sim, a perception-oriented maritime simulation and data-generation platform built on Unreal Engine 5 and Project AirSim. It provides island, open-sea, and port environments; configurable weather, time of day, and wave conditions; a diverse vessel asset library; and spline-based multi-vessel motion. MMUSV-Sim acquires RGB, depth, semantic, LiDAR, and radar observations across multiple USVs and captures a common world state for per-agent annotation export. Experiments verify that the configured wave settings produce the intended changes in vessel heave, roll, and pitch, and evaluate the geometric consistency between projected annotations and semantic renderings. In LiDAR-based cooperative BEV vessel detection experiments on the generated multi-USV dataset, Early Fusion achieves an AP@0.5 of 72.74, compared with 45.54 using a single USV.

29. 【2608.14178】LightTeaNet: A Weakly Supervised Lightweight CNN for Multi-Label Tea Leaf Disease Detection and Localization

链接https://arxiv.org/abs/2608.14178

作者:Naif Haider Chowdhury,Md Rahim,Syed Farhan Hasan,Murad Hasan,Prithwiraj Bhattacharjee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Southeast Asia, South and Southeast, parts of South, quantity and quality, important crop

备注: 24 pages, 11 figures, 5 tables

点击查看摘要

Abstract:Tea is known as an important crop in many parts of South and Southeast Asia, yet the production of tea is still hampered by the multiple diseases that decrease the quantity and quality. Traditional methods of inspection, which are manual, are not consistent, labor-intensive, and depend on extensive monitoring. This paper introduces a lightweight convolutional neural network (CNN) designed for weakly supervised multi-label classification and disease localization in tea leaves called LightTeaNet. LightTeaNet learns directly from image-level labels and employs Class Activation Mapping (CAM) to localize disease-affected regions automatically, unlike conventional object detection models such as YOLO, which require extensive bounding box annotations. For Parameter efficiency, the network integrates Depthwise Separable Convolutions, and for enhanced feature discrimination, it integrates Channel Attention. LightTeaNet has achieved a Precision of 0.9615, a Recall of 0.8772, and an F1-score of 0.9179, while it shows mAP@0.50=0.1810 without any manual annotations, which delivers a competitive localization performance in the experimental results. These results validate the model as an interpretable as well as a resource-efficient framework for intelligent disease monitoring in agriculture.

30. 【2608.14172】Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

链接https://arxiv.org/abs/2608.14172

作者:Nikolai Röhrich,Isabell Hans,Felix Krause,Björn Ommer

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)

关键词:high local coherence, tasks requiring high, requiring high local, standard models lack, practical utility

备注: Accepted at GCPR 2026 (Oral). 28 pages, includes supplementary material

点击查看摘要

Abstract:Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.g., for precisely controlling how aesthetically pleasing an image looks), and (2) they lack reliability for tasks requiring high local coherence (e.g., generating text or human hands). To tackle these issues, we introduce a novel notion of concept-wise mutual information and find large, concept-dependent differences between individual layers, demonstrating that the generation of specific structures is localized in distinct parts of the network. We exploit this insight by reinforcing the impact of concept-relevant layers in Concept Guidance (CoG), a precise, target-specific guidance method that works for models out-of-the-box without additional training, external models, gradients, or prompt engineering. CoG first quantifies each layer's concept-specific impact and then guides denoising using a weighted combination of predictions generated with concept-relevant layers skipped. We demonstrate performance increases across various targets and popular models like PixArt-alpha, SD3, SD3.5, and FLUX.1-dev. Code is available at this https URL

31. 【2608.14148】SCVIB: Editable State-Conditioned Visual Instance Binding forMulti-Turn Personalized Localization

链接https://arxiv.org/abs/2608.14148

作者:Xiongtai Yang,Ziyan He,Tao Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:introduce editable state-conditioned, editable state-conditioned visual, visual instance binding, visual evidence, state-conditioned visual instance

备注

点击查看摘要

Abstract:We introduce editable state-conditioned visual instance binding, a multi-turn localization setting in which several support-defined instances are introduced across turns and protocol-defined state events determine the final target. We instantiate this setting as SCVIB, comprising 1,050 manually verified support--query base pairs and 1,500 episodes spanning five visual domains, three difficulty levels, and four target-state dependency groups. Direct Seq-free inference reaches only 60.13\% Joint@0.5, indicating that resolving the final reference does not ensure effective use of the corresponding visual evidence for query-side localization. We address this gap with TT-VG (Transition-Tree Visual Grounding), which combines a Target-State Transition Tree (TSTT) with Visual Evidence Grounding Adaptation (VEGA). TSTT compiles the visible interaction into protocol-defined events, executes them over versioned target states, and resolves the final-query reference to the corresponding support evidence. Adapted on trajectory-derived same-instance pairs, VEGA performs support-conditioned grounding of the resolved instance using a Visual Evidence Package. TT-VG reaches 70.27\% Joint@0.5; under matched target resolution, VEGA exceeds the strongest comparison method by 16.20 points. Gains over direct inference are largest on Counter-Recency and Rollback, which require routing to non-latest or restored support evidence. Together, these results establish SCVIB as a controlled testbed and highlight the effective use of resolved support evidence for query-side same-instance localization as a central challenge in multi-turn personalized localization.

32. 【2608.14146】CSG-Mamba: A Convolutional Scoring Gating Vision State Space Network for Endoscopic Polyp Segmentation

链接https://arxiv.org/abs/2608.14146

作者:Yuliang Wang,Jiaqi Wu,Jiaye Song,Shuxia Ren

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:mucosal texture interference, device-dependent appearance shifts, Vision State Space, Accurate polyp segmentation, State Space Models

备注: 14 pages, 6 figures, and 5 tables. Accepted by ICONIP 2026

点击查看摘要

Abstract:Accurate polyp segmentation is critical for computer-aided colonoscopy, yet endoscopic images often contain low-contrast boundaries, mucosal texture interference, specular highlights, and device-dependent appearance shifts. Vision State Space Models (SSMs) provide efficient long-range modeling with linear complexity, but existing Vision Mamba segmentation models typically convert 2D features into 1D scanning sequences, which may weaken local geometric continuity and over-smooth irregular contours. We propose CSG-Mamba, a convolutional scoring gating Vision State Space network for endoscopic polyp segmentation. Built on a VM-UNet-style asymmetric U-shaped encoder-decoder, CSG-Mamba inserts a Convolutional Scoring Gating (CSG) module at the semantically rich bottleneck. CSG generates a local spatial score map through pointwise and large-kernel depthwise convolutions and recalibrates state-space features by multiplicative gating. Experiments with three random seeds show that CSG-Mamba achieves 0.9220 Dice and 15.87 HD95 on Kvasir-SEG, and 0.7418 Dice and 0.6570 mIoU on CVC-ColonDB, outperforming the baselines on most overlap and recall metrics while maintaining competitive boundary accuracy.

33. 【2608.14144】Self-Supervised Visual On-Policy Distillation

链接https://arxiv.org/abs/2608.14144

作者:Yijiang Li,Yijun Liang,Yunjie Tian,Bingyang Wang,Ke Zhang,Zhenfei Yin,Di Fu,Philip Torr,Nuno Vasconcelos

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:distillation relies heavily, informative teacher-student asymmetry, Visual on-policy distillation, on-policy distillation relies, regions of interest

备注

点击查看摘要

Abstract:Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S$^2$VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S$^2$VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S$^2$VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at this https URL

34. 【2608.14142】PISA: A Pseudo-Individual Source-Domain Feature Adaptation Framework for Test-Time Open-Vocabulary Object Detection

链接https://arxiv.org/abs/2608.14142

作者:Ziyan He,Xiongtai Yang,Tao Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:encountering image-domain shifts, object detection test-time, Open-vocabulary object detection, source-free OVOD-TTA methods, source-free OVOD-TTA method

备注

点击查看摘要

Abstract:Open-vocabulary object detection test-time adaptation (OVOD-TTA) aims to address the performance degradation that pre-trained base models suffer when encountering image-domain shifts. Existing source-free OVOD-TTA methods rely either on refined test-time information for re-scoring or on pseudo-labels for self-training, leading to significant accuracy degradation when initial predictions are poor. Meanwhile, most conventional source-domain estimation methods recover abstract, sparse representations suitable for the classification task, but fail to capture the dense, concrete features required for detection. To address these issues, we propose PISA, a novel source-free OVOD-TTA method that can be seamlessly integrated into open-vocabulary visual backbones. The core components of our method are the Corruption-Invariant Feature Extractor (CIFE), the Feature Alignment Module (FAM), and a multi-scale alignment framework (BAA). To capture detection-suitable features, we develop CIFE to exploit the invariance of CLIP's visual features across corrupted images, ensuring robustness against various corruptions. We further develop FAM and BAA for the pre-training and adaptation to transform the corruption-invariant features into pseudo-individual source-domain features that are close to the original source-domain features. In this way, dense and concrete pseudo-individual source-domain features are used for supervision instead of unreliable pseudo-label signals. Experiments on the corrupted VOC-C, COCO-C, and LVIS-C benchmarks across three base models demonstrate that PISA substantially improves both the localization precision and the category recognition accuracy of the original models. Notably, PISA achieves state-of-the-art performance without requiring access to source-domain data, surpassing existing methods by 3.92% in AP@50% on COCO-C.

35. 【2608.14138】SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

链接https://arxiv.org/abs/2608.14138

作者:Jinsheng Quan,Jianhua Li,Siyi Xie,Xuanke Shi,Kewang Deng,Zukai Chen,Feifei Shao,Lei Yang,Quan Wang,Yawei Luo

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:visual observations require, observations require recovering, recovering geometric structure, understanding spatial relations, require recovering geometric

备注

点击查看摘要

Abstract:Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.

36. 【2608.14136】HiCo-GS: Hierarchical Context Aggregation and Geometric Consistency for Octree Gaussian Splatting

链接https://arxiv.org/abs/2608.14136

作者:Wei Zhang,Shengkai Yu,Shiqiang Gong,Qi Zhang,Qiang Li,Qi Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Octree-based anchor Gaussian, anchor Gaussian Splatting, Gaussian Splatting, adaptively capture scene, capture scene content

备注: 21 pages, including supplementary material. To appear in the Proceedings of the 34th ACM International Conference on Multimedia (MM '26)

点击查看摘要

Abstract:Octree-based anchor Gaussian Splatting has emerged as a scalable representation for city-scale novel view synthesis, where multi-level anchors adaptively capture scene content from coarse building structures to fine architectural details. However, we identify a fundamental limitation in existing methods: cross-level feature isolation, where each level's anchor features are optimized independently with no inter-level communication, causing color drift on building facades and over-smoothing in textured regions. We present HiCo-GS, a high-fidelity reconstruction framework with two complementary modules. Cross-Level Context Aggregation (CLCA) enables bidirectional hierarchical prior injection by leveraging the octree's spatial containment structure to aggregate per-level context vectors into parent-self-child triplets, fused via a lightweight MLP with residual connection. Coarse-level structural priors flow down to inform fine-level anchors, while fine-level detail statistics feed back to prevent over-smoothing, at negligible computational overhead. Depth-Normal Geometric Consistency (DNGC) regularization enforces agreement between rendered normals and depth-derived normals through an alpha-weighted consistency loss, complemented by edge-aware smoothness losses with progressive warmup that exploit the strong planar priors ubiquitous in urban geometry to suppress floating artifacts. We further introduce the China-Pagoda dataset comprising 8 ancient Chinese pagodas with over 1,200 images each, featuring dense ornamental carvings, curved multi-layer eaves, and repetitive fine-grained textures. Extensive experiments on Mill19, UrbanScene3D, MatrixCity, and China-Pagoda demonstrate that HiCo-GS achieves state-of-the-art rendering quality and substantially cleaner geometry across real-world and synthetic urban this http URL: this https URL.

37. 【2608.14130】AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

链接https://arxiv.org/abs/2608.14130

作者:Ying Huang,Wencan Zhang,Brian Y. Lim

类目:Multimedia (cs.MM); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC)

关键词:increasingly affect people, Computer vision models, Computer vision, generated facial content, privacy protection

备注: 10 pages, 10 figures, 2 tables, ACM MM 26

点击查看摘要

Abstract:Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debugging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face similarity perception: dependence on facial featural and configural attributes, nonlinear psychophysical response scaling, and own-group biases. We introduce the FACETS dataset and propose AlignFace, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face images and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and human cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images.

38. 【2608.14112】Fixed-Budget Gaussian Volume Encoding with Structure-Aware Allocation

链接https://arxiv.org/abs/2608.14112

作者:Michael R. Martin,Joseph Insley,Victor A. Mateevitsi,Silvio Rizzi,Kwan-Liu Ma

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Graphics (cs.GR); Machine Learning (cs.LG)

关键词:Scientific simulations, simulation resources, situ reduction, produce scalar volumes, scalar volumes faster

备注: 10 pages, 6 figures, 6 tables

点击查看摘要

Abstract:Scientific simulations often produce scalar volumes faster than they can be stored, transferred, and loaded, while in situ reduction must use only a limited share of simulation resources. This work encodes scalar fields as anisotropic Gaussian primitives under a fixed budget. The complete primitive set is allocated analytically from local field structure, including position, orientation, and shape, then refined directly against the scalar field without densification, pruning, or count changes. The selected budget determines encoded storage before refinement and, together with the iteration schedule, provides a controllable refinement-time budget. In a controlled benchmark, truncation-aware field evaluation reduces encoding time by up to 51x; 1.4 million Gaussians encode a billion-voxel volume in at most four minutes on one desktop GPU, with reduced-iteration refinement completing in under one minute. Across five datasets spanning 2.1 million to 1.1 billion evaluated voxels, compression-useful configurations achieve 15.0-38.7 dB PSNR at compression ratios from 2.2x to over 40,000x. Pre-encoding structure statistics characterize fields for which one-shot allocation yields limited gains from additional capacity. Because primitives retain scalar attributes rather than baked appearance, a single compact model serves every subsequent visualization state - supporting post-hoc transfer-function, colormap, lighting, and viewpoint changes without re-encoding.

39. 【2608.14085】CoDS: Robust Collaborative Perception via Expert-driven Detection and BEV Segmentation

链接https://arxiv.org/abs/2608.14085

作者:Jinlong Wang,Yuang Jia,Junhong Lin,Nannan Li,Wei Gao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multi-agent information exchange, Collaborative perception breaks, information exchange, breaks through single-view, single-view limitations

备注: 10 pages, 6 figures

点击查看摘要

Abstract:Collaborative perception breaks through single-view limitations via multi-agent information exchange. However, multi-source noise such as pose errors and communication delays degrades fusion feature quality, constraining perception performance. Joint training of detection and BEV segmentation provides a natural remedy, where segmented road regions help constrain target distributions and detection bounding boxes help recover ambiguous segmentation boundaries. To this end, we propose a robust Collaborative perception framework with expert-driven Detection and bev Segmentation (CoDS). To address spatial inconsistency in fusion quality, we first introduce the Collaborative Reliability Map (CoRM) to explicitly quantify feature quality distribution. Based on CoRM, we design the Semantic Mixture-of-Experts (S-MoE) module to extract differentiated features for inconsistent feature demands. Finally, to further mitigate feature noise degradation, the Bidirectional Task Complementary Interaction (BTCI) refines task-aware features through bidirectional injection. Extensive experiments on OPV2V and V2V4Real datasets show that our CoDS surpasses existing baselines on both tasks and maintains stable robustness under multi-source noise. Code: this https URL and this https URL.

40. 【2608.14078】Owner3D: Ownership-Guided Style Writing for Training-Free Localized 3D Stylization

链接https://arxiv.org/abs/2608.14078

作者:Suchang Tao,Kaifeng Shi,Zhiyan Liu,Zhuoyuan Jiang,Yuqi Ouyang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:aims to modify, part while preserving, preserving the remaining, remaining surfaces, appearance

备注

点击查看摘要

Abstract:Localized 3D stylization aims to modify the appearance of a specified object part while preserving the remaining surfaces. In large reconstruction models (LRMs), this task is challenging because style is injected into intermediate appearance representations before rendering, while compact triplane features are shared across target and non-target surfaces, causing style leakage and boundary ambiguity. We propose Owner3D, a training-free framework for localized 3D stylization that integrates localized appearance control directly into the LRM reconstruction process. Specifically, Owner3D introduces ownership-guided style writing to restrict reference-style injection to target regions, producing a single localized stylized triplane without additional training while avoiding separate global style and appearance representations. To resolve appearance ambiguity near semantic boundaries, we further introduce boundary dual slots that maintain separate local feature sources for target and non-target regions. Finally, a surface-first texture readout hierarchically combines surface, 3D, and triplane ownership evidence to robustly recover appearance under incomplete visibility. Experiments on a benchmark constructed from Google Scanned Objects and PartNet demonstrate that Owner3D consistently outperforms existing 3D stylization methods in target-region style fidelity and non-target appearance preservation, reducing appearance leakage by 86.4% and 89.9% compared with StyleSplat and LAENeRF, respectively.

41. 【2608.14075】A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

链接https://arxiv.org/abs/2608.14075

作者:Jennifer D'Souza,Fahad Ahmed,Cecilia Andrea Bustamante Andrade,Lina Frolova,Poorani Gnanasambandan,Dilshad Hussain,Muhammad Uzair Khan,Nkembeng Kevin Nkengfoa,Paul Praveen J.,Fabio Priante,Sjoerd Franciscus van der Werf,Thomas Frederik Jan van Roeden

类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Digital Libraries (cs.DL)

关键词:essential experimental evidence, encode essential experimental, tables encode essential, Atomic Layer Deposition, Etching Scientific Figures

备注: 15 pages, 1 figure

点击查看摘要

Abstract:Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.

42. 【2608.14070】InstructVVT: Instruction-Driven Video Virtual Try-On without Auxiliary Spatial Priors

链接https://arxiv.org/abs/2608.14070

作者:Dingbao Shao,Song Wu,Xinyu Chen,Qian Wang,Jiahang Li,Kuai Jiang,Jiang Lin,Yuhang Liu,Ziyu Chen,Duo Li,Jiaxin Hu,Shengrong Gu,Ziheng Tang,Rongrong Liu,Yanlun Peng,Liang Li,Junlan Feng,Lujia Jin,Ting Zhang,Jian Yang,Zili Yi

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:highly constrained editing, constrained editing task, target person clothing, video spatial structure, original video spatial

备注: 23 pages, 10 figures. Dingbao Shao and Song Wu contributed equally. Zili Yi is the corresponding author

点击查看摘要

Abstract:Video virtual try-on is a highly constrained editing task requiring the precise replacement of a target person's clothing while strictly preserving the original video's spatial structure and temporal dynamics. Existing methods heavily rely on auxiliary handcrafted spatial priors (e.g., masks, poses) for editing control. However, these priors are prone to failure in unconstrained real-world videos and often compress rich visual context into incomplete structural signals. Furthermore, standard reconstruction objectives fail to fully capture try-on-specific human preferences. To address these challenges, we propose InstructVVT, an instruction-driven and reference-guided video virtual try-on framework based on a Diffusion Transformer (DiT) that operates without inference-time spatial priors. Our core insight is to recover fine-grained control directly from the input triplet (source video, reference garment, and instruction) via a dual-level reference conditioning scheme. Specifically, an MLLM infers semantic edit tokens for target disambiguation and structural preservation, while a lightweight conditioning pathway explicitly injects fine-grained visual garment details. Finally, we design a try-on-specific reward and utilize the DiffusionNFT algorithm to align the model with human preferences. Extensive experiments on ViViD-S and TripVVT-Bench demonstrate that InstructVVT outperforms state-of-the-art open-source methods in garment fidelity, structural preservation, and temporal consistency, despite requiring fewer inference-time controls.

43. 【2608.14058】Voxel-based 3D Facies Segmentation from Seismic Data: A Comparative Study

链接https://arxiv.org/abs/2608.14058

作者:Duc-Thanh Pham,Minh-Tan Pham,Anh Nguyen,Van Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:requiring robust methods, effectively identify geologically, identify geologically analogous, geologically analogous facies, Seismic facies segmentation

备注: 5 pages

点击查看摘要

Abstract:Seismic facies segmentation has emerged as a significant challenge in geophysics, requiring robust methods and systems to effectively identify geologically analogous facies with limited labeled data. Although existing studies have shown promising results in 2D facies segmentation, they often preprocess the original 3D seismic volumes into sets of 2D slices, typically the inline and crossline directions, and treat this problem as a purely 2D segmentation task. This simplification introduces discontinuities across slices and fails to preserve the spatial and structural continuity in 3D seismic data, thus limiting the model's ability to learn coherent geological patterns. In this work, we present a comparative and reproducible benchmark for voxel-based 3D seismic facies segmentation, built upon publicly available seismic volumes including the Netherlands F3 and the Parihaka datasets, with standardized data splits and evaluation metrics. By evaluating the three representative families of modern 3D segmentation architectures, we establish strong baseline results that highlight the potential and remaining challenges for future research in this domain.

44. 【2608.14051】Discovery and Spatial Characterisation of Multiple Shortcut Groups for Auditing Vision Model Bias

链接https://arxiv.org/abs/2608.14051

作者:Akshit Achara,Vishnunarayan Manickam,Thomas Day,Esther Puyol Anton,Alexander Hammers,Andrew P. King

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieve high average, high average accuracy, Deep learning models, average accuracy whilst, accuracy whilst relying

备注

点击查看摘要

Abstract:Deep learning models trained on datasets with spurious correlations can achieve high average accuracy whilst relying on shortcut features that do not generalise out of distribution. Whilst out-of-distribution testing highlights subgroup performance disparities arising from shortcut learning, it does not localise the regions within images that are associated with it. Existing research mostly uses attribution maps from interpretability methods to understand the spatial nature of spurious correlations. For example, conditional alignment methods separate task-relevant evidence from evidence tied to spurious correlations by comparing attribution maps from a task model, a sensitive attribute model, and a bias-reduced reference model. This yields shortcut-aligned and task-aligned contribution maps for each image. However, existing methods aggregate these maps across the dataset, potentially masking recurring spatial shortcut patterns that occur only in subsets of images. We address this limitation by grouping per-image shortcut and task contribution maps into recurring spatial patterns using K-means and non-negative matrix factorisation, and visualising the resulting shortcut groups through contribution maps and representative examples. Across CelebA, CheXpert, Waterbirds, Camelyon17, and ISIC2019, and across ResNet and ViT models, the discovered shortcut groups reveal both shared and distinct spatial patterns of shortcut and task contribution, with varying subgroup composition and error rates, enabling targeted inspection of image subsets with higher error rates. We perform input occlusion and internal test-time interventions to show that masking or suppressing task contribution regions substantially degrades the model classification performance and propose a combined shortcut suppression and task amplification feature intervention approach which generally reduces performance disparities.

45. 【2608.14047】Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

链接https://arxiv.org/abs/2608.14047

作者:Yi Ding,Yanzhao Yu,Xili Dai,Xianbiao Qi,Peiwen Sun,Xueqian Wang,Xiangyu Yue,Jianan Wang

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:propose Agentic Robot, Agentic Robot, propose Agentic, paper integrates, vanilla VLA models

备注: 12 pages, 4 figures, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern (CVPR) Findings

点击查看摘要

Abstract:This paper integrates end-to-end Visual-Language-Action (VLA) models with agentic tool-use to propose Agentic Robot with Tool-use (ART). ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement. Compared to vanilla VLA models with a whole continuous action solution space, ART reduces the complexity of the action solution space through tool-use, which not only improves generalizability across different tasks but also reduces data dependency. To demonstrate the advantages (high generalizability and low data dependency) of this framework, we first built a dataset of 30K tool-use trajectories and action demonstrations, which is much smaller than those used by baseline methods. We then designed a training regimen for long-trajectory tool-use reasoning in challenging environments. Experiments show that ART achieves a 20% higher success rate than mainstream baselines on simulation and real-world tasks, such as pick-and-place in the dark at novel viewpoints. Empirical results highlight the benefits of an agent-based approach: modular tool utilization enables more efficient training, lightweight deployment, and scalable integration of new tools. This design fosters robustness, adaptability, and extensibility, paving the way for the practical deployment of VLA systems in complex real-world scenarios.

46. 【2608.14046】Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking

链接https://arxiv.org/abs/2608.14046

作者:Tomislav Dobrički,Byung-Woo Hong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reverse diffusion process, latent time step, diffusion model prediction, pretrained diffusion model, reverse diffusion

备注: This paper has been accepted for publication at the European Conference on Computer Vision (ECCV), 2026

点击查看摘要

Abstract:In this work, we propose a source-agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the discrepancies of a pretrained diffusion model's prediction for each latent time step. Rather than relying on a fixed threshold, our method introduces a time-dependent statistical thresholding scheme derived from the empirical mean and standard deviation of prediction discrepancies across the latent noisy images from the target distribution. This allows the mask to adapt to the model's varying predictive confidence at different noise levels, effectively isolating domain-specific regions while preserving global structural coherence. Experimental results on the AFHQ and Celeba-HQ datasets demonstrate that our approach outperforms state-of-the-art unsupervised Image-to-Image methods in both realism (FID, KID) and faithfulness (SSIM, LPIPS). By requiring only a pretrained model of the target domain, our approach enables precise, automated localization and seamless translation across diverse source distributions without any specialized training. The project source code is available at: this https URL

47. 【2608.14043】Beyond Text Conditioning: A Systematic Study of MLLM-DiT Fusion for Video Generation

链接https://arxiv.org/abs/2608.14043

作者:Yanbo Ding,Yijia Fan,Caihua Shan,Yifan Yang,Yifei Shen,Weijie Wang,Xirui Hu,Dongsheng Li,Lili Qiu,Yuqing Yang,Yali Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Diffusion Transformers, perform high-level semantic, high-fidelity video generation, planning remains limited, dominant paradigm

备注

点击查看摘要

Abstract:Diffusion Transformers (DiTs) have become the dominant paradigm for high-fidelity video generation, yet their ability to perform high-level semantic planning remains limited. While hybrid architectures integrating MLLMs with diffusion backbones have shown strong advantages in image synthesis, such designs remain underexplored in video generation, where existing approaches often treat MLLMs primarily as frozen feature encoders rather than semantic generators. To fill this gap, we systematically study how an MLLM should be integrated with a DiT for video generation by answering three questions: what intermediate representation should bridge the MLLM and DiT, how the MLLM should generate it, and how the DiT should incorporate it during diffusion rendering. Our analysis reveals three key findings: (1) discrete semantic visual tokens produced by an EMA-based tokenizer provide a stable and expressive interface, (2) autoregressive causal modeling is effective for generating these tokens, and (3) explicit visual-token conditioning is more effective than prompt refinement or latent bridging. Based on these findings, we propose BiVidGen, a hybrid framework where an MLLM first generates semantic visual tokens and a DiT renders videos conditioned on both text and these tokens via multi-layer cross-attention. Extensive experiments show that BiVidGen improves semantic alignment and temporal coherence over a fine-tuned DiT baseline, achieving stronger performance on VBench-Long. These results demonstrate that explicit MLLM-based visual planning provides an effective intermediate interface for text-to-video generation beyond text-only conditioning.

48. 【2608.14027】E-S2Feat:Semantic-Guided Spiking Local Feature Detection and Description for Event Cameras

链接https://arxiv.org/abs/2608.14027

作者:Yang Yi,Juntao Hua,Jinpu Zhang,Liangwei Fan,Hui Shen,Dewen Hu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:attracted increasing attention, high temporal resolution, Benefiting from high, event-based local feature, dynamic range

备注

点击查看摘要

Abstract:Benefiting from high temporal resolution and dynamic range, event-based local feature methods have attracted increasing attention. However, event sparsity, noise, and limited texture still hinder robust local feature learning. Deploying such methods on resource-constrained platforms such as unmanned aerial vehicles also requires balancing accuracy and energy efficiency. To address these challenges, this paper proposes \textbf{E-S2Feat}, a spiking neural network framework for event-based local feature detection and description. The framework jointly optimizes local feature learning from the perspectives of feature representation and selection. First, a module-specific spiking activation mechanism preserves fine-grained structural cues and discriminative information under low-bit, energy-efficient inference, thereby improving overall representation fidelity. Furthermore, a semantic-guided feature modulation mechanism leverages semantic priors to refine keypoint response distributions and enhance local descriptor discriminability, thereby guiding the model to extract local features with greater geometric stability and stronger discriminative capability. Experiments on the ECD and EDS datasets show that the proposed method significantly outperforms baseline methods such as SuperEvent in pose estimation accuracy. It also achieves accuracy comparable to its artificial neural network counterpart while delivering an approximately 4.8-fold improvement in theoretical computational energy efficiency. Visual-inertial odometry experiments on the TUM-VIE dataset further verify the effectiveness and practical application potential of the proposed method in complete SLAM systems.

49. 【2608.14024】SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models

链接https://arxiv.org/abs/2608.14024

作者:Haojie Feng,Peizhi Zhang,Xinrui Zhang,Zhuoren Li,Junpeng Huang,Xiurong Wang,Dongxiao Yin,Yuxiang Zhang,Junfan Zhu,Lu Xiong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:autonomous driving jointly, driving jointly produce, produce scene interpretation, jointly produce scene, driving trajectories

备注

点击查看摘要

Abstract:Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.

50. 【2608.14022】ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

链接https://arxiv.org/abs/2608.14022

作者:Xinye Li,Lingshuai Lin,Lei Wang,Liuzhou Zhang,Jialin Cui,Qingshan Li,Guanchu Wang,Qingbin Liu,Xi Chen,Jiang Bian,Wai Lam

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:models require low-latency, require low-latency causal, world models require, Action-conditioned video world, world models

备注

点击查看摘要

Abstract:Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirectional action-conditioned video generator into efficient few-step world models through domain adaptation, teacher-forced causal training, causal consistency distillation, and on-policy distribution matching with a bidirectional teacher. The resulting budget-specialized students operate at steady-state denoising budgets of 1, 2, and 4 steps. ForgeWM further supports a dual-path deployment protocol combining latency-critical interaction with optional replay-time refinement, where the one-step student re-noises and refines its saved draft. On paired Minecraft trajectories, ForgeWM leads the evaluated systems in Imaging Quality, reference-aligned motion-profile agreement, action-sign accuracy, and mouse-control accuracy, while achieving the lowest reference LPIPS; the same four-stage recipe transfers to gamepad-controlled FPS gameplay. Replay-time refinement matches four-step reference quality while remaining roughly three times closer to the experienced trajectory than regeneration from noise. These results demonstrate ForgeWM's effectiveness for controllable few-step video generation.

51. 【2608.14016】Content Based Video Narration of Gameplay with Vision Language Models

链接https://arxiv.org/abs/2608.14016

作者:Mathew Varghese

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:professional esports broadcasts, Live game commentary, exists for professional, professional esports, esports broadcasts

备注: [this https URL](https://mathewvarghese.space/ai-powered-game-commentary-auto-narrating-gameplay-videos-with-gpt-4o/)

点击查看摘要

Abstract:Live game commentary is scarce: it exists for professional esports broadcasts and almost nowhere else. We present a content-based video narration system that produces spoken, esports-style commentary for arbitrary gameplay recordings using a general-purpose vision-language model (VLM) and a text-to-speech back end, with no game-specific instrumentation, no engine telemetry, and no task-specific training. Three mechanisms carry the system. Temporal mosaic packing arranges nine uniformly sampled frames into a single 3x3 image, letting an image-native VLM reason about motion while consuming one image payload per segment instead of nine. Context-conditioned prompting replays the K most recent narrations as assistant-role history, suppressing the repetition that dominates per-segment captioning of static scenes. Duration-conditioned generation and elastic alignment constrain narration length in the prompt, then time-scale or symmetrically pad the synthesized audio so each utterance fills its segment slot exactly, giving frame-accurate muxing without a forced aligner. The implementation supports either cloud TTS or a 6-bit quantized 4B-parameter on-device TTS model on Apple silicon, making the speech stage fully local. We report a qualitative case study on real-time strategy footage, a cost model showing the mosaic reduces per-minute image payloads by 9x, and a candid account of observed failure modes - hallucinated game state, resolution loss from mosaicking, and prosody artifacts from time-scaling. We release the system as a reproducible baseline, with an evaluation protocol for the quantitative study a full version will report.

52. 【2608.14015】MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning

链接https://arxiv.org/abs/2608.14015

作者:Yingying Fan,Penghui Du,Leyan Zhu,Runze He,Zimeng Wu,Yuxuan Zhang,Liang Chen,Jiahao Xie,Jiangtang Wang,Shuai Shao,Anchao Yang,Yutong Bai,Yan Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:long-horizon temporal reasoning, visual evidence spread, requires long-horizon temporal, spread across time, surgical videos requires

备注

点击查看摘要

Abstract:Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the question in visual evidence spread across time. Existing approaches handle this poorly: a one-shot vision-language model (VLM) compresses the whole procedure to fit its context window and loses the detail a "before" or "after" question depends on, while video agents that train the model where to look are data-hungry and transfer poorly to out-of-domain surgery. We build an agent harness that separates reasoning from perception and improves by evolving context rather than optimizing weights. A text-only orchestrator plans which evidence to gather and issues an auditable sequence of tool calls, while frozen vision-language sub-agents execute each call over the pixels, viewing, cropping, inspecting frames, and retrieving external knowledge. We further propose a gradient-free, reward-gated Heuristic Skill Distillation loop that mines the agent's own low-scoring traces and keeps a candidate skill only when it raises a validation reward, yielding reusable retrieval skills, notably directed re-look. Growing an external skill library rather than tuning weights, the loop adapts from only about 100 labeled examples, far fewer than supervised or reinforcement fine-tuning requires. To evaluate this agent, we introduce MedClawBench, a de-leaked, doctor-grounded benchmark of 1,123 questions over self-built long neurosurgery recordings and a held-out public lecture-video test split. Across both datasets and all four evaluation dimensions, our agent consistently outperforms one-shot VLMs and general video-agent frameworks, with the largest gains on the long, out-of-domain neurosurgery videos. Project page: this https URL.

53. 【2608.13980】FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation

链接https://arxiv.org/abs/2608.13980

作者:Weidong Tang,Kaiyu Li,Yikai Wang,Yanan Wu,Haotian Gan,Shihong Wang,Xiangyong Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:large language models, requires multimodal large, multimodal large language, translate implicit instructions, precise pixel-level masks

备注

点击查看摘要

Abstract:Reasoning segmentation requires multimodal large language models (MLLMs) to translate implicit instructions into precise pixel-level masks. MLLMs encode an image as visual tokens, each of which merges a group of image patches. In remote sensing images, small targets, thin structures, and adjacent instances can occupy different parts of the same visual token. Assigning a single binary mask label to such a token loses its internal spatial structure, causing nearby targets to merge and object boundaries to become coarse. To bridge this representational gap, we introduce FIRM, a Fine-grained Intra-token Representation of Masks. For each visual token, FIRM predicts a mask code that specifies an $r\times r$ binary sub-cell pattern rather than a single foreground/background label. Given a target identified by the MLLM, the complete grid of mask codes is predicted in one mask pass. Fixed lookup converts the predicted codes into a discrete sub-cell mask, while marginalizing the code distribution yields a soft structural field. To further recover fine-grained boundaries within each sub-cell, we introduce a lightweight continuous renderer that refines this field using pre-merge visual features and image details. Across five reasoning and referring segmentation benchmarks on satellite and UAV images, FIRM achieves leading results, including $70.5/80.5$ gIoU/cIoU on LaSeRS and a $3.0$-point average gain on EarthReason. These results demonstrate the value of explicitly representing intra-token mask patterns for fine-grained MLLM segmentation.

54. 【2608.13974】ProFocus: Interpreting Affective Experience in Artistic Images with Progressive Visual Focusing

链接https://arxiv.org/abs/2608.13974

作者:Zhiyan Zhang,Zicheng Yan,Jianqi Chen,Peipei Song,Shanshan Wang,Xun Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieving emotional intelligence, emotional responses triggered, emotional responses, central to achieving, emotional intelligence

备注

点击查看摘要

Abstract:Interpreting the emotional responses triggered by images is central to achieving emotional intelligence. Compared with natural images, visual art is intentionally created to elicit emotional responses from its viewers through abstract concepts and visual metaphors, making affective interpretation particularly challenging. However, most existing methods rely on general-purpose visual embeddings (e.g., CLIP), failing to capture the nuanced cues underlying artistic emotion. To address this gap, we propose \textbf{ProFocus}, a novel framework that models affective experience in artistic images via progressive visual focusing. The key idea is to model visual representation learning inspired by a hierarchical cognitive theory of human aesthetic appreciation. Technically, ProFocus contains two core components: a Hierarchical Art Critic (HAC) and a Progressive Hint Fusion (PHF) module. HAC leverages multimodal large language models to generate structured linguistic priors at three cognitive levels--atmospheric style, narrative subjects, and concrete details--thereby translating artistic perception into coherent semantic guidance. Building upon these priors, PHF departs from conventional cross-modal fusion by sequentially injecting the hierarchical hints into visual features, enabling a progressive focusing process that mirrors human perception. This design allows the model to capture subtle affective cues and produce more faithful explanations. Extensive experiments on the ArtEmis v1.0 and v2.0 datasets demonstrate that ProFocus consistently outperforms state-of-the-art methods in both emotion recognition and affective explanation. Project page: this https URL.

55. 【2608.13973】Rethinking Auxiliary Modalities in Multimodal Zero-shot Anomaly Detection: From Semantic Fusion to Conditional Modulation

链接https://arxiv.org/abs/2608.13973

作者:Peng Wu,Xin Ge,Yujia Sun,Guansong Pang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Recent foundation model-based, endowed RGB images, zero-shot anomaly detection, Recent foundation, RGB

备注

点击查看摘要

Abstract:Recent foundation model-based methods have endowed RGB images with strong zero-shot anomaly detection (ZSAD) through vision-language pretraining. However, RGB observations alone remain limited in perceiving anomalies dominated by geometric deformation, depth variation, or subtle surface changes. Auxiliary modalities can provide complementary structural information, but existing multimodal methods typically fuse them directly into a shared semantic space, which may disturb the text-aligned anomaly semantics established by RGB foundation models and often requires modality-specific architectures. To address this issue, we propose a plug-and-play auxiliary-conditioned enhancement framework for zero-shot anomaly detection. Instead of reconstructing a joint multimodal anomaly semantic space, our framework preserves the original RGB image-text anomaly matching pathway and uses auxiliary observations as conditional signals for RGB feature refinement, allowing auxiliary modalities to seamlessly enhance existing RGB-based zero-shot anomaly detectors. Specifically, a lightweight meta-learning module takes global RGB and auxiliary representations as input and generates sample-adaptive low-rank residual updates to determine how RGB features should be refined. We further construct uncertainty-aware spatial modulation from the initial RGB anomaly response and auxiliary reliability, which determines where local residual updates are strengthened or suppressed. This global-to-local conditional modulation enables selective multimodal enhancement while preserving the original RGB anomaly semantics. Extensive experiments on MVTec 3D-AD and Eyecandies demonstrate that our framework consistently improves multiple popular RGB-based zero-shot anomaly detectors, achieving state-of-the-art performance for multimodal zero-shot anomaly detection.

56. 【2608.13969】PPOM: Marginalizing Patch-Grid Phase for CLIP-Based Generalizable Vision-Language Prompt Tuning

链接https://arxiv.org/abs/2608.13969

作者:Liang Wang,Haoyang Li,Chao Wang,Guodong Long,Jing Jiang,Yan Peng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:frozen vision transformer, tuning adapts CLIP-based, adapts CLIP-based vision-language, CLIP-based vision-language models, spatial sampling imposed

备注

点击查看摘要

Abstract:Prompt tuning adapts CLIP-based vision-language models with few trainable parameters, yet its predictions remain sensitive to the spatial sampling imposed by a frozen vision transformer. In particular, non-overlapping patch tokenization makes predictions depend on the alignment (phase) between image and the patch lattice. To reduce prediction sensitivity to patch-grid alignment, we introduce Patch-Phase Orbit Marginalization (PPOM), a training-free inference operator that treats phase shift as a nuisance variable. Given a patch stride, PPOM evaluates the identity view and reflection-padded translations, pairs opposite shifts into horizontal, vertical, and diagonal antithetic families, and assigns equal mass to these families and the identity prediction to avoid view-count bias during phase integration. In summary, PPOM provides a deterministic interface between prompt adaptation and patch-grid sensitivity. Across multiple prompt-learning hosts, PPOM improves host performance without re-training.

57. 【2608.13967】SAFE: Scene-Aware Feature Modulation for Color Constancy with Learned Color Space in Pure-Color Scenes

链接https://arxiv.org/abs/2608.13967

作者:Yuan-Kang Lee,Kuan-Lin Chen,Chih-Heng Chang,Jian-Jiun Ding

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Learned Color Space, chromaticity-based cue collapses, band of hues, estimators become ambiguous, Color constancy

备注: Project page: [this https URL](https://ntuneillee.github.io/research/safe/)

点击查看摘要

Abstract:Color constancy on pure-color scenes is challenging: when most pixels share a narrow band of hues, every chromaticity-based cue collapses to a single point and standard estimators become ambiguous. We propose a compact framework that couples two innovations: (i) SAFE, a Scene-Aware FeaturE modulation network that organizes illumination cues into a structured four-token representation, which is then selectively reweighted based on scene complexity features; (ii) the Learned Color Space (LCS), a scene-dependent chromaticity normalization that directly addresses the chromaticity collapse problem for pure-color scenes. Experiment results show that SAFE consistently improves performance in pure-color scenes. Compared to the best-performing baseline in each metric, it reduces the mean angular error by 10%, the best-25% error by 20%, and the worst-25% error by 5.8%.

58. 【2608.13949】Fast Implicit Neural Light Field Representation via Geometric Decomposition and Multi-Resolution Low-Rank Features

链接https://arxiv.org/abs/2608.13949

作者:Yao Guo,Ligen Shi,Shuchen Sun,Jun Qiu,Chang Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reconstruct dense light, dense light fields, light field, sampled ray coordinates, neural representations provide

备注: 6 pages, 2 figures, submitted to IEEE CYBER 2026

点击查看摘要

Abstract:Implicit neural representations provide a compact and continuous way to reconstruct dense light fields from sampled ray coordinates. However, fast light field reconstruction remains challenging because a light field is a high-dimensional signal with strong spatial-angular redundancy and structured disparity variations. Directly fitting 4D ray coordinates with a neural network often requires considerable optimization time to recover both view appearance and cross-view consistency. To address this issue, this paper proposes a fast implicit light field representation based on geometric decomposition and multi-resolution low-rank features. The proposed method decomposes a 4D light field into a horizontal disparity plane, a spatial texture plane, and a vertical disparity plane. Each plane is represented by a low-rank structure that combines a low-resolution 2D grid with the element-wise product of two high-resolution 1D line features at multiple resolution levels. The fused features are decoded by a lightweight multilayer perceptron to predict RGB values. Experiments on public light field datasets show that the proposed method achieves competitive reconstruction quality while providing a better trade-off among model parameters, training time, and inference efficiency.

59. 【2608.13939】CMCNet: Aligning Ultrasound Image Embeddings with Textual TI-RADS Representations for Fine-Grained Thyroid Classification

链接https://arxiv.org/abs/2608.13939

作者:Bingxin Yu,Xueli Wang,Jerry Zhou,Wenyan Wang,Li Wen,Lan Huang,Xin Feng,Fengfeng Zhou,Kewei Li

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:ACR TI-RADS framework, primary imaging modality, framework standardizes diagnosis, TI-RADS framework standardizes, ACR TI-RADS

备注

点击查看摘要

Abstract:Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5). Although widely adopted in clinical practice, most deep learning approaches focus on binary malignancy classification, while multi-class prediction and explicit utilization of feature-level supervision remain underexplored, largely due to limited annotated data. In this study, we introduce the STN dataset of 600 thyroid nodules with paired transverse and longitudinal ultrasound images, bounding box annotations, and complete labels for all five TI-RADS feature categories. Following the clinical decision process, we investigate how structured feature information can guide representation learning during training while requiring only images at inference. We demonstrate that text embeddings derived from standardized feature descriptions form a stable surrogate representation for TI-RADS risk levels. Based on this observation, we propose CMCNet, which aligns image embeddings to fixed textual embeddings via a Center-Margin Contrastive Loss that simultaneously promotes intra-class compactness and inter-class separation. Experimental results show that this embedding alignment strategy is more data-efficient and robust than direct multitask learning, and consistently outperforms InfoNCE, center loss, a strong multitask baseline, and a VQA-style multimodal model, particularly in imbalanced settings. The dataset is freely available at doi: https://doi.org/10.5281/zenodo.19125693 and the source code is available at: this https URL.

60. 【2608.13938】CoANeRV: Coordinate-Aware Token-Space Neural Video Representation

链接https://arxiv.org/abs/2608.13938

作者:Jialong Guo,Ke Liu,Mengxuan Li,Jiajun Bu,Haishuai Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Neural representations, storing video-specific information, strong reconstruction fidelity, shown strong reconstruction, shown strong

备注: 27 pages, 13 figures. Code is available at [this https URL](https://github.com/jialong2023/CoANeRV)

点击查看摘要

Abstract:Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video-specific information in network weights. However, existing formulations typically require either costly per-video optimization or video-specific weight generation, making it difficult to scale to efficient amortized video representation. We propose CoANeRV, a coordinate-aware token-space framework that adapts the broader token-conditioned neural-field paradigm to amortized video representation. CoANeRV forms compact video tokens in one feed-forward pass and uses a shared coordinate-conditioned decoder to reconstruct continuous spatio-temporal queries, avoiding per-video decoder optimization or generation while retaining coordinate-level reconstruction flexibility. To make token-space reconstruction effective, CoANeRV introduces a coordinate-aware decoding architecture that aligns spatio-temporal queries with video tokens through axis-adaptive positional encoding and temperature-modulated cross-attention. Block-wise coordinate querying further reduces peak attention memory, making high-resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization. These results support the proposed video-specific combination of feed-forward token formation, spatio-temporal coordinate retrieval, and memory-bounded dense querying. The code is available at this https URL.

61. 【2608.13929】RGBX-Next: Towards Realistic Generative Rendering from G-Buffers

链接https://arxiv.org/abs/2608.13929

作者:Zheng Zeng,Marco Salvi,Lifan Wu,Jan Novák,Daqi Lin,Saeed Hadadan,Yichen Sheng,Robert Pottorff,Shiqiu Liu,Ravi Ramamoorthi,Ling-Qi Yan,Miloš Hašan

类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)

关键词:achieved impressive results, streaming generation, achieved impressive, impressive results, G-buffers

备注

点击查看摘要

Abstract:Diffusion models have achieved impressive results in image, video, and streaming generation. However, compared to traditional 3D rendering, they still lack precise control over the generated output. We believe a viable path forward is to use generative models as learned renderers conditioned on traditionally rendered G-buffers. We introduce RGBX-Next, a unified generative framework for forward and inverse rendering, which allows estimating G-buffers from images, videos, and streams, and rendering realistic images, videos, and streams from G-buffers. Our key contribution is a general recipe for finetuning diffusion transformer (DiT) models into generative forward and inverse renderers. We show that the resulting models achieve high quality in both realistic generative rendering and intrinsic decomposition. We will make all our models publicly available. We believe that the design principles presented in this paper will benefit future research on controllable generative forward and inverse rendering.

62. 【2608.13923】OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation

链接https://arxiv.org/abs/2608.13923

作者:Dinh Tuan Nguyen,Anh Dao,Phuong Nam Dang,Quan-Dung Pham,Tuyen P. Le,Truong Nguyen,Quan Nguyen

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:committed semantic label, provide compact semantic, graphs provide compact, single fused feature, compact semantic memory

备注

点击查看摘要

Abstract:Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove minority yet task-relevant hypotheses from the task-time interface. We present OpenBelief-Nav, an evidence-preserving object memory that retains observation-level phrases, reliability cues, and frame-mask provenance while maintaining separate aggregate geometric and visual representations. Semantically related phrases are consolidated into a vocabulary-independent object belief from which task-specific readouts perform fixed-vocabulary projection or free-form retrieval. On five ScanNet200 and eight Replica scenes, full-belief projection achieves mIoU scores of 0.2742 and 0.2912, compared with 0.2393 and 0.2701 for a matched early-commit readout. Across 78 HM3D-YCB navigation trials, consensus and early-commit retrieval each achieve 60/78 successes, compared with 58/78 for belief-weighted retrieval and 55/78 for DualMap. Across 20 Unitree G1 runs organized as 10 matched evaluation cases, a correction policy permitting at most two verified candidate attempts improves target-confirmation success from 6/10 to 8/10 relative to top-1-only execution. Code will be released upon acceptance at this https URL.

63. 【2608.13918】Beyond Control Points: Arcsecond Relative-Motion Estimation of Vision Measurement Platforms With Incomplete or Absent Control Fields

链接https://arxiv.org/abs/2608.13918

作者:Meng Lian,Jian Wang,Shuixin Pan,Haibo Liu,Yueqiang Zhang,Yulan Guo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Long-range vision-based deformation, vision-based deformation monitoring, Long-range vision-based, vision-based deformation, deformation monitoring

备注: 13 pages, 14 figures. Submitted to IEEE Transactions on Image Processing

点击查看摘要

Abstract:Long-range vision-based deformation monitoring is highly sensitive to motion of the camera platform. Absolute-pose differencing typically relies on dedicated control data and propagates two independent pose errors into the relative-motion estimate. We develop a control-adaptive differential framework that estimates inter-frame platform motion directly from image displacements and known 3D points. With no dedicated control point, the framework recovers platform rotation from measurement-point observations. One surveyed control point enables prior-constrained translation recovery, while two nonparallel control rays recover full 3D translation. The framework requires neither nonlinear optimization nor an initial pose estimate. Excluding control data from the rotation stage makes the rotation estimate exactly immune to contamination confined to the control field. The inherited differential formulation also cancels translational extrinsic errors exactly. We derive the rotation observability condition, a leakage bound for unmodeled translation and nonrigid point motion, and the single-point axial-prior bias law. Under 0.5-pixel image noise, attitude changes of up to 30~arcmin, and 3D point perturbations of up to 2~mm, the multi-camera estimator achieves a rotation RMSE of 2.97~arcsec and an average runtime of 0.46~ms. With one surveyed control point, its prior-constrained translation RMSE is 1.19~mm. In a bridge experiment without a stable control field, the median coordinate-wise displacement RMSE relative to total-station measurements is 0.85~mm. The estimator also maintains zero divergence under the tested 3D coordinate perturbations on public RGB-D and stereo sequences. These results establish state-of-the-art accuracy, calibration robustness, and computational efficiency among the evaluated methods.

64. 【2608.13889】Consensus-gated Multi-Agent Neural Architecture Search for Seismic Fault Segmentation

链接https://arxiv.org/abs/2608.13889

作者:Shehram Baig,Ahmad Mustafa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:medical imaging domains, labeled data resources, larger labeled data, seismic fault segmentation, labeled data budgets

备注

点击查看摘要

Abstract:Neural networks for seismic fault segmentation are often borrowed from computer vision and medical imaging domains where they train under relatively much larger labeled data resources. Optimizing their architecture under tight labeled data budgets as are common in geophysical applications is not a trivial problem. Manually designing data-optimal architectures is time-consuming while classical neural architecture search (NAS) is restricted to hand-crafted search spaces and large compute budgets. We present an agentic NAS system in which a panel of three large language models (Claude, GPT-5.1, and Gemini~2.5~Pro) debates each candidate architecture to unanimous consensus, authors the complete PyTorch implementation, cross-reviews it, and submits it to an automated validate-train-score loop with a hard 450K parameter budget, keep-or-revert lineage, and a memory of failed mechanisms. Operating on source code rather than a predefined operation menu, the search ran on a single consumer GPU and trained only eight candidates. It discovered \ours{}: a 425K-parameter encoder-decoder with a strip-pooling bottleneck, squeeze-and-excitation gating, an asymmetric one-conv decoder, and a feature-pyramid fusion neck. Trained under a protocol identical to all baselines on sections derived from the Thebe fault dataset, it attains the highest F1 (0.578) and IoU of all models tested while being the smallest, outperforming a published-capacity U-Net (31M parameters, F1 0.484), DeepLabV3-ResNet50 (39.6M, 0.516), an Attention U-Net(1.83M, 0.502). The search cost 101 LLM calls ($\sim$1.15M input / 0.39M output tokens) and roughly one GPU-day, making consensus-gated LLM panels a practical, low-cost route to domain-specific architecture discovery.

65. 【2608.13865】Attention Capture Is Not Detection: A Two-Stage Account of How Humans Miss Localized AI Image Edits

链接https://arxiv.org/abs/2608.13865

作者:Chiao-Chieh Deng

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:image edits proliferate, AI-generated image edits, resulting disinformation treat, undifferentiated property, AI-generated image

备注: 9 pages, 4 figures

点击查看摘要

Abstract:As AI-generated image edits proliferate, the platforms meant to curb the resulting disinformation treat detectability as a single, undifferentiated property: an edit either gets a warning or it does not. We show this is the wrong model. Across a controlled eye-tracking study ($N=59$, Latin-square design, four conditions crossing edit area and semantic plausibility), a mixed-effects analysis reveals that whether an edit is noticed and whether it is correctly judged as fake are dissociable stages, governed by different factors: edit area drives attention capture ($p0.001$) while semantic plausibility drives judgment accuracy and look-but-fail-to-see (LBFS) error rates ($p0.001$). This dissociation survives correction for multiple comparisons; a secondary interaction between the two factors does not. This two-stage account extends a long-standing distinction in visual attention research (between pre-attentive capture and effortful recognition) into the new domain of AI-edit detectability. We then test whether a generative eye-movement model can computationally operationalize the attention-capture stage: a Transformer trained to generate scanpaths tracks per-image attention with strong discriminative power (Pearson $r=0.77$--$0.82$ across held-out stimuli) and, on the harder task of predicting LBFS incidence, modestly outperforms a two-parameter linear baseline even without access to the plausibility label ($r=0.52$ vs. $r=0.48$). We report this comparison, our ablations, and our method's limitations (a single fixed train/validation split, not leave-one-subject-out) without inflation, consistent with responsibly communicating what a machine learning system can and cannot do to help curb AI-driven disinformation.

66. 【2608.13861】XSA-MAD: Cross-modal Semantic Alignment for Morphing Attack Detection

链接https://arxiv.org/abs/2608.13861

作者:Jie Jin,Mahiro Tokumasu,Yu Makino,Masakatsu Nishigaki,Tetsushi Ohki

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:face recognition systems, Morphing attacks pose, recognition systems, morphing attack detection, image-based morphing attack

备注: accepted to ICIP2026

点击查看摘要

Abstract:Morphing attacks pose a serious threat to face recognition systems. However, existing image-based morphing attack detection (MAD) methods often generalize poorly to unseen generation techniques because they rely solely on visual cues. We propose XSA-MAD, a CLIP-based multimodal framework that explicitly models semantic inconsistencies between bona-fide and morphed faces. Morphing concepts are decomposed into four interpretable attributes, including identity, facial geometry, texture, and consistency, and are encoded as structured and attribute-aware textual representations. The image encoder is progressively aligned with this discriminative textual space, resulting in a unified semantic representation that captures generation-invariant and concept-level discrepancies between bona-fide and morph images. Experiments on MAD22 and MorDIFF, following training on SMDD, demonstrate strong generalization across diverse morphing principles. In particular, XSA-MAD achieves an equal error rate of 2.92% on GAN-based morphs and consistently outperforms existing methods under high-fidelity generative attacks.

67. 【2608.13858】Face Re-morphing: Differential Morphing Attack Detection via Feature-Space Similarity Changes

链接https://arxiv.org/abs/2608.13858

作者:Jie Jin,Masakatsu Nishigaki,Tetsushi Ohki

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:single morphed document, morphed document image, document image, morphing attacks pose, trusted live image

备注: accepted to IJCB2026

点击查看摘要

Abstract:Face morphing attacks pose a serious threat to face recognition systems because a single morphed document image can be matched to multiple contributors. Differential morphing attack detection (D-MAD) addresses this threat by comparing a document image with a trusted live image, but existing methods often rely on static feature differences, constituent-face reconstruction, or multi-cue fusion. This paper proposes Face Re-morphing, a D-MAD method that uses the feature-space response to an additional morphing operation as a detection cue. Given a document image and a trusted live image, the proposed method generates a re-morphed image and uses the change between the document--live and live--re-morphed cosine similarities as the detection score. Experiments on FRLL-Morphs and FEI Morph show that the proposed cue is effective across different morphing conditions, re-morphing methods, and face recognition models. Comparisons with existing methods show favorable results on AMSL and indicate that the proposed method performs well under the Criminal condition on FEI Morph Version~1, particularly when using MorDIFF. These results indicate that re-morphing-induced similarity change provides a complementary cue for D-MAD.

68. 【2608.13783】Doomed to Re-Annotate, Forever: The ImageNet Story

链接https://arxiv.org/abs/2608.13783

作者:Illia Volkov,Nikita Kisel,Tetiana Mishkina,Klara Janouskova,Jiri Matas

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:commonly reported metric, visual recognition, metric in visual, commonly reported, reported metric

备注: 25 pages, 16 figures, 8 tables. Project page: [this https URL](https://vrg.fel.cvut.cz/reimagenet)

点击查看摘要

Abstract:Top-1 accuracy on ImageNet-1k remains the most commonly reported metric in visual recognition. Quality issues with the dataset have been repeatedly reported, yet the original 2012 noisy labels are still predominantly used. The paper presents a comprehensive effort, which goes well beyond prior correction attempts, towards obtaining accurate and complete ImageNet-1k validation set annotations. The result, ReImageNet, includes multilabel correction, object localization, revised class definitions, and semantic attributes (text-recognition, rendition, reflection, crowd, dominant). The reannotation reveals that approximately 12% of the original ImageNet-1k labels are incorrect, 33.3% of images are multilabel and 3.8% contain no object from an ImageNet-1k class. With the new labels, top-1 accuracy increases by up to 1.2% for supervised models and by 5-6% for MLLMs. We argue that annotation at ImageNet scale cannot realistically be completed in one pass, as errors and definitional issues are discovered only through annotating, and we build our pipeline around repeated refinement and error checking. We observed that human and LLM collaboration with appropriate tooling represents the current quality ceiling for annotation at this scale. ImageNet-1k issues propagate into its derivative test sets, indicating that the problem is structural rather than specific to any single benchmark. All annotations, class definitions, guidelines, and analysis code have been publicly released. Project page: this https URL Annotations: this https URL Code: this https URL

69. 【2608.13769】Kolmogorov-Arnold Networks for Spatially Independent Multispectral Land Classification

链接https://arxiv.org/abs/2608.13769

作者:Katherine L. Bauer,Teemu Harkonen,Simo Sarkka,Arturo Sanchez-Azofeifa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:environmental monitoring, Kolmogorov-Arnold network, land management, random forest, Land classification

备注: 10 pages, 10 figures

点击查看摘要

Abstract:Land classification from satellite imagery is important for land management, environmental monitoring, and urban planning. Machine learning methods such as random forests and multilayer perceptrons have shown strong performance on multispectral data, while the Kolmogorov-Arnold network has emerged as an alternative architecture with compact model structures. This study evaluates the Kolmogorov-Arnold network for land classification using Landsat 8 imagery and compares it with random forest and multilayer perceptron models. The models were trained and tested on data from Edmonton, Alberta and evaluated on an independent dataset from Calgary, Alberta across five land classes: agriculture, urban, water, forest, and bare ground. For the Calgary dataset, the Kolmogorov-Arnold network matched the accuracy of the random forest and outperformed the multilayer perceptron, while requiring substantially fewer trainable parameters and providing greater interpretability.

70. 【2608.13766】ChartProbe: A Diagnostic Study on Visual Reasoning through Perception, Grounding, and Simple Reasoning

链接https://arxiv.org/abs/2608.13766

作者:Mahsa Khoshnoodi,Sarah Adel Bargal

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision-language models, remain unreliable, deficit and addressed, reasoning, Vision-language

备注

点击查看摘要

Abstract:Vision-language models (VLMs) remain unreliable on chart questions that require reasoning over visual quantities, and this weakness is usually attributed to a reasoning deficit and addressed with more reasoning supervision. We ask whether the difficulty lies in reasoning itself, or in the simpler skills that reasoning operates on: reading the plotted elements (\emph{perception}), locating them and binding them to their labels (\emph{grounding}), and performing single-step computations such as ranking, totals, and differences (\emph{simple reasoning}). We introduce \textbf{ChartProbe}, a diagnostic framework whose probes are generated directly from the code that renders each chart, so every gold answer is exact by construction, needs no human annotation, and attributes each failure to a single skill. ChartProbe enables an intervention prior work does not attempt: instead of synthesizing complex-reasoning data, we withhold complex questions and reasoning traces entirely, fine-tune on one simple skill at a time, and measure transfer to held-out complex-reasoning questions. Across three open-weight VLMs, supervising the simpler skills alone produces large gains on complex-reasoning questions the model never trained on: where these skills are weak and the model can be taught to read the image, training them recovers much of complex reasoning at no reasoning-data cost. The gains hold across three out-of-distribution settings: an unseen chart type (pie charts), a human-written benchmark disjoint from our images and templates (ChartQA), and a non-chart visual domain (CLEVR). Complex visual reasoning can therefore improve without complex-reasoning supervision.

71. 【2608.13760】Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

链接https://arxiv.org/abs/2608.13760

作者:Jean de Dieu Nyandwi,Leena Mathur,Yonatan Bisk,Robert Hawkins,Graham Neubig

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:reasoning-oriented training, correct answers, behaviors, Behavioral Lift, reasoning-oriented training amplify

备注: Published in COLM 2026

点击查看摘要

Abstract:Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

72. 【2608.13751】CAST: Closed-form Analytic Semantic Transfer for Zero-Shot Classifier Extension

链接https://arxiv.org/abs/2608.13751

作者:William Heyden,Habib Ullah,Muhammad Salman Siddiqui,Fadi Al Machot

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Large pre-trained models, machine learning systems, modern machine learning, Large pre-trained, foundational components

备注

点击查看摘要

Abstract:Large pre-trained models have become foundational components of modern machine learning systems. Yet adapting these models to novel categories typically requires examples from the target distribution. In many domains, however, such data are unavailable. Zero-shot learning (ZSL) permits recognition under these limitations through relying on auxiliary semantic information such as textual descriptions. We introduce CAST (Closed-form Analytic Semantic Transfer), a training-free, image-free framework for extending a pre-trained classifier to previously unseen classes through weight injection. We provide a theoretical foundation for CAST and derive a finite-sample error decomposition that identifies the \emph{semantic extrapolation residual} $\rho_u$. The residual is a computable, model-agnostic measure and provides a principled criterion for dataset curation and benchmark design. Experiments on standard zero-shot learning benchmarks demonstrate that CAST matches or exceeds existing image-free approaches and approaches the performance of few-shot adaptation methods, while requiring neither iterative optimization nor examples from the target distribution.

73. 【2608.13729】Limitations of Synthetic Data Generation in Specialized Data-Scarce Domains

链接https://arxiv.org/abs/2608.13729

作者:Edward Zhang,Marcel Hussing,Tanay Tandon,Shenbagaraj Kannapiran,Jason Hughes,Youkang Wang,Joshua Caswell,Agelos Kratimenos,Yi Fan Li,Milan Manoj,Ethan Sanchez,Sumukh Shrote,Camillo Jose Taylor,Daniel A. Hashimoto,Eric Eaton

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Advances in diffusion-based, diffusion-based generative models, alleviate data scarcity, models have motivated, scarcity in vision

备注

点击查看摘要

Abstract:Advances in diffusion-based generative models have motivated the use of synthetic image generation to alleviate data scarcity in vision tasks. While this strategy has shown promise in natural image benchmarks such as ImageNet, its effectiveness in sparse, high-variance real-world domains remains unclear. In this work, we focus on domains where images differ substantially from common image datasets and additional data are expensive to obtain. Against non-generative data augmentation baselines, we evaluate the downstream classifier performance improvements yielded by two schools of generative sparse data extension: distribution modeling and sample perturbation. Across five trauma classification tasks using subject-wise train--validation splits, no generative approach consistently outperforms a strong non-generative baseline. Feature-space analysis reveals recurring failure modes: memorization or collapse, distributional drift, and generation of visually plausible but simplified canonical instances that are easier to classify than real data.

74. 【2608.13702】SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers

链接https://arxiv.org/abs/2608.13702

作者:Kiran Nair,Rodrigue Rizk,KC Santosh

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)

关键词:Spiking neural networks, deep neural networks, conventional deep neural, sparse event-driven computation, exploiting sparse event-driven

备注: In-Review at a Conference

点击查看摘要

Abstract:Spiking neural networks (SNNs) offer an energy-efficient alternative to conventional deep neural networks by exploiting sparse event-driven computation, but their training remains challenging because the non-differentiable spike function requires surrogate gradients whose fixed shape may be suboptimal across layers and training stages. In this work, we introduce SAGE, an uncertainty-modulated surrogate-gradient mechanism for Transformer-based SNNs. SAGE estimates block-level uncertainty from normalized self-attention entropy and uses this signal to adapt the surrogate-gradient slope during training while leaving the inference model unchanged. By modulating only the training-time surrogate parameter, the proposed method preserves the original architecture and deployment cost while improving optimization flexibility. Experiments on CIFAR-10/100 demonstrate that SAGE achieves improved accuracy over fixed-surrogate baselines, with results up to 1-2\% consistent gains across multiple simulation time steps. These results highlight the potential of attention-derived uncertainty as a lightweight training signal for adaptive surrogate-gradient learning in transformer-based SNNs.

75. 【2608.13690】MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation

链接https://arxiv.org/abs/2608.13690

作者:Rafi Ibn Sultan,Hui Zhu,Chengyin Li,Dongxiao Zhu

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Medical image segmentation, Medical image, vision-only problem, knowledge of anatomy, surrounding context

备注: Accepted By BMVC-2026

点击查看摘要

Abstract:Medical image segmentation is still largely treated as a vision-only problem, although clinical interpretation often relies on textual knowledge of anatomy, location, appearance, and surrounding context. Existing text-guided segmentation methods within the Vision-Language Model (VLM) paradigm often use language only as a late conditioning signal, limiting its influence on visual representation learning. We introduce MedPlex (Medical Plexus of Vision and Language), an end-to-end VLM framework that makes text guidance a continuous, clinically grounded component of segmentation learning. Through Bi-Fusion (Bidirectional Fusion), visual and textual representations evolve jointly across the encoding hierarchy. MedPlex further introduces class-level and region-level concept alignment to organize the shared representation at complementary granularities. Class-level alignment anchors each anatomical target to an aggregated clinical concept profile, while region-level alignment preserves individual concepts, such as shape, location, appearance, and texture, through class-specific visual evidence. In this way, language provides structured supervision throughout the encoder rather than serving only as a late-stage cue. MedPlex achieves state-of-the-art performance across CT and MR benchmarks for multi-organ, cardiac substructure, and tumor segmentation, including settings with real free-text clinical supervision. Code: this https URL.

76. 【2608.13671】PROVE: Training-Free Prompt Recovery using Verifiable Evidence

链接https://arxiv.org/abs/2608.13671

作者:Rupayan Mallick,Mahsa Khoshnoodi,Sarah Adel Bargal

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:generate highly realistic, highly realistic images, models can generate, generated outputs, raising new concerns

备注

点击查看摘要

Abstract:Modern text-to-image models can generate highly realistic images from natural-language prompts, while recent advances in prompt inversion have made it increasingly feasible to recover those prompts from generated outputs, raising new concerns for copyright protection and content ownership. As prompt marketplaces emerge, recovered prompts can enable both the unauthorized reproduction and redistribution of copyrighted creative works, and the exposure of the prompts that encode an artist's creative recipe in AI-generated content. Existing prompt inversion methods rely on gradient-based optimization, autoregressive captioning, or reinforcement learning. However, optimization-based methods often produce unreadable prompts, captioning methods hallucinate unverified details, and RL-based approaches frequently overfit to specific generators while introducing evaluation circularity. We introduce PROVE (Prompt Recovery with Verified Evidence), a training-free, black-box prompt inversion attack that reconstructs prompts by composing verifiable scene descriptions rather than optimizing token sequences, targeting both original copyrighted works and AI-generated content. The resulting prompts are fully auditable, with every recovered claim grounded in explicit image evidence, and are formalized through a precision-constrained recall maximization objective. Across MS-COCO, Flickr30K, and Lexica, using state-of-the-art text-to-image generators, PROVE consistently outperforms optimization, captioning, and RL-based baselines on image similarity (DINO, LPIPS) and text-image alignment (CLIP), without any training, generator access, or fine-tuning, demonstrating a stronger and more practical prompt inversion attack.

77. 【2608.13669】Multiphase-Diff: Diffusion-Based Generative Modeling for High-Contrast Multiphase Physical Systems with Sharp Interfaces

链接https://arxiv.org/abs/2608.13669

作者:Yining Huang,Zhenyu Liang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:sharp-interface multiphase fields, multiphase fields faces, Physics-constrained diffusion, coupled difficulties, fields faces

备注

点击查看摘要

Abstract:Physics-constrained diffusion for high-contrast, sharp-interface multiphase fields faces three coupled difficulties. At coefficient jumps, expanded pointwise strong-form PDE residuals contain singular gradient terms that can penalize physical interfaces. Under extreme contrast, low-magnitude phases may fall below the diffusion noise floor and be erased, misscaled, or generated with negative coefficients, while a global likelihood scale allows high-magnitude phases to dominate supervision. We therefore propose Multiphase-Diff, which makes three corresponding contributions: (i) a conservative flux residual that avoids differentiating discontinuous coefficients and enforces discrete conservation; (ii) an analytic bijective representation that maps low-amplitude signals to order-one latent scales and guarantees coefficient positivity through exponential decoding; and (iii) a Jacobi-preconditioned likelihood that normalizes local residual scales for balanced supervision. Experiments on three complementary multiphase benchmarks demonstrate the superiority of Multiphase-Diff over seven baselines in both physical and distributional fidelity and its robustness across phase contrasts and compositions, establishing its effectiveness for scientific sample generation in this challenging regime.

78. 【2608.13660】What to Preserve, Where to Adapt: A Depth-Wise Analysis of Forgetting in Continual Gynecological Image Segmentation

链接https://arxiv.org/abs/2608.13660

作者:Amal Saqib,Tausifa Jan Saleem,Numan Saeed,Mohammad Yaqub

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:Medical image segmentation, Medical image, typically trained, Medical, continual learning

备注

点击查看摘要

Abstract:Medical image segmentation models are typically trained under the assumption that all data are available simultaneously. However, in clinical practice, datasets often arrive sequentially, requiring models to adapt continuously to evolving data distributions. We study this problem in gynecological image segmentation, where substantial heterogeneity across imaging modalities, anatomical structures, and annotation protocols creates a particularly challenging continual learning setting. Under these large distribution shifts, existing continual learning methods struggle to preserve previously learned knowledge, leading to catastrophic forgetting. To better understand forgetting in this setting, we investigate how different encoder--decoder regions influence segmentation performance and forgetting during continual gynecological segmentation. Through block-wise ablation analysis, we observe that ablating early encoder and late decoder regions results in the largest performance degradation, indicating that segmentation performance depends unevenly across the network hierarchy. Using controlled adaptation experiments, we further show that forgetting remains limited when updates are restricted to bottleneck-adjacent regions, but increases sharply once shallower encoders and decoders become trainable, even when only a small subset of parameters is updated. These findings suggest that forgetting in the encoder-decoder architecture is strongly influenced by where updates occur across network depth during continual learning. Full code and analysis pipelines will be made publicly available upon acceptance.

79. 【2608.13602】Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Visual Avatar Generation

链接https://arxiv.org/abs/2608.13602

作者:Lunjie Zhu,Xingtong Ge,Fangyu Lin,Yi Zhang,Zhening Liu,Mengfei Li,Yumeng Zhang,Guanglu Song,Yu Liu,Jun Zhang

类目:Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD)

关键词:generative models serve, interactive digital-human generation, Joint audio-video generative, audio-video generative models, serve as foundation

备注

点击查看摘要

Abstract:Joint audio-video generative models serve as foundation for immersive and interactive digital-human generation. Nevertheless, most existing models rely on bidirectional attention and multi-step denoising and can generate only short clips, making them unsuitable for real-time interaction over extended durations. We present Omni-LiveAvatar, the first framework for minute-level, real-time streaming joint audio-video avatar generation. Specifically, we propose (1) a progressive autoregressive distillation pipeline that transfers a large bidirectional joint audio-video diffusion model into a few-step autoregressive generator without auxiliary stabilization mechanisms; (2) a synchronized audio-video long-short-term memory that preserves global consistency under a bounded memory budget; and (3) a hierarchical rolling prompt planning strategy that enables coherent semantic evolution and seamless prompt transitions. Extensive experiments show that Omni-LiveAvatar generates high-quality, synchronized minute-level avatars in real time. In terms of speed, it achieves a 33$\times$ generation speedup over its teacher, LTX-2, on a single NVIDIA H200 GPU; in terms of generation quality, it outperforms accelerated baselines across visual quality, audio quality, cross-modal synchronization, and human fidelity. Our code is available at this https URL.

80. 【2608.13584】UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction

链接https://arxiv.org/abs/2608.13584

作者:Mikhail Kiselev,Aleksandr Marukhin,Ivan Snegirev,Elizaveta Semenyakina,Miguel Altamirano Cabrera,Dzmitry Tsetserukou

类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:real-time marker-based tracking, lightweight multilingual library, mobile augmented reality, real-time marker-based, augmented reality

备注

点击查看摘要

Abstract:UltraArUco - a lightweight multilingual library and framework for low-latency, real-time marker-based tracking in mobile augmented reality. Unlike standard OpenCV-based implementations, UltraArUco introduces an optimized multilingual wrapper that reduces per-frame latency by five times, while maintaining high accuracy. Distributed Wi-Fi architecture provides portability, connects a mobile device (camera input) with a PC-based visual application, enabling responsive interactions. The framework is validated through an interactive piano simulation, where static ArUco markers on keys enable occlusion-based note triggering, and hand-mounted markers provide spatial gesture recognition. UltraArUco's system requirements make it perfect for resource-constrained mobile AR applications, demonstrating a viable AR music application without specialized equipment.

81. 【2608.14422】UMPIRE-Net: Unrolled Magnitude-Phase Regularization Network for Accelerated MRI

链接https://arxiv.org/abs/2608.14422

作者:Mahdi Saberi,Toygan Kiliç,Mehmet Akçakaya

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP); Medical Physics (physics.med-ph)

关键词:ill-posed inverse problem, inverse problem, ill-posed inverse, MRI, PD-DL

备注: IEEE International Workshop on Machine Learning for Signal Processing (MLSP)

点击查看摘要

Abstract:MRI reconstruction from undersampled k-space measurements is an ill-posed inverse problem. Physics-driven deep learning (PD-DL) methods have shown strong performance for this task by combining the MRI forward model with learned image regularization within algorithm-unrolling frameworks. However, most existing PD-DL methods reconstruct complex-valued images directly, thereby implicitly coupling magnitude and phase within a single learned representation. This coupled regularization may be suboptimal in reconstruction settings where accurate phase modeling plays an important role, such as partial Fourier (PF) imaging, where recovery of the omitted asymmetric k-space measurements depends on the underlying image phase. In such scenarios, explicit modeling of magnitude and phase as separate components may reduce the reliance on externally estimated or predefined phase information. To this end, we propose UMPIRE-Net (Unrolled Magnitude-Phase In REgularization Network), a PD-DL method that introduces separate learned regularizers for magnitude and phase components, together with a novel data-fidelity formulation that enforces measurements consistency. We evaluate UMPIRE-Net for accelerated MRI with PF across different datasets and acceleration factors. Experimental results demonstrate that our proposed method improves reconstruction quality compared with a conventional complex-valued PD-DL baseline, yielding sharper images and reduced artifacts. Code available at: this https URL

82. 【2608.13989】A Subjective Study on a New Sharpness Informed Class of Metrics

链接https://arxiv.org/abs/2608.13989

作者:Uditangshu Aurangabadkar,Vibhoothi Vibhoothi,Darren Ramsook,Anil Kokaram

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Deep Neural Network, Neural Network, Deep Neural, deblurring architectures improve, Perceptual loss functions

备注: Accepted at IEEE MMSP 2026, 6 pages

点击查看摘要

Abstract:Perceptual loss functions in Deep Neural Network (DNN) deblurring architectures improve the overall quality of restored images. However, few focus on explicitly targeting sharpness in the restorations. We conduct a subjective study of models trained with and without losses which explicitly target sharpness using a four-protocol approach, exploring preferred sharpness levels and effects on image quality. We introduce a novel dataset of images with uniform sharpness increments along with Difference Mean Opinion Scores (DMOS). Additionally, we propose a novel class of Sharpness Informed (SI) Image Quality Assessment (IQA) metrics which properly penalize over-sharpening. Our new SI-PSNR metric outperforms all other PSNR variants in terms of correlation statistics on IQA benchmarking datasets. We show that, on average, images restored using a sharpness-aware composite loss are preferred in 67% of binarized comparisons, as opposed to losses that do not explicitly target sharpness.

83. 【2608.13897】Practical Lossless Volumetric Medical Image Compression via Tri-plane Context Tree Learning

链接https://arxiv.org/abs/2608.13897

作者:Yuanchao Bai,Yifan Zhao,Kai Wang,Yuanbo Du,Jie Cheng,Teng Fang,Xianming Liu,Wen Gao

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:fidelity is essential, paramount importance, importance for clinical, clinical and research, research applications

备注

点击查看摘要

Abstract:Lossless compression of volumetric medical images is of paramount importance for clinical and research applications where data fidelity is essential. Traditional compression methods are often limited in efficiency due to rigid, handcrafted models. Conversely, deep neural network (DNN)-based compression methods, while effective, demand substantial computational resources, hindering deployment in resource-constrained settings. To address these challenges, we propose a novel tri-plane context tree (TCT)-based method for lossless volumetric medical image compression that delivers high performance without relying on DNNs or external training data. To exploit intra-slice and inter-slice redundancies, we introduce a compact tri-plane context representation that decomposes complex 3D context modeling into efficient 2D modeling on three orthogonal planes. By integrating this representation with a context tree framework, we develop an input-specific TCT model employing an adaptive binary tree structure. At each tree node, the model dynamically selects from a suite of tri-plane based predictors and contextual feature extractors, enabling data-adaptive context modeling tailored to local structural characteristics. Instead of offline training, we sample a subset of the input volume to learn the TCT model by optimizing the minimum description length (MDL) through iterative construction and pruning. With the learned TCT model, each pixel retrieves its corresponding context, computes the prediction residual using the predictor dictated by the context, and performs entropy encoding based on the associated histograms. Experimental results demonstrate that the proposed method achieves compression performance on par with recent DNN-based methods on multiple datasets, while maintaining low computational cost and fast coding speeds, making it highly applicable in practice.

84. 【2608.13856】From crown candidates to neighborhood screening: integrating optical GeoAI and spatial modeling for urban-canopy assessment in Davis, California

链接https://arxiv.org/abs/2608.13856

作者:Mohammadreza Narimani,Shreyan Mitra,Parastoo Farajpoor

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Applications (stat.AP)

关键词:Timely urban-canopy information, linking remote sensing, National Agriculture Imagery, Timely urban-canopy, Agriculture Imagery Program

备注: 17 pages, 10 figures + 5 supplementary figures, 7 tables. Preprint submitted to Taylor Francis. Code: [this https URL](https://github.com/MohammadrezaNarimaniUCDavis/Davis_Urban_Canopy_GeoAI) Data: [this https URL](https://doi.org/10.5281/zenodo.21925527)

点击查看摘要

Abstract:Timely urban-canopy information is essential for linking remote sensing with heat, mobility, and neighborhood planning. We developed an optical GeoAI workflow for Davis, California, using 2022 National Agriculture Imagery Program imagery (0.6 m RGB+NIR). DeepForest generated crown candidates; an NDVI threshold, non-maximum suppression, and box-prompted Segment Anything Model (ViT-B) produced a crown-anchored canopy surface. Analyses used the 25.92 km2 Census TIGER municipal boundary and a 100 m grid. The workflow retained 11,741 candidate crowns and mapped 2.43 km2 of canopy (9.37% of the city). On the identical extent, 87.8% of mapped canopy pixels and 97.4% of candidate centers agreed with the 2022 USDA/CAL FIRE LiDAR-assisted canopy product; the optical surface represented 34.2% of the reference canopy area (IoU 0.288; Dice 0.448). Approximately 49% of candidates occurred within 15 m of a road. Canopy was inversely associated with Landsat land-surface temperature (Spearman rho = -0.293; partial rho = -0.370 controlling for built probability), and spatial-lag modeling confirmed clear neighborhood structure. Two transparent attention surfaces combined canopy need with thermal and contextual indicators. The framework provides a reproducible, updateable screening layer that complements structural canopy products and municipal inventories while retaining assumptions, data provenance, and spatial diagnostics for planning interpretation.

85. 【2608.13807】Label-Free Deep-Tissue Peripheral Nerve Detection with a Handheld Multimodal OCT Probe and NerveDetNet

链接https://arxiv.org/abs/2608.13807

作者:Yihan Wang,Ruilin You,Shaobai Li,Jiabin Chen,Bofan Song,Anh D. Le,Rongguang Liang

类目:Optics (physics.optics); Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph)

关键词:optical imaging methods, buried beneath intact, Peripheral nerves buried, white light wide-field, difficult to visualize

备注

点击查看摘要

Abstract:Peripheral nerves buried beneath intact tissue are difficult to visualize during surgery and remain inaccessible to white light wide-field imaging and other surface optical imaging methods. Existing OCT nerve studies have largely relied on exposed nerves or polarization contrast with limited depth penetration, restricting their value for subsurface intraoperative guidance. Here, we introduce, to our knowledge, the first label-free framework for detecting peripheral nerves beneath unopened tissue and resolving their depth using intensity-based OCT structural signatures alone. The framework combines a handheld multimodal probe, integrating swept-source OCT with co-registered white light and autofluorescence imaging, with a ``confirm-then-capture'' workflow designed for practical surgical use. To enable efficient analysis of sparsely sampled OCT volumes, we develop NerveDetNet, a lightweight 2.5D segmentation network that recovers weak and spatially displaced nerve signals by incorporating spatial context, frame-order information, and shift-tolerant correlations across frames through a dedicated nerve feature correlation module. In ex vivo tissue experiments, NerveDetNet consistently outperformed six representative 2D baselines across all frame spacings, achieving a Dice score of 0.725 under the sparsest sampling condition while using approximately half the model parameters. End-to-end validation demonstrated localization of nerves invisible at the surface and depth-resolved detection up to 1.3--1.4~mm below the tissue surface, with OCT derived depth maps overlaid directly onto the surgical view. Together, these results establish a practical label-free approach for subsurface nerve visualization that supports intraoperative compatibility, enables efficient sparse-volume analysis, and provides depth-resolved guidance without tissue opening, contrast agents, or nerve exposure.

86. 【2608.13791】VLM- and LLM-Driven Multi-Agent System for PET Image Denoising

链接https://arxiv.org/abs/2608.13791

作者:Boxiao Yu,Savas Ozdemir,Yang Xing,Fumio Hashimoto,Jiong Wu,Yizhou Chen,Axel Rominger,Ruogu Fang,Kuangyu Shi,Tinsu Pan,Kuang Gong

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Positron emission tomography, limited spatial resolution, Positron emission, compromise quantitative accuracy, PET image quality

备注

点击查看摘要

Abstract:Positron emission tomography (PET) imaging suffers from limited spatial resolution and low signal-to-noise ratio, which can compromise quantitative accuracy and lesion detectability. Deep learning-based denoising methods have demonstrated strong potential for improving PET image quality. However, their practical deployment in real-world settings remains challenging, often requiring multiple specialized models and expert interventions, such as identifying motion-induced misregistration artifacts, estimating noise levels to select an appropriate denoiser, and performing lesion-focused quantitative assessment after denoising. Recent advances in vision-language models (VLMs) for image quality understanding and large language models (LLMs) for contextual reasoning provide new opportunities for automated, decision-driven workflows. Inspired by expert workflows for PET image quality enhancement, we propose an VLM- and LLM-driven multi-agent PET denoising framework that dynamically assesses image quality and lesion status, autonomously selects optimal denoising models and parameters, and enables closed-loop feedback with rollback mechanisms. Experiments were conducted on Siemens Biograph Vision Quadra PET/CT data with 1/20 and 1/50 low-dose settings. Individual module evaluations demonstrated the reliability of the agentic components, while the complete framework achieved higher PSNR and SSIM than UNet, GAN, and DDPM baselines at both dose levels. These preliminary results demonstrate the feasibility of using a closed-loop multi-agent framework to adapt PET denoising strategies to different image conditions.

87. 【2608.13752】An Interactive, Automated 4D-STEM data acquisition and analysis routine for Scanning Electron Nanobeam Diffraction and Ptychography experiments

链接https://arxiv.org/abs/2608.13752

作者:Mohsen Danaie,Max England,Yiming Xu,Ruomu Zhang,Ed Darnbrough,Josh Willem De Boer,Frederick Allars,Zaeem Najeeb,Aakash Varambhia,Jinseok Ryu,Benjamin Bradnick,Damien McGrouther,Manfred E. Schuster,Christopher S. Allen

类目:Instrumentation and Detectors (physics.ins-det); Computer Vision and Pattern Recognition (cs.CV)

关键词:Modern transmission electron, Modern transmission, transmission electron microscopes, atomic scale, transmission electron

备注: 41 pages, 7 figures, 11 SI figures

点击查看摘要

Abstract:Modern transmission electron microscopes are versatile instruments which have become indispensable tools for understanding structure and chemical composition at the nano- and atomic scale. In the physical sciences these instruments are still largely manually controlled, requiring significant operator expertise, limiting throughput, and precluding statistical analysis of large datasets. Recent technical advances in both hardware and in control software now allow for the interaction with almost every functionality of the microscope through a programming interface. This enables better experimental design and data collection automation while also reducing operator collection bias and required expertise. In this study, we present an automated data collection routine with machine-driven decision-making to enable the collection of hundreds of 4D-STEM nanobeam diffraction and ptychography data from a large distribution of size-selectively deposited Pt nanoparticles. We present a semi-automated data analysis workflow to extract pertinent information from the large volumes of collected data. For the nanobeam diffraction data, reducing each dataset to its azimuthal variance profile and combining automated crystal orientation mapping with per-particle morphology descriptors reveals the orientation, shape and phase distributions across the ensemble, including a weak {110} texture. For the ptychographic data, an automated screening pipeline identifies on-zone-axis particles and enables atomic-resolution phase imaging and lattice-strain mapping of individual grains. Together these demonstrate how automation turns instrument throughput into statistically meaningful, atomic-scale microstructural information.

88. 【2608.13711】RUE-Colon: Exposing a Consistent Transfer Asymmetry in Real-Time Polyp Detection

链接https://arxiv.org/abs/2608.13711

作者:Sebastian Doerrich,Andreas Franz Schwab,Francesco Di Salvo,Shyam Nandan Rai,Hanh Huyen My Nguyen,Christian Ledig

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:clinical miss rates, reduce clinical miss, deployment remains elusive, reliable real-world deployment, real-world deployment remains

备注: Accepted to EndoLINA @ MICCAI 2026 (The International Workshop on Endoluminal Intervention Navigation and Autonomy)

点击查看摘要

Abstract:Computer-aided detection (CADe) systems for colonoscopy promise to reduce clinical miss rates, yet reliable real-world deployment remains elusive. This translational gap stems in part from a structural flaw in model development: the reliance on curated datasets that under-represent the long negative stretches and procedure-related artifacts characteristic of routine examinations. Training and evaluating architectures strictly on these lesion-centric benchmarks creates an illusion of success, since such benchmarks cannot capture clinically crucial metrics. To expose this gap, we establish TRUE-Colon, a standardized benchmarking protocol that measures key deployment characteristics alongside localization accuracy, and evaluate four real-time architectures (Faster R-CNN, YOLOv8, YOLOv11, RT-DETR) across curated benchmarks (SUN, PICCOLO) and 60 unedited, full-length procedures (REAL-Colon). We observe a consistent transfer asymmetry: models trained strictly on curated clips suffer a severe performance collapse when evaluated on full procedures, whereas procedure-trained models substantially improve rejection of non-polyp content on REAL-Colon, and largely retain their accuracy on curated benchmarks. Beyond transferability, we find that the Transformer detector attains the strongest sensitivity and the earliest, most persistent detections, while the convolutional detectors stay competitive at a higher throughput. Together, these results indicate that both training and benchmarking for deployable CADe should shift from curated, lesion-centric clips toward full-procedure data and deployment-relevant operating points. Source code is available at this https URL.

89. 【2608.13597】Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography with Diffusion Models

链接https://arxiv.org/abs/2608.13597

作者:Hongxin Xu,Jianping Mei,Can Wang,Defang Chen

类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Coverless image steganography, enabling authorized recipients, Coverless image, Existing diffusion-based CIS, existing cover image

备注

点击查看摘要

Abstract:Coverless image steganography (CIS) synthesizes a stego image rather than modifying an existing cover image, enabling authorized recipients to reconstruct the original secret image from the stego. Existing diffusion-based CIS methods can generate natural-looking stego images but preserve substantial visual similarity to the secret image. This resemblance risks exposing structural and semantic cues, giving rise to security vulnerabilities that cannot be evaluated solely via recovery fidelity. Achieving substantial visual dissimilarity between the secret and stego images without compromising stego quality and recovery fidelity remains challenging. To address this issue, we propose InvCISD, an invertible diffusion framework that couples the latent representations of the secret and an irrelevant reference image with an invertible network called LIMNet. We first train LIMNet in diffusion latent space, followed by end-to-end fine-tuning of the entire network, i.e., LIMNet integrated diffusion inversion and generation modules. Experiments demonstrate that the proposed method substantially reduces secret-stego visual similarity, improves stego quality, and retains satisfactory secret reconstruction quality. Our further investigation shows that all evaluated methods are highly detectable by the CIS-oriented steganalysis model, indicating that resistance against targeted steganalysis constitutes a critical direction for future CIS research.