本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。

统计

今日共更新622篇论文,其中:

  • 自然语言处理72
  • 信息检索12
  • 计算机视觉103

自然语言处理

1. 【2608.18072】Multi-Agent AI System for Radiology Report Structuring and Quality Assurance with Independent Radiologist Evaluation

链接https://arxiv.org/abs/2608.18072

作者:Iryna Hartsock,Cesar Lam,Christopher Otteni,Aliya Qayyum,Robert Gatenby,Cyrillo Araujo,Ghulam Rasool

类目:Computation and Language (cs.CL)

关键词:structuring and quality, reports, develop and evaluate, quality assurance, Purpose

备注: 14 pages, 2 figures, 4 tables

点击查看摘要

Abstract:Purpose: To develop and evaluate a locally deployed multi-agent AI system for radiology report structuring and quality assurance. Materials and Methods: This retrospective study included 638 radiology reports from CT examinations of the chest, abdomen, and pelvis dictated by 15 board-certified radiologists in 2023 and 2024. A multi-agent AI pipeline was developed to perform report structuring and quality assurance (QA). The system structured the report into standardized anatomical sections at the sentence level using regex rules and local large language models. It also detected mismatches between the Findings and Impression sections, or within sections; gender-anatomy conflicts; and undocumented communication of critical findings. Two board-certified radiologists independently evaluated a 45-report subset. Results: The multi-agent system structured the Findings sections of all reports (22,270 sentences) into a predefined anatomical format while retaining the original report content. The system flagged 90 (14.1%) reports, most commonly for section mismatches (80 reports, 12.5%). In the radiologist evaluation, both reviewers agreed that 31 (69%) were correctly restructured, 2 reports (4%) were incorrectly restructured, and disagreed on the remaining 12 reports (27%). Both reviewers agreed that no clinically important information was omitted and no fabricated content was introduced. Overall QA performance was rated as "excellent" or "good" in 84% of the evaluated reports, with the remaining reports rated as "fair". Conclusion: A locally deployed multi-agent AI system combined radiology report structuring and quality assurance within a single workflow. The system demonstrated favorable performance in radiologist evaluation. Such systems may support standardization of reporting and quality assurance in radiology practice.

2. 【2608.18066】On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

链接https://arxiv.org/abs/2608.18066

作者:Qinyuan Ye,Yu Li,Yada Pruksachatkun,Jiaxin Zhang,Chien-Sheng Wu

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:shown great promise, textual memory bank, recent literature, online stream, improve over time

备注: Code: [this https URL](https://github.com/SalesforceAIResearch/self-improve-fragility) Data: [this https URL](https://huggingface.co/datasets/Salesforce/self-improve-fragility)

点击查看摘要

Abstract:Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self-improving agents by reporting results across multiple runs and stress-testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.

Comments:
Code: this https URL Data: this https URL

Subjects:

Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

Cite as:
arXiv:2608.18066 [cs.AI]

(or
arXiv:2608.18066v1 [cs.AI] for this version)

https://doi.org/10.48550/arXiv.2608.18066

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
3. 【2608.18062】okEval: A Tokenizer Evaluation Suite

链接https://arxiv.org/abs/2608.18062

作者:Clara Meister

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:design choices directly, choices directly impact, typically selected, selected with minimal, design choices

备注: Published as a conference paper at COLM 2026; Library hosted at [this https URL](https://github.com/cimeister/tokenizer-intrinsic-evals)

点击查看摘要

Abstract:Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.

4. 【2608.18041】Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

链接https://arxiv.org/abs/2608.18041

作者:Hollis Robbins(University of Utah)

类目:Computation and Language (cs.CL)

关键词:phase, parameter, meaning stays, meaning, Abstract

备注: 23 pages; 0 figuresCC

点击查看摘要

Abstract:Language has two parameters. Count how often words occur together and you estimate amplitude, the strength of association. Word embeddings and attention weights refine that count, which sums every writer in the corpus together. This paper claims a second parameter, phase, which signed weights learned from a corpus do not supply. Phase exists only between meanings: it determines how coactivated meanings combine, and it can reverse what a meaning contributes while that meaning stays fully present. A speaker can set phase in the signal through linguistic form; encounters install phase relations and history distributes them. Population averaging deletes history-indexed phase: agent-deindexed corpora identify the population marginal state and determine no individual or dyadic state, at any scale. The standard transformer has no explicit representation for phase in frozen inference, and the interpretability program measuring progress by monosemanticity is optimizing against it: the coexistence it treats as a defect is the condition of allusion, irony, and quotation. Six predictions test whether a suppressed meaning stays active, whether encounter order changes what a phrase does, whether marking the signal changes how a shared phrase is taken, and whether a model given a history is changed by it or only informed about it. The claim defended is the weak version: interpretation requires a second relational parameter, signed, persistent, and indexed to individuals and dyads. Quantum probability is one notation for the parameter; nothing in the formalism claims quantum processes in the brain. The strong version, that the quantum calculus constrains these phenomena as signed classical models do not, rests on an encounter-order constraint not yet derived. The architecture the theory calls for is a language model with agent-indexed, phase-bearing semantic states.

5. 【2608.18027】Chain-of-Experience for Continual LLM Improvement

链接https://arxiv.org/abs/2608.18027

作者:Haoqin Tu,Yunhao Fang,Yizhong Wang,Cihang Xie,Shen Yan

类目:Computation and Language (cs.CL)

关键词:Humans continuously learn, conventional large language, Humans continuously, large language model, evaluations ignore

备注: H.T. and Y.F. contributed to this work equally

点击查看摘要

Abstract:Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.

6. 【2608.18011】he IOL-AI Challenge: An Open Challenge towards Advancing Linguistic Reasoning

链接https://arxiv.org/abs/2608.18011

作者:Eduardo Sánchez,Rita Berrada,Dan-Mircea Mirea,Sara Rajaee,Alexander Piperski,Ana Meta Dolinar,Boris Iomdin,Andrey Nikulin,Mariya Shmatova,Marzieh Fadaee,Julia Kreutzer

类目:Computation and Language (cs.CL)

关键词:International Linguistics Olympiad, mathematics and code, linguistic reasoning, LLMs is overwhelmingly, overwhelmingly studied

备注

点击查看摘要

Abstract:Reasoning in LLMs is overwhelmingly studied in domains that provide a model with rules: mathematics and code. Linguistic puzzles invert this: the solver must first discover the system before reasoning within it. We present the IOL-AI Challenge, an open-science competition run on the unseen problems of the International Linguistics Olympiad (IOL) 2026 Individual Contest, evaluated both automatically and, for the first time, by members of the official IOL Jury under the same rubrics applied to human contestants. The challenge drew 731 submissions from 46 teams under a strict compute budget (one T4, 30 mins). We additionally benchmark 15 unconstrained frontier and open models, with Claude Opus 4.8 earning a jury score equivalent to a gold medal, while both resource-constrained systems we submitted for jury grading scored in the range of the bottom 5% of contestants. Capability was not determined by scale: 14B submissions outperform models twice their size, and gains come from decoding and output-handling rather than model capacity. We also found that automatic metrics rank systems exactly as the jury does, but compress the scale, upscoring weak systems by ~13 points and understating strong ones. Our analysis shows that while frontier models might have prior knowledge about some of the problem languages, it does not significantly help them solve the linguistic reasoning tasks, leaving linguistic reasoning as a strong benchmarking proxy for generalizable reasoning skills.

7. 【2608.17994】Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

链接https://arxiv.org/abs/2608.17994

作者:Sher Badshah,Ali Emami,Hassan Sajjad

类目:Computation and Language (cs.CL)

关键词:evaluating model outputs, standard practice, practice for evaluating, evaluating model, model outputs

备注: Accepted at Conference on Language Modelling 2026

点击查看摘要

Abstract:Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$\alpha$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.

8. 【2608.17987】Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

链接https://arxiv.org/abs/2608.17987

作者:Yijie Xu,Chao Wang,Hui Xiong

类目:ocial and Information Networks (cs.SI); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:influenced political discourse, greatly influenced political, understand individual political, individual political ideologies, rapid growth

备注: Accepted by ACM Transactions on Intelligent Systems and Technology

点击查看摘要

Abstract:The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.

9. 【2608.17979】When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts in Genre, Time and the AI-Era

链接https://arxiv.org/abs/2608.17979

作者:Lotta Kiefer,Brisca Balthes,Christoph Leiter,Yamen Ajjour,Elena Schmidt,Steffen Eger

类目:Computation and Language (cs.CL)

关键词:Authorship verification, author writing style, style remains sufficiently, writing style remains, Authorship

备注

点击查看摘要

Abstract:Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by distribution shifts caused by changes in genre, time, and AI-assisted writing. Existing AV benchmarks typically study these factors in isolation and focus predominantly on English, limiting our understanding of model robustness under realistic conditions. We introduce AVShift, the first German benchmark for systematically evaluating AV under multiple distribution shifts. AVShift comprises over 150,000 text pairs spanning three genres and 21 years, enabling controlled evaluation of cross-genre, temporal, and AI-era shifts within a unified framework. We benchmark representative feature-based, embedding-based, and LLM-based approaches. Our experiments show that fine-tuned LLMs generalize best across genres and benefit substantially from stylistically diverse training data. We further demonstrate that temporal drift is one of the strongest factors affecting AV, with performance degrading significantly as the time gap between documents increases. In contrast, we find no evidence of a measurable AI-era distribution shift within AVShift. Finally, our feature analysis reveals stylistic features that remain stable across genres, while their relative importance varies depending on the specific genre transition. We release AVShift and our code for future research.

10. 【2608.17950】Do Large Language Models Play Six Degrees of Separation? Measuring Topological Compression in Long-Context Manifolds

链接https://arxiv.org/abs/2608.17950

作者:Md. Faiyaz Abdullah Sayeedi

类目:Computation and Language (cs.CL)

关键词:Large Language Models, Large Language, Language Models, internal mechanisms enabling, distant cognitive leaps

备注

点击查看摘要

Abstract:Large Language Models (LLMs) demonstrate remarkable multi-hop reasoning capabilities over long contexts, yet the internal mechanisms enabling these distant cognitive leaps remain poorly understood. Traditional attention-based interpretability often fails to capture true semantic proximity due to routing artifacts like attention sinks. In this paper, we bypass attention weights to directly analyze the dynamic geometry of the hidden state manifold, proving that deep LLM latent spaces natively organize into Small-World networks. By sparsifying the continuous similarity matrices of long-context representations into unweighted graphs, we trace the connectivity between highly disjoint semantic anchors across two distinct architectures. Our findings reveal a sharp topological phase transition: while early syntactic layers remain entirely fractured, deep reasoning layers abruptly compress massive conceptual distances into highly navigable pathways strictly bounded by the "Six Degrees of Separation" limit (= 6 semantic hops). Furthermore, we demonstrate the practical efficacy of this framework by applying it to zero-shot hallucination detection within Retrieval-Augmented Generation (RAG) using the RAGognize dataset. We show that factually grounded generations maintain structural integrity with their source context (approximately 3 hops), whereas hallucinations induce severe topological collapse. Ultimately, this work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability.

11. 【2608.17941】Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

链接https://arxiv.org/abs/2608.17941

作者:Zhizhao Liu,Zhiliang Tian,Xi Wang,Zhihua Wen,Yihang Xiong,Zhiquan Lai,Dongsheng Li

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Reinforcement learning, large language models, verifiable rewards, learning with verifiable, capabilities of large

备注

点击查看摘要

Abstract:Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates remains challenging: dedicated probing adds substantial generation overhead, whereas history-based estimators face a cold start with no initial observations and stale feedback, and typically ignore relations among samples. To address these limitations, we propose a plug-and-play graph-based online difficulty estimator that shares rollout feedback across related samples and continuously updates their difficulty estimates, mitigating cold start and staleness without dedicated probing. Specifically, we first construct a difficulty-aware sample graph based on semantic and reasoning similarities. Based on this graph, we introduce latent difficulty states and use a Potts prior to encourage neighboring samples to share the same state. We then employ a state-level Beta-Binomial model to aggregate the rollout outcomes associated with each state. Finally, we use an online mean-field variational algorithm to continuously update the latent-state assignments and state-level difficulty as new feedback arrives. Our framework can be integrated into sample-selection and rollout-allocation schedulers, enabling difficulty-adaptive exploration without dedicated probing. Experiments across multiple base models, RL schedulers, and benchmarks demonstrate that our framework achieves better performance.

12. 【2608.17938】Grading Needs a Rubric, Not Intelligence

链接https://arxiv.org/abs/2608.17938

作者:Jhen-Ke Lin

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Small language models, Small language, reliably as substantially, substantially more expensive, Small

备注

点击查看摘要

Abstract:Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.

13. 【2608.17931】SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis

链接https://arxiv.org/abs/2608.17931

作者:Shicheng Ma,Wenqian Cui,Irwin King

类目:Computation and Language (cs.CL); Multimedia (cs.MM); Sound (cs.SD)

关键词:understanding requires discerning, Speech Sentiment Analysis, Recent advances, revolutionized speech processing, effective speech understanding

备注: 7 pages, 2 figures, 5 tables. Accepted to ACM Multimedia 2026 (Dataset Track). Dataset and code: [this https URL](https://github.com/Sher13cked/SpeechSense)

点击查看摘要

Abstract:Recent advances in AI have revolutionized speech processing, yet effective speech understanding requires discerning not just what is said, but how it is said. Speech Sentiment Analysis plays a critical role in decoding these paralinguistic cues for diverse real-world applications such as recruitment and customer service. However, existing Speech Sentiment Analysis research faces two primary limitations. First, dominant approaches rely on text-centric pipelines that cascade Automatic Speech Recognition with text analysis. This process inevitably discards essential acoustic features like prosody and tone, failing to capture attitudinal meanings in acoustically ambiguous utterances. Second, current benchmarks suffer from a mismatch in label granularity, prioritizing basic emotions (e.g., happy, sad) over the nuanced interpersonal stances (e.g., confident, impatient) necessary for social sensitivity. To address these limitations, we propose a novel dataset, SpeechSense, for fine-grained speech sentiment analysis. Specifically, we define a specialized 8-class taxonomy of interpersonal stances detectable primarily through prosodic cues beyond lexical content alone. We then construct a curated dataset based on this taxonomy, built from high-fidelity speech synthesis and rigorous human validation. Comprehensive experiments across multi-modal LLMs, text-only LLMs, and speech encoders demonstrate that models with acoustic access consistently outperform text-only baselines. These results empirically validate the primacy of acoustic cues in detecting subtle speaker attitudes, highlighting the necessity of SpeechSense. Dataset and supplementary materials are available at this https URL.

14. 【2608.17911】CABLE: Extending the Reach of Memory Retrieval via Complementary Antecedent-Based Linking and Expansion

链接https://arxiv.org/abs/2608.17911

作者:Zheling Tan,Jin Gao,Dequan Wang

类目:Computation and Language (cs.CL)

关键词:LLM agents operate, preserving long-term history, LLM agents, bounded memory interface, agents operate

备注: Accepted by COLM 2026

点击查看摘要

Abstract:As LLM agents operate across structured workflows and sessions, preserving long-term history does not ensure that later contexts can recover relevant evidence through a bounded memory interface. We study this evidence-reachability problem in long-term conversational memory, where retrieval still relies heavily on semantic similarity. This works well for topical recall, but it often misses earlier experiences, plans, or motivations that are semantically distant from the later events they help explain. Existing memory graphs provide cross-memory structure, yet links driven mainly by semantic overlap can duplicate what the host retriever already recovers. We argue that link construction should instead prioritize a sparse set of retriever-complementary associations. We present CABLE (Complementary Antecedent-Based Linking and Expansion), a plug-in augmentation that constructs links designed to extend the host retriever's direct semantic reach. For each new memory, CABLE generates antecedent-oriented queries, retrieves prior memories, subtracts candidates in the direct semantic neighborhood, and verifies the remainder before adding the accepted complementary associations into a sparse directed graph. At retrieval time, CABLE expands the host system's retrieved seeds along these links to surface implicit supporting evidence. We evaluate CABLE with A-MEM on LoCoMo and MA-LongMemEval, and further integrate it into SimpleMem and Mem0g on LoCoMo, using Qwen3.5-27B, DeepSeek-chat, and GPT-4o-mini. CABLE yields higher mean LLM-judge scores in every evaluated system-level setting, with the largest gains in categories where useful evidence is distributed across memories or sessions, including open-domain, multi-session, and preference-oriented questions. These results support prioritizing sparse, reasoning-relevant associations that complement rather than duplicate the host retriever.

15. 【2608.17895】BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

链接https://arxiv.org/abs/2608.17895

作者:Liubov Chubarova,Alexandra Kuleshova,Daniil Volkov,Kirill Sultanov,Alexey Zaytsev

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Multimodal Large Language, remains incompletely evaluated, Multimodal Large, made significant strides, Large Language Models

备注

点击查看摘要

Abstract:While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.

16. 【2608.17866】BayesPrompt: human readable prompts that make sense

链接https://arxiv.org/abs/2608.17866

作者:Franky Kevin Nando Tezoh,Ali Hussaini Umar,Alessandro Laio,Guido Sanguinetti,Riccardo Rende

类目:Computation and Language (cs.CL)

关键词:important research topic, research topic, Reconstructing prompts, elicit a desired, open and important

备注

点击查看摘要

Abstract:Reconstructing prompts that can elicit a desired answer or behaviour in an LLM is an open and important research topic. Optimisation methods which aim at minimising the perplexity of a given answer, however, consistently yield so-called pseudoprompts, unintelligible strings of tokens which can lack human interpretability. We argue that this is a consequence of the ill-posedness of the prompt optimisation task. By reframing the task as a Bayesian posterior inference over prompts, we propose an efficient algorithm to sample prompts which are both efficient (in terms of perplexity) and human readable. We compare our approach with state of the art alternatives showing on a real data set a marked improvement over a range of metrics.

17. 【2608.17843】Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

链接https://arxiv.org/abs/2608.17843

作者:Man Liang,Xinzhao Cheng,Faizan Wajid

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Large language models, behavior remain unclear, informs model behavior, model behavior remain, structured reasoning tasks

备注: 13 pages, 7 figures, 8 tables, including appendices

点击查看摘要

Abstract:Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it informs model behavior remain unclear. We investigate this question through geometric reasoning, using parametric CAD constraints as a controlled testbed for separating local pairwise relations from sketch-level constraint status. By probing the hidden states of six frozen decoder-only LLMs, we examine four properties: linear decodability, forced-choice generation, activation-level influence, and behavioral steerability. Pretraining substantially improves the decoding of local geometric relations, and this advantage persists after accounting for positional cues with shuffled-order controls. In contrast, sketch-level DOF status is already highly decodable from randomly initialized representations and improves only modestly with pretraining, indicating that much of its probe performance is available without learned weights. Further analyses show that decodable information is not always actionable. Generation often fails to express this information, and on the two intervention-tested backbones, activation-restoration effects at the patched entity position vanish while decodability persists across depth. Mean-difference steering also does not reliably control outputs. These results show that decodability, generation, activation-level influence, and steerability can diverge in the tested setting. The audit provides a controlled way to distinguish failures to encode geometric structure from failures to express or control encoded information.

18. 【2608.17827】From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector

链接https://arxiv.org/abs/2608.17827

作者:Camilla Dalerci,Thilo Michael,Robin Schaefer,Daniel Weinland

类目:Computation and Language (cs.CL)

关键词:selecting LLMs suited, face a persistent, persistent challenge, challenge in selecting, selecting LLMs

备注: Accepted as non-archival paper at Eval4SD (co-located with KONVENS 2026)

点击查看摘要

Abstract:Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of MÖVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.

19. 【2608.17810】Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

链接https://arxiv.org/abs/2608.17810

作者:Alona Strugatski,Licol Zeinfeld,Jason Cooper,Shelley Rap,Gil Schwarts,Giora Alexandron

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC)

关键词:large language models, employ similar underlying, similar underlying cognitive, humans employ similar, underlying cognitive constructs

备注: Accepted for publication at AIME 2026

点击查看摘要

Abstract:The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.

20. 【2608.17809】Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

链接https://arxiv.org/abs/2608.17809

作者:Quang Minh Nguyen,Luis Frentzen Salim

类目:Computation and Language (cs.CL)

关键词:Humans naturally form, Humans naturally, daily communication, naturally form, Humans

备注: In submission

点击查看摘要

Abstract:Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of the weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on "I vaguely remember" to -14% on "I seriously doubt". We further show that the phenomenon stems from task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief; chains of thought that explicitly fact-check show lower accuracy on false information than those that do not; and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs. Our code is available at this https URL.

21. 【2608.17804】An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

链接https://arxiv.org/abs/2608.17804

作者:Rubén Balbastre,Juan Manuel Orduña,Mariano Pérez

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Practical LLM unlearning, preserve non-target utility, Practical LLM, suppress target-specific knowledge, non-target utility

备注: 32 pages, 4 figures. Code and artifacts linked in the paper

点击查看摘要

Abstract:Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and rewards that can select broad-topic answering with low semantic leakage during optimization.

22. 【2608.17795】raceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

链接https://arxiv.org/abs/2608.17795

作者:Neelesh Kumar Shukla,Debasmita Panda,Srutanik Bhaduri,Aditya Banerjee,Viji Krishnamurthy

类目:Computation and Language (cs.CL)

关键词:ground-truth SQL queries, reference execution results, real-world deployments, commonly evaluated, evaluated using ground-truth

备注: 9 pages main paper with 6 pages supplementary material

点击查看摘要

Abstract:Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification problem: given only a user question, database context, and generated SQL, can a system estimate whether the generated query is likely to correctly answer the question? Recent approaches use LLMs as judge or specialized agents to inspect generated SQL, but their decisions can be difficult to trace. Outcome Reward Models (ORMs) address this by learning from execution-labeled candidate SQLs and assigning correctness scores to unseen queries, yet they still provide limited visibility into the signals behind each verification. To address this limitation, we propose TraceSQL, a lightweight and traceable verification model built on explicit diagnostic features. TraceSQL combines 67 features capturing question ambiguity, question requirements, question-schema-SQL consistency, SQL structure, and intent alignment. These signals remain available for examining which factors influence each prediction and for tracing decisions back to diagnostic evidence. On BIRD development databases, TraceSQL achieves 66.47% F1 and 64.48% ROC-AUC, compared with 61.87% F1 and 58.26% ROC-AUC for the GradeSQL-7B ORM baseline on the same generated-SQL evaluation. Feature attribution further shows that the model relies on both semantic grounding and deterministic SQL-structure signals. These results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions.

23. 【2608.17781】Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

链接https://arxiv.org/abs/2608.17781

作者:Shi Zhou

类目:Computation and Language (cs.CL)

关键词:downstream model identity, systems increasingly condition, model-specific differences form, differences form reusable, form reusable structure

备注: 16 pages, 6 figures, 11 tables

点击查看摘要

Abstract:ML systems increasingly condition decisions on downstream model identity, but this is useful only if model-specific differences form reusable structure rather than input-local interactions. We test this in retrieval-augmented generation (RAG), where evidence utility can be measured under controlled interventions. Holding query, evidence, task, scoring, and intervention fixed, nine readers disagree on effect sign in 33\% of jointly affected cells; reader$\times$query interaction explains 29.8\% of utility variance versus an 8.4\% permutation null; and self-selected evidence improves F1 by $+0.031$ ($t=3.39$). We then ask the sharper question: \emph{which components of this heterogeneity are stable reader properties across queries?} Separating three measurable objects---evidence \emph{activity}, \emph{ordinal preference}, and \emph{conditional signed direction}---we find ordinal reader geometry stable across four independent settings (split-half $\rho=0.60$--$0.83$): leave-one-out interventions, PRISM preferences, RAMDocs, and RAGuard. Signed geometry is task-bounded: weak in open-ended QA (0.14, 0.35), especially for misleading and irrelevant evidence, but strong in binary fact-checking (0.75) with no significant ordinal gap, though still below its sparsity-matched ceiling. Sparsity, decoding noise, and metric artifacts do not explain the main ordinal--signed gap. Finally, stable ordinal similarity fails to predict cross-reader intervention transfer (oracle-distance $\rho=-0.27$; regret reliability $-0.28$). Reader-specific utility exists, but preference is not intervention: stable ranking similarity does not license transfer of help/harm decisions.

24. 【2608.17744】hinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

链接https://arxiv.org/abs/2608.17744

作者:Ayoub Kirouane,Christos Petrocheilos

类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Robotics (cs.RO); Machine Learning (stat.ML)

关键词:active parameters, Alibaba, NVIDIA, Greek, random seed moves

备注

点击查看摘要

Abstract:Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. Base models never think in Greek: 0 of 1,000 reasoning traces, even when the question is Greek, so the model answers correctly while reasoning in a form its user cannot read, audit, or correct. After supervised fine-tuning (SFT), every released checkpoint reasons in the language of the question on ~98% of items, one family at 3x fewer tokens, with judged grammaticality improving on all four models and general ability within a few points of each base: nothing was forgotten, and fluency was gained. We propose six behavioural dimensions that make such changes measurable, each gated to reject any metric that correlates with output length, and we report how our own instruments lied: six failures, each caught by a control. What SFT cannot do is fix its own defects: a quarter of answers skip the requested format, answers leak into the reasoning channel, and an explicit "think in English" is obeyed under half the time. Reinforcement learning with verifiable rewards, pre-registered before training, fixes the first two outright (fallback 24% to 2.5%, leak 3.5% to 0.0%, both against a flat random-reward control) and moves the third (+9.1pp), while the Greek reasoning habit survives an accuracy-only gradient untouched. We release five checkpoints. The instruments, the controls and the pre-registration travel to any low-resource language; Greek is the case that let us measure them.

25. 【2608.17719】What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

链接https://arxiv.org/abs/2608.17719

作者:Xiaonan Xu,Wenjing Wu

类目:oftware Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:commercial large language, vendors deprecate older, Software systems, language model APIs, large language model

备注: 25 pages, 1 figure, 10 tables (including 8 appendix tables)

点击查看摘要

Abstract:Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to GPT-5.6 Sol product sequence, we query 900 public benchmark items (graduate-level knowledge, olympiad mathematics, instruction following) 50 times per item per model, classify each item as reliably improved, reliably regressed, practically equivalent, or inconclusive under false-discovery-rate control and a practical-significance threshold, and calibrate the results against a label-permutation null. Results: Across all nine migration-benchmark cells, reliable improvements and reliable regressions coexist. Edges with aggregate gains of up to 7.3 percentage points contain up to 8.3% reliably regressed items; edges with aggregate losses contain up to 10.7% reliably improved items. On the instruction-following benchmark, the gap between strict and loose scoring widens by 3.9 percentage points on the latest migration: a 3.9-point regression under strict scoring shrinks to 0.04 points under loose scoring. Conclusion: Migration decisions based on aggregate scores alone miss substantial bidirectional item-level change. The complete response-level archive and per-item scoring outputs are released.

26. 【2608.17644】LLM-Derived Preference Judgments Are Not Self-Consistent

链接https://arxiv.org/abs/2608.17644

作者:Matthew T. Ford,Francis Bahk,Jingjing Wang,Adam S. Jovine,Tinghan Ye,David B. Shmoys,Peter I. Frazier

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:Agents increasingly interpret, Agents increasingly, increasingly interpret, Agents, utility function

备注: 16 pages, 4 figures; includes appendices

点击查看摘要

Abstract:Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.

27. 【2608.17616】MoNe: Modular Neural Memory for Efficient Long Context Inference

链接https://arxiv.org/abs/2608.17616

作者:Wonguk Cho,Kyubyung Chae,Tribhuvanesh Orekondy,Sunghyun Park,Hyoungwoo Park,Jeongho Kim,Arash Behboodi,Kyuwoong Hwang,Sungrack Yun

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:frozen pretrained Transformer, lightweight modular neural, pretrained Transformer, Transformer to enable, enable long-context inference

备注

点击查看摘要

Abstract:We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.

28. 【2608.17605】Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

链接https://arxiv.org/abs/2608.17605

作者:Syeda Faiza Ahmed,Zien Sheikh Ali,Hunzalah Hassan Bhatti,Firoj Alam,Shammur Absar Chowdhury

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD)

关键词:isolated text prompts, prompts toward sustained, moving beyond isolated, isolated text, text prompts

备注: Multi-turn Conversational AI; Multimodal Dialogue; AudioLLMs; Conversational Memory; Tool-Augmented Agents; Dialogue Evaluation

点击查看摘要

Abstract:Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (this https URL)

29. 【2608.17587】Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

链接https://arxiv.org/abs/2608.17587

作者:Kang Peng,Zhiwei Zhang,Yichen Zhang,Zezhong Wang,Yiming Du,Geng Tu,Baojun Wang,Bin Liang,Ruifeng Xu,Kam-Fai Wong

类目:Computation and Language (cs.CL)

关键词:Expert-written natural language, Expert-written natural, natural language skills, agent-authored skills perform, skills perform 8-11

备注

点击查看摘要

Abstract:Expert-written natural language skills can improve tool-using agents, yet agent-authored skills perform 8-11 points worse than using no skill. This gap suggests that following procedural guidance and improving it from execution evidence are distinct capabilities. Inference time loops can repair skills but do not improve the model that writes the next one. We study how to organize execution experience from intermediate skills into training states for an optimizer. We introduce WER (Write, Execute, and Refine), a multi-phase framework that trains a Skill Optimizer outside a frozen executor. The optimizer proposes skills, a frozen agent executes each repeatedly, and a programmatic verifier scores the outcomes. The scores provide relative credit and select mixed-outcome records. Matched successful and failed trajectories from these records form the next phase's refinement states, so the optimizer learns from the consequences of its earlier outputs. On BFCL v4 multi-turn and tau2-bench, WER improves average Pass@1 over the no-skill baseline by 7.80 and 3.85 points, respectively. Under an identical refinement workflow, it outperforms the same backbone without optimizer training by 9.35 and 10.29 points. The trained 4B optimizer reaches 76.63 percent on BFCL v4, outperforming all evaluated off-the-shelf general-purpose models used as skill optimizers on average.

30. 【2608.17583】Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study

链接https://arxiv.org/abs/2608.17583

作者:Hamidreza Saffari,Francesco Pierri

类目:Computation and Language (cs.CL)

关键词:Online video platforms, expose young users, moderation judgments vary, independent audits remain, audits remain difficult

备注: 20 pages, 16 figures, 14 tables. Accepted to Findings of EMNLP 2026

点击查看摘要

Abstract:Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.

31. 【2608.17567】Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

链接https://arxiv.org/abs/2608.17567

作者:Henrik Wille,Luis-Finley Schütz,Felix Strieth-Kalthoff

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:learning structure-property relationships, Pretrained molecular language, Pretrained molecular, molecular language, molecular language models

备注

点击查看摘要

Abstract:Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.

32. 【2608.17556】Reflex-Guard: A Low-Latency Guardrail for LLM Prompt Safety Using Dense Semantic Embeddings

链接https://arxiv.org/abs/2608.17556

作者:Istiaque Ahmed,Afia Anjum Borsha,Ranat Das Prangon,Abu-fuad Ahmad,Thi Hong Tran

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Large Language Models, Large Language, Language Models, specially crafted prompts, crafted prompts designed

备注

点击查看摘要

Abstract:Large Language Models (LLMs) in real-world applications often face the risks of specially crafted prompts designed to bypass the safety controls. Existing guardrail methods, such as LLM-as-a-judge and cloud-based safety APIs are able to detect unsafe content. However, they often add a delay of about 250-900 ms to each request. This delay is too high for real-time applications, when the system usually needs to respond in less than 100 ms. Furthermore, routing user prompts through external moderation endpoints raises significant data privacy concerns. This paper introduces Reflex-Guard, a lightweight guardrail that runs locally. It uses jailbreak-aware preprocessing, compact sentence-transformer embeddings, and seven fast binary classifiers. Together, these components enable high-accuracy prompt safety filtering with much lower latency than existing solutions. Through systematic evaluation on a strategically balanced dataset of 30,568 samples drawn from five complementary sources, we demonstrate that Reflex-Guard achieves 95.9% recall on harmful prompts at 37.6 ms end-to-end latency. It is faster than existing baselines, including Llama Guard 2 at 255 ms and SafeDecoding at 723 ms. It can detect 100% of GCG suffix attacks and Base64-encoded prompts using the default threshold. However, DrAttack structured prompts required lowering the threshold to 0.03 for optimal detection, as they produced a distinct probability distribution. Reflex-Guard achieves Reflex Efficiency Score (RES) scores up to 16.79, significantly outperforming Llama Guard 2 (11.90) and SafeDecoding (9.80). This analysis offers practical deployment advice and shows that different attack types occupy distinct regions in the embedding probability space.

33. 【2608.17550】Code as Representation: A Compilable Parsing Paradigm for Academic Documents

链接https://arxiv.org/abs/2608.17550

作者:Rihui Jin,Jun Wang,chengyuan zhu,Liang Mingyu,Yue Gao,Li Yunxuan,Kuicai Dong,Guilin Qi,Lin Ren,Yongrui Chen,Xinbang Dai,Jiaqi Li,Tongtong Wu,Gholamreza Haffari

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:knowledge remains locked, Multimodal Large Language, Structured Academic Elements, Large Language Models, knowledge remains

备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core challenge is not only perception, but representation: scientific pages interleave text with Structured Academic Elements (SAEs) such as tables, formulas, charts, and pseudocode, whose structure, data, and logic are poorly preserved by common surrogates like Markdown. We therefore propose Compilable Academic Document Parsing (CADP), a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page. To support this setting, we introduce CADP-Bench, an expert-verified benchmark of full academic pages containing tightly coupled text and multiple SAE types, evaluated through a re-injection compilation protocol. We further study current capabilities using SOTA MLLMs and an exploratory multi-agent baseline that incorporates common agentic techniques. Results show that even frontier models still struggle to produce high-fidelity executable reconstructions, highlighting substantial room for improvement in structure-aware scientific document parsing. CADP-Bench is released for future research.

34. 【2608.17536】CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

链接https://arxiv.org/abs/2608.17536

作者:Jin Su,Zhuofeng Zhao,Huanhuan Wang,Hao Chen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:exhibit multi-level complexity, consultation questions exhibit, questions exhibit multi-level, exhibit multi-level, Legal consultation questions

备注

点击查看摘要

Abstract:Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions and poor interpretability for complex ones, making it difficult to meet the requirements for both answer quality and efficiency in high-risk scenarios. To address this issue, this paper proposes CoAL-RAG, a complexity-aware legal retrieval-augmented generation method, which constructs a multi-dimensional evaluation mechanism based on ``question essence'' and ``retrieval consistency'' to enable adaptive routing of retrieval strategies. First, the reasoning demand is quantified according to the logical structure of the question. Then, the discrepancy between semantic retrieval and keyword retrieval is utilized to indirectly reflect problem complexity, thereby selecting the most appropriate retrieval strategy and dynamically filtering contextual information. Experimental results demonstrate that the proposed method significantly outperforms baseline models not only on Chinese legal benchmarks (SocialLawQA, LawBench) but also demonstrates strong cross-jurisdictional generalization on English datasets (LexGLUE, CaseHold). Specifically, on Chinese datasets, the BLEU score improves by 42.5\% and ROUGE-L reaches 3.6 times that of knowledge graph-based methods. On English benchmarks, CoAL-RAG maintains highly competitive accuracy, achieving an optimal balance between generation quality, deep logical reasoning, and system efficiency across different legal systems.

35. 【2608.17534】ArborMem: Navigating Interaction States with Memory Forests

链接https://arxiv.org/abs/2608.17534

作者:Zongwei Lv,Yuemeng Xu,Yilun Yao,Siyi Ding,Xinyu Tan,Yaoming Li,Guangxiang Zhao,Weihong Lin,Lin Sun,Xiangzheng Zhang,Tong Yang

类目:Computation and Language (cs.CL)

关键词:Large language models, language models increasingly, models increasingly serve, persistent conversational assistants, Large language

备注: 24 pages, 2 figures

点击查看摘要

Abstract:Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains continuity across interactions. Existing methods improve access to conversational history through long-context processing, selective retrieval, and structured memory organization. However, most systems treat memory access as retrieving relevant past information without first determining which prior interaction state the current turn resumes. This limitation becomes particularly important when conversations interleave multiple tasks, people, and plans that may be interrupted and later revisited. We introduce ArborMem, an online memory framework that represents a long-running conversation as a navigable forest of interaction states. Each branch preserves a locally coherent trajectory, while the forest maintains multiple trajectories that may later be resumed. For each new input, ArborMem localizes the relevant state, restores its branch-local context, and augments it with reusable evidence retrieved across branches, preserving interaction continuity without conflating semantically related but structurally distinct trajectories. Existing long-term memory benchmarks cover diverse memory and reasoning capabilities but do not explicitly isolate branch-structured challenges. We therefore introduce BranchMemEval, a controlled diagnostic benchmark for interleaved and resumable interaction trajectories. Experiments on LongMemEval, LoCoMo, BEAM 100K, and BranchMemEval show that ArborMem outperforms the strongest baselines by 3.36 to 10.31 percentage points on the three established benchmarks and by 5.0 points on BranchMemEval. Its advantage grows under constrained read budgets, while complete memory queries remain below half a second.

36. 【2608.17516】Effects of Answer Format Variation on Gender Bias in Large Language Models

链接https://arxiv.org/abs/2608.17516

作者:Ksenia Merzlyakova,Sebastian Padó,Franziska Weeber

类目:Computation and Language (cs.CL)

关键词:large language models, answer format, social biases, biases in large, large language

备注: 6th Workshop on Computational Linguistics for the Political and Social Sciences (CPSS 2026)

点击查看摘要

Abstract:Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.

37. 【2608.17454】From Entity Mentions to Tone: An LLM-Based Pipeline for Media Bias Analysis

链接https://arxiv.org/abs/2608.17454

作者:Klesti Hoxha,Olti Qirici

类目:Computation and Language (cs.CL)

关键词:analyzing media bias, paper presents, analyzing media, pipeline groups articles, GDELT automated annotations

备注

点击查看摘要

Abstract:This paper presents a pipeline for analyzing media bias and framing in online news. The pipeline groups articles into topics and events, adds named-entity and sentiment annotations, and compares news sources through people mentions, source-level tone, and event-level coverage patterns. We apply it to 8,358 Albanian news articles collected from GDELT and compare the resulting annotations with GDELT's automated annotations. The results show moderate agreement for sentiment and entity extraction, as well as additional person-entity pairs that can potentially support the bias analysis. We compare two annotation prompts and find that stricter sentiment-validation rules remove label-score inconsistencies but increase execution time and reduce annotation coverage. Based on these results, the simpler prompt is used for the rest of the analysis. We have provided sample analysis on source-level framing pro les, person-level tone differences across sources, and event-level gatekeeping and coverage indicators. These outputs show how the same news collection can be used to examine what sources cover, how they describe public figures, and where coverage is concentrated. The approach is particularly useful in settings where manually verified datasets or specialized language tools are limited.

38. 【2608.17445】Decomposition Attacks Across Unlinkable Identities: Limits of Stateful Defenses for LLM Services

链接https://arxiv.org/abs/2608.17445

作者:Bowen Sun,Zhengyue Zhao,Xiaogeng Liu,Yinzhi Cao,Chaowei Xiao

类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)

关键词:large language model, language model services, refuse harmful tasks, large language, language model

备注

点击查看摘要

Abstract:Most large language model services use stateless defenses, which judge only the current request, to refuse harmful tasks. Decomposition attacks exploit this limitation by splitting a harmful task into individually permissible requests and combining their answers. Defending against them therefore requires a stateful monitor that considers requests together. If it can group all requests for one attacker task, it can stop the attack. However, attackers can use unlinkable identities and combine answers elsewhere, leaving no reliable grouping signal. We ask whether decomposition attacks can still be stopped under this setting. For a fixed attack strategy without retries, we prove that the achievable security and utility tradeoff depends entirely on how benign requests for the same capabilities are grouped. Persistent, recognizable groups permit a useful defense; fresh, indistinguishable groups do not. When attackers can retry and learn from Allow/Block decisions, this useful operating point disappears: the feedback reveals what passes but not whether a block was correct. Experiments on 91 executable tasks and 11,393 capability-matched benign requests support these results. Under a 1% denial cap for these requests and a 0.5% cap for unrelated background traffic, all ten tested policies, including one privileged policy with an exact request-to-operation map, either fail to stop attacks or exceed the budget. On defense-unseen task families, attack success is at least 99% after one attempt and 100% after two. Effective defenses therefore require additional evidence or mechanisms tied to grouping, such as reliable identity linkage, costs for fresh identities, or control over answer use.

39. 【2608.17399】An Investigation of Translationese in the Generations of Multilingual Large Language Models

链接https://arxiv.org/abs/2608.17399

作者:Maria Valentini,Téa Wright,Julisa Granados,Eliana Colunga,Katharina von der Wense

类目:Computation and Language (cs.CL)

关键词:translationese, Text, MLLMs, languages, translation

备注: Accepted to COLM 2026

点击查看摘要

Abstract:Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs' generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.

40. 【2608.17379】PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

链接https://arxiv.org/abs/2608.17379

作者:Genghan Zhang,Yixin Dong,Chengze Fan,Zhichen Zeng,Yueming Yuan,Shaowei Zhu,Kunle Olukotun

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:adapting large language, GPU kernel optimization, large language models, kernel optimization, architecture-specific PTX

备注

点击查看摘要

Abstract:We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.

41. 【2608.17356】ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generation

链接https://arxiv.org/abs/2608.17356

作者:Weiran Wang,Hongxiang Shi,Huitao Tang,Wenjuan Qin

类目:Computation and Language (cs.CL)

关键词:single holistic score, introduce data privacy, automated essay scoring, cost barriers, automated essay

备注

点击查看摘要

Abstract:Most automated essay scoring (AES) systems output a single holistic score without interpretable evidence and rely on closed APIs that introduce data privacy and cost barriers. We present ArguLens, an opensource, locally deployable system that decomposes AES into three decoupled components: a discourse-move classifier (Qwen2.5-7B-Instruct fine-tuned with LoRA on PERSUADE 2.0), a grade-independent LightGBM scorer over 31 linguistic and discourse features, and a label-aware feedback generator served through vLLM with a Qwen2.5-14BInstruct backbone. A Gradio web UI exposes pluggable inference backends and supports single-essay and batch scoring with downloadable per-essay breakdowns. On an essaydisjoint PERSUADE 2.0 test split, the logitprobe classifier achieves 82.6% accuracy and 0.727 macro-F1; under prompt-grouped 5-fold cross-validation the scorer reaches a mean QWK of 0.813 under an oracle discoursefeature protocol, and an ablation shows that adding gold discourse annotations yields an increment of +0.055 QWK over the lexical+syntactic configuration (paired t-test, p = 0.010). This is a component-level diagnostic rather than an end-to-end classifier-to-scorer result. The feedback generator ships with a structured evaluation protocol; its human-rater study is left to future work. The system is released under Apache 2.0 at this https URL.

42. 【2608.17330】LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

链接https://arxiv.org/abs/2608.17330

作者:Yining Hua,Cyrus Ayubcha,Hongbin Na,Levi Lian,Alon Gorenshtein,Yiftach Barash,Eyal Klang

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)

关键词:Large language models, Large language, made clear, misframed concern, medical consultation

备注: 17 pages, 3 tables. Code, cases, prompts, complete transcripts, and results: [this https URL](https://github.com/ningkko/preformulation-gap)

点击查看摘要

Abstract:Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.

43. 【2608.17325】What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?

链接https://arxiv.org/abs/2608.17325

作者:Saketh Reddy Vemula,Parameswari Krishnamurthy

类目:Computation and Language (cs.CL)

关键词:fundamental component, language modeling pipelines, modeling pipelines, language modeling, language

备注

点击查看摘要

Abstract:Tokenization is a fundamental component of language modeling pipelines. Despite its importance, it is often fixed, even though it significantly impacts model performance across languages. In this work, we analyze what tokens are learned when tokenization is jointly optimized with language modeling. We compare tokenizer-free approaches such as SSLMs and H-Nets with fixed tokenizers across 18 typologically and script-diverse languages. Our results show that joint optimization fundamentally alters token structure. SSLMs recover morphologically aligned and contextually efficient tokens, whereas H-Nets prioritize byte-level efficiency, producing longer tokens with very low overlap with standard subword vocabularies. We further show that tokenization behavior varies across language typologies. Agglutinative languages exhibit more dynamic segmentation patterns while learning. Through downstream evaluation, with pretrained-then-finetuned BERT models, we find that SSLM-based pretokenization consistently reduces language modeling perplexity and achieves competitive downstream performance despite distinct vocabularies. Overall, tokenizer-free approaches optimize for contextual and computational efficiency rather than strict morphological structure, resulting in fundamentally different yet effective vocabularies for downstream NLP.

44. 【2608.17288】Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention

链接https://arxiv.org/abs/2608.17288

作者:Emama Nahid,Tahmid Imtiaz Imu,Huayue Gu,Liran Ma,Zhipeng Cai,Honghui Xu

类目:Computation and Language (cs.CL)

关键词:measures token compatibility, compatibility through dot-product, GPT attention measures, attention measures token, measures token

备注: Preprint

点击查看摘要

Abstract:GPT attention measures token compatibility through dot-product similarity. This mechanism is simple, effective, and memory-efficient. But it does not explicitly model whether strong token features should reinforce or suppress one another. We introduce Q-Interference, a fully classical quantum-inspired attention mechanism for autoregressive language modeling that augments each query and key feature with an amplitude and a learned phase. The resulting attention score is phase-aware which aligned phases contribute constructively while conflicting phases contribute destructively. Although Q-Interference yields a richer interaction rule than similarity alone, a naive implementation of Q-Interference requires a large token-pair-feature interaction tensor, making it memory-intensive and often impractical. To address this limitation, we propose an exact trigonometric factorization that computes the same score using two standard matrix multiplications avoiding materialization of the large intermediate tensor. Q-Interference fits directly into a Transformer block in GPT and leaves the remainder of the model architecture and next-token prediction objective unchanged. Experiments on public benchmark datasets and baseline models show that the proposed reformulation trains stably in a controlled GPT-style setting and provides a consistent memory advantage over naive phase-aware interference attention. These results support the specific contribution of this work: an exact memory-efficient reformulation that makes phase-aware interference attention practical within a standard GPT pipeline.

45. 【2608.17223】mporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific MA Signal

链接https://arxiv.org/abs/2608.17223

作者:Chenhao Xue,Raslen Guesmi,Siwei Feng,Yucheng Gong,Jacob Xavier Sundram,Jordan Pang,Lan Wang,Julian Kaljuvee

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Financial-news direction prediction, reported gains depend, gains depend critically, Financial-news direction, popular NLP benchmark

备注

点击查看摘要

Abstract:Financial-news direction prediction has become a popular NLP benchmark, yet reported gains depend critically on whether the train-test split is chronological or random, i.e., on temporal leakage. We audit this dependence on a 49,799-article corpus across 16 feature-model combinations spanning TF-IDF, MiniLM, FinBERT, and fine-tuned RoBERTa-large / DeBERTa-v3-large, plus separate zero/few-shot and LoRA probes of Llama-3 and Qwen2.5 LLMs: random splits inflate MCC by $1.1\times$ to $6.5\times$, tracking model capacity and feature richness, and end-to-end FinBERT fine-tuning re-amplifies rather than closes the gap (size-matched ratio $1.75\times$). Conditioning on event type, mergers and acquisitions (MA) is the only audited category with a positive locked-test signal under near-temporal chronological evaluation (TF-IDF MCC $= 0.138$ train-only, $0.068$ under train$\cup$val refit; 10,000-permutation $p 10^{-3}$); the signal does not transfer to FNSPID's 2009-2020 U.S. corpus, localising the headline to our 2024-2025 European-tilted MA semantics rather than a universal predictor. Three independent role labellers converge on acquirer-tagged articles as the signal locus, a power-limited qualitative convergence rather than a hypothesis-tested asymmetry. Chronological splitting plays for financial NLP the role characteristics-purging plays for asset pricing: it strips the predictable, stale component of news and leaves a residual that is small, event-localized, and lexically shallow. We advocate leakage audits as a required disclosure for financial-NLP benchmarks.

46. 【2608.17218】he Plot Thins: Uniformity and Linearity in Literary Summaries

链接https://arxiv.org/abs/2608.17218

作者:Rebecca M. M. Hicke,Sil Hamilton,David Mimno,Ross Deans Kristensen-McLachlan

类目:Computation and Language (cs.CL)

关键词:artistic expression, literature prioritize plot, balance plot, plot, literature are complicated

备注

点击查看摘要

Abstract:Works of literature are complicated; they balance plot, suspense, surprise, and artistic expression. Summaries of literature prioritize plot, and therefore may deviate from their sources. Using a combination of manual and LLM-based annotation, we construct a dataset mapping sentences from 150 novel summaries to their respective source chapters. We find the task unexpectedly difficult for both human and model annotators. Using the sentence-to-chapter mappings, we then measure summary linearity, the degree to which it maintains the source's order of events, and uniformity, the degree to which a summary spreads attention equally across a source. By examining when and how summaries break linearity and uniformity, we identify differences in how literary works and summaries express plot, particularly with regard to the clarity and prominence with which narrative details are described.

47. 【2608.17205】Which Source Wins? Task-Dependent Reliance in Vision-Language Models

链接https://arxiv.org/abs/2608.17205

作者:Rodela Ghosh,Aviral Gupta,Guangjing Wang

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision-language models, harder to read, Vision-language, combine images, text

备注: 20 pages. Under review

点击查看摘要

Abstract:Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the text across four levels of legibility while keeping the other clean, and track how the model's preference changes. We build conflicts from GSM8K and SVAMP by pairing the rendered image of one arithmetic problem with the text of another, so the two sources support different answers. We also introduce ChartQA-Conflict, a manually reviewed benchmark of 229 chart-report conflicts with matched chart and table-image representations. We evaluate six open-weight VLMs using both generated answers and a length-normalized conditional log-likelihood margin. On GSM8K and SVAMP, five of six models shift more strongly away from degraded text than from degraded images. On ChartQA-Conflict, all six likelihood-scored models exhibit the opposite pattern, shifting more strongly away from the degraded visual source. This reversal persists after calibrating for unimodal accuracy loss and after replacing charts with plain table images. Two frontier API models, GPT-5.6-Luna and Gemini-3.5-Flash, behaviorally replicate the ChartQA-Conflict reversal, with GPT-5.6-Luna also matching the arithmetic direction. These results show that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings. The source code is available at this https URL.

48. 【2608.17188】oken Optimization and Context Window Management in Multi-Agent AI Workflows

链接https://arxiv.org/abs/2608.17188

作者:Dvir Shamay

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:context-window quality, token cost, Multi-agent, Multi-agent AI workflows, quality

备注: 29 pages (main paper + technical appendix), 3 figures. Also archived on Zenodo: [https://doi.org/10.5281/zenodo.21924612](https://doi.org/10.5281/zenodo.21924612)

点击查看摘要

Abstract:Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents a practitioner framework for token optimization and context-window management, grounded in an internal production dashboard that extracts structured work items from meetings, email, and chat with LLMs and routes summaries across workstreams. Six patterns are described: context stratification, fetch-once/process-locally architecture, schema-contracted prompts, token-aware fallback chains, semantic caching, and inter-agent communication compression. In production they cut measured cold-load latency to 61-116 seconds (six timed runs) from an operational baseline of roughly 3.5-10.5 minutes, with an estimated 60-70% token reduction. It also reports a controlled context-composition study: 2,420 confirmatory trials across 11 model configurations, using 661 anonymized workplace items scored for relevance. Holding the prompt at a fixed ten items, replacing some high-relevance items with same-domain low-relevance items improves the model's relevance-score concordance on the target items, versus high-relevance items only; we call this relevance-contrast context. In the all-11 paired analysis, the 50:50 signal/noise condition improved relevance accuracy by +0.077 over the 100% condition (naive 95% CI [+0.056, +0.098], Cohen's d = 0.49, Holm-adjusted p .001, n = 220). These cells are not independent; by the nine model families the effect is +0.084 (95% interval [+0.064, +0.103]), reported as a within-corpus descriptive comparison, not a population inference. A Fusion-of-N follow-up found that learned synthesis did not beat the mechanical set union of item IDs. The contribution is a measured engineering layer between model research and production agent practice: repeatable patterns and evaluation methods for faster, cheaper, more reliable workflows.

49. 【2608.17184】AISA: AI Safety Assistant Framework for Continuous Improvement of Highway Construction

链接https://arxiv.org/abs/2608.17184

作者:Mason Smetana,Trevor Neece,Lev Khazanovich

类目:Computation and Language (cs.CL)

关键词:Job Safety Analysis, Job Safety, Safety Analysis, benefit from prior, stored as unstructured

备注: 17 pages, 5 figures

点击查看摘要

Abstract:Job Safety Analysis (JSA) and pre-task planning can benefit from prior incident records, yet historical accident data is often stored as unstructured narratives that are difficult to consult at the point of planning. A novel framework centered on large language models (LLMs) for highway construction safety reporting and planning is proposed as a foundation for future agentic applications, prioritizing deterministic, local inferencing. The first aim is to enable classification and quality scoring of incident narratives for existing and future reporting purposes. The second is to evaluate retrieval of relevant historical accidents, related imagery, and trusted industry documents for incorporation into daily safety plans. Neural probes were trained to classify incidents along four multiclass and two binary Occupational Injury and Illness Classification System (OIICS) fields and to derive an overall quality score, evaluated on a test set of over 15,000 narratives and a held-out set of 100 author-labeled records, benchmarked against a majority-vote LLM ensemble. The retrieval of historical accidents, reference imagery, and industry documents was benchmarked across embedding models using standard information retrieval metrics. OIICS classification reached 75% held-out accuracy, though the two binary flags were degenerate. The quality score, while meaningful on one database, was distorted on out-of-distribution fatalities in the held-out dataset. Accident retrieval recovered relevant incidents far above chance, performing best on lexically distinct construction activities. On document question answering, an open-weight decoder embedding model surpassed proprietary models. Overall, this work provides a new framework rooted in local inferencing and text embedding models for future agentic applications, with emphasis on bridging external data to JSA reports.

50. 【2608.17171】Polaris: Learning to Generate Table Descriptions from Retrieval Feedback

链接https://arxiv.org/abs/2608.17171

作者:Ting Cai,Tuan Minh Phan,AnHai Doan

类目:Computation and Language (cs.CL); Databases (cs.DB)

关键词:table-centric NLP tasks, table-centric NLP, retrieve relevant tables, NLP tasks, keyword search

备注: 22 pages, 6 figures

点击查看摘要

Abstract:Many table-centric NLP tasks such as NL2SQL first retrieve relevant tables from large collections using keyword search. Recent work uses LLMs to generate natural-language table descriptions to improve retrieval, but they are typically optimized for fluency rather than retrieval effectiveness. We present Polaris, a system that trains an LLM to generate table descriptions directly from retrieval feedback. Our key insight is that existing table retrieval benchmarks already contain the supervision needed for this task: given query-table relevance judgments, we generate multiple candidate descriptions for each table, rank them by their BM25 retrieval effectiveness, and use the resulting preference pairs to fine-tune the LLM with Direct Preference Optimization (DPO). Polaris further expands abbreviated table and column names before generation to reduce vocabulary mismatch. Extensive experiments show that Polaris outperforms the state-of-the-art AutoDDG solution, often by a significant margin. More broadly, our results demonstrate that retrieval benchmarks can be repurposed as supervision for training LLMs to generate retrieval-oriented metadata.

51. 【2608.17168】Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

链接https://arxiv.org/abs/2608.17168

作者:Amogh Raina,Ilias Chalkidis,Daniel Hershcovich,Henrik Palmer Olsen

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:legal case forecasting, case forecasting, demanding legal-oriented tasks, remain under explored, legal case

备注: 24 pages, 4 figures, 4 tables, Submitted to AI4LAW Workshop at ICML 2026

点击查看摘要

Abstract:Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.

52. 【2608.17153】owards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

链接https://arxiv.org/abs/2608.17153

作者:Mehrdad Ghassabi

类目:Computation and Language (cs.CL)

关键词:Retrieval-Augmented Generation, systems remain vulnerable, knowledge-poisoning attacks, significantly enhanced, enhanced the performance

备注

点击查看摘要

Abstract:Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it. Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsible for final answer synthesis from directly accessing raw evidence. Although effective, this strict isolation can introduce substantial computational overhead. In this work, we propose a refined security principle: only agents capable of deliberative System 2 reasoning may access untrusted documents. To evaluate this principle, we introduce novel metrics that quantify the discrepancy between misinformation detection and downstream influence. We then empirically compare state-of-the-art reasoning language models with standard language models across these metrics. Our results show that reasoning-capable models are substantially more robust to corrupted evidence, without requiring the strict isolation imposed by the Cordon Principle. These findings provide empirical support for our refined principle and suggest a more practical foundation for secure RAG system design.

53. 【2608.17150】KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

链接https://arxiv.org/abs/2608.17150

作者:Yoonjoo Lee,Hyoungwook Jin,Tae Soo Kim,Shaoyang Zhang,Philippe Laban,Q. Vera Liao

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)

关键词:Large Language Models, Large Language, perform information calibration, user evolving understanding, Language Models

备注: 30 pages, 6 figures, 16 tables

点击查看摘要

Abstract:To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.

54. 【2608.17120】Children, but not language models, show accelerating returns in word learning

链接https://arxiv.org/abs/2608.17120

作者:Michael C. Frank

类目:Computation and Language (cs.CL)

关键词:picks up speed, hundreds of words, begins slowly, slowly but quickly, quickly picks

备注

点击查看摘要

Abstract:Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models -- even those trained on child-directed speech -- do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.

55. 【2608.17102】Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

链接https://arxiv.org/abs/2608.17102

作者:Xiutian Zhao,Luqi Sun,Björn Schuller,Berrak Sisman

类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS); Image and Video Processing (eess.IV)

关键词:Modern multimodal foundation, multimodal foundation models, made rapid progress, tasks requiring integrated, requiring integrated perception

备注: 9 pages, 4 figures

点击查看摘要

Abstract:Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways. We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces. Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.

56. 【2608.17096】A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not

链接https://arxiv.org/abs/2608.17096

作者:Liudmila Rozanova,Alexander Temerev

类目:Computation and Language (cs.CL)

关键词:unstated assumptions, Voynich manuscript, Beinecke, Voynich, tokens

备注: 33 pages, 7 figures, 3 appendices. Analysis code and data are included as ancillary files and mirrored at [this https URL](https://github.com/lrozanova/voynich-units)

点击查看摘要

Abstract:The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls and quire-level resampling. None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of tokens themselves. Glyph regularity is too strong for one-to-one substitution of any tested plaintext (conditional entropy 2.7 bits against about 3.5 for Latin, Italian, and English) and resolves instead onto a quire-stable scale of recurrent multi-symbol units. Tokens form a plausible vocabulary, yet the identity of one token predicts the next by under 1% of token entropy, below every matched control (2-10%), while the glyphs at token edges share 0.2 bits of mutual information, more than in any prose control. Blanks fall into two regimes: the separators transcribers marked uncertain behave like word-internal junctures, are physically narrower on the page (AUC 0.905 from independent image coordinates, with the same sign in a small blind ink audit), and are crossed by learned units even when every space is erased before learning. This profile is also what discriminates. A published Voynich-imitating cipher and a self-citation text generator both reproduce the low entropy, the unit scale, the weak token order, and the null result of a calibrated substitution attack; neither reproduces the edge-glyph coupling or the open, hapax-rich vocabulary (70% singleton types against 41% and 59-60%). Any account of the manuscript must therefore earn, rather than assume, the step from glyphs, tokens, and separators to letters, words, and word spaces, and these are the measurements on which to do so.

57. 【2608.17088】here is No Theoretical Curse of Multilinguality For Embedding Space Structure

链接https://arxiv.org/abs/2608.17088

作者:Niyati Bafna,Neha Verma,Vilém Zouhar,Philipp Koehn,David Yarowsky

类目:Computation and Language (cs.CL)

关键词:achieve high monolingual, high monolingual performance, large-scale language coverage, multilingual NLP, multilingual model performance

备注

点击查看摘要

Abstract:A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model performance as we increase language coverage, posing a threat to the above goal. This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required capacity. We first formalize the goal of "perfect multilinguality", embodied in two multilinguality conditions. We then prove that the minimum dimensionality required for perfect multilinguality grows only logarithmically in the number of languages. That is, we show that there is no theoretical curse of multilinguality for embedding space structure. This suggests that the empirical curse of multilinguality is a result of real world data and training conditions. We back this understanding with a small-scale empirical study. Our paper provides the first theoretical and intrinsic perspective on the curse of multilinguality, with implications for the scientific understanding of this phenomenon.

58. 【2608.17084】Uncertainty-Aware Decision Making in Multimodal Large Language Models

链接https://arxiv.org/abs/2608.17084

作者:Abderrahmene Boudiaf,Irfan Hussain,Sajid Javed

类目:Computation and Language (cs.CL)

关键词:large language models, Multimodal large language, increasingly answer questions, language models, depends on visual

备注

点击查看摘要

Abstract:Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Their failures are therefore not only linguistic. A fluent answer may conceal poor input quality, a perceptual error, weak grounding, conflict between modalities, unstable reasoning, distribution shift, or a question that is not answerable from the supplied evidence. This survey organizes the literature on uncertainty-aware MLLMs around a decision-centered framework: uncertainty sources give rise to observable signals, signals must be calibrated or controlled for risk, and calibrated uncertainty should determine the system action. We review work on token and logit uncertainty, semantic disagreement, perturbation instability, grounding and attribution scores, verbalized confidence, verifier and judge scores, conformal prediction, selective answering, abstention, clarification, retrieval, self-checking, and escalation. The central argument is that uncertainty should not be evaluated only as a confidence number; it should be evaluated by whether it improves behavior under insufficient, conflicting, shifted, or high-risk multimodal evidence. We position this survey against text-only uncertainty and abstention surveys, broad MLLM surveys, MLLM hallucination surveys, and safety-oriented reviews. We conclude with open problems in source-aware decomposition, action-aware benchmarks, calibration under shift, black-box uncertainty estimation, broader modality coverage, reproducible reporting, and human-centered uncertainty communication.

59. 【2608.17075】Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

链接https://arxiv.org/abs/2608.17075

作者:Junda Wang,Meysam Ghaffari,Akshat Choube,Mohsen Sharifi Renani,Hong Yu,Carlos Morato

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Next-encounter ICD forecasting, ICD forecasting predicts, standardized diagnosis codes, Next-encounter ICD, forecasting predicts

备注

点击查看摘要

Abstract:Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct. Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses. We introduce ICD-Deepresearch, a DeepResearch workflow that composes these predictive foundation models with medical search and ICD dictionaries. Because no source reveals the future code set, research evaluates candidate transitions by linking patient evidence, external clinical relations, and exact code semantics under a fixed top-K budget. Candidate Generation uses SparseEHR to produce an EHR Prior that initializes two bounded Research Expansion rounds; an independent GPT-5 Direct Forecast supplies complementary candidates. Final Selection validates, deduplicates, and jointly ranks both paths, after which a separate module writes rationales without changing predictions. Finally ICD-Deepresearch achieves patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV. Physicians rate 51% and 68% of its retrieved documents useful, compared with 22% and 39% for standalone GPT-5 web search and 32% and 41% for Medical Deep Research. ICD-Deepresearch therefore improves over the registered local comparators while retrieving evidence with higher physician-rated usefulness than the standalone research systems

60. 【2608.17063】J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

链接https://arxiv.org/abs/2608.17063

作者:Yunfan Gao,Xinyi Huang,Tao Sheng,Haorui Song,Yun Xiong,Haofen Wang

类目:Machine Learning (cs.LG); Computation and Language (cs.CL)

关键词:Large language models, make complex judgments, Large language, language models, complex judgments

备注: 19 pages, 12 figures, and 13 tables; includes appendices

点击查看摘要

Abstract:Large language models can be fine-tuned into specialized classifiers that perform well across diverse text tasks and make complex judgments, but they typically expose only final labels, leaving the decision knowledge acquired through fine-tuning implicit within the model. We study how to mine this internal decision knowledge from a fine-tuned classifier and encode it in an executable representation that can be inspected, validated, and reused beyond the source classifier. We introduce J-Miner, which mines text-level named concepts by aggregating vocabulary-aligned internal signals across layers and token positions, and uses the classifier's own predictions to learn executable decision rules over them. This process distills local internal readouts into an explicit classifier-level knowledge representation. Across multiple classification tasks, J-Miner rules reproduce up to 98.3\% of source-classifier decisions and achieve 6.0--29.5 percentage points higher behavioral fidelity than equally compact rules learned from input words. Further analysis shows that the named concepts reflect internal semantic evidence associated with task decisions, while the learned rules consolidate these distributed signals into inspectable decision structures. The resulting decision knowledge also transfers to lightweight standalone students: using about 1/24 as many parameters as the source classifiers, they reconstruct and execute the representation from raw text while retaining 99.8\% of the source classifiers' mean task accuracy. These findings show that task-specific decision knowledge can be faithfully represented in an explicit, executable form and reused beyond the classifier in which it was learned.

61. 【2608.17053】Memory Is Communication: The Frontier Between Remembering and Signaling

链接https://arxiv.org/abs/2608.17053

作者:Yashar Talebirad,Eden Redman,Ali Parsaee,Osmar R. Zaiane

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Theory (cs.IT); Multiagent Systems (cs.MA)

关键词:bounded agent, history, Abstract, agent, peer

备注

点击查看摘要

Abstract:A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under limits on both resources, how should an agent allocate its information budget? Given a fixed task and decision rule, the memory and message rate pairs attaining a performance threshold form an achievable region under specified rules for using history and peer observations. We call its efficient boundary the remembering--signaling frontier. Across conditions where history permits the same maximum reduction in task loss, we hypothesize that a bounded agent will need less peer communication when it obtains a larger loss reduction from history. In preliminary referential games, target repetition coincided with shorter successful messages, while predictability from a hidden cyclic rule did not shorten them. Experiments varying memory and message rates can estimate the frontier and test this prediction across cooperative tasks.

62. 【2608.17051】Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

链接https://arxiv.org/abs/2608.17051

作者:Daniel Palacios,Matthew Brady Neeley,Angel Adetomike Otto,Shalini Dhamodharan,John P. Woodhouse,Chi-fan Lin,Mark Zobeck,Zhandong Liu,Hyun-Hwan Jeong

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:protected health information, electronic health records, health records requires, Toggle, Texas Children Hospital

备注

点击查看摘要

Abstract:Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.

Subjects:

Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

Cite as:
arXiv:2608.17051 [cs.CL]

(or
arXiv:2608.17051v1 [cs.CL] for this version)

https://doi.org/10.48550/arXiv.2608.17051

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Hyun-Hwan Jeong [view email] [v1]
Mon, 17 Aug 2026 18:56:04 UTC (1,381 KB)

Full-text links:
Access Paper:

View a PDF of the paper titled Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss, by Daniel Palacios and Matthew Brady Neeley and Angel Adetomike Otto and Shalini Dhamodharan and John P. Woodhouse and Chi-fan Lin and Mark Zobeck and Zhandong Liu and Hyun-Hwan JeongView PDFHTML (experimental)TeX Source

view license

Current browse context:
cs.CL

prev

|
next

new
|
recent
| 2026-08

Change to browse by:

cs
cs.AI

References Citations

NASA ADSGoogle Scholar
Semantic Scholar

export BibTeX citation
Loading…

BibTeX formatted citation

loading…

Data provided by:

Bookmark

checked="checked"class=“labs-tab-input”>
Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

Author
Venue
Institution
Topic

    About arXivLabs

arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.

Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)

mathjaxToggle();

    We gratefully acknowledge support from
    our major funders,
    member institutions, ,
    and all contributors.

About

Help

Contact

Subscribe

Copyright

Privacy

Accessibility

Operational Status (opens in new tab)

Major funding support from

63. 【2608.17050】Cross-Model Memory Transfer via Target-Side Reader Adaptation

链接https://arxiv.org/abs/2608.17050

作者:Mingyuan Li,Guangsheng Yu,Xu Wang,Shaoxiong Ji

类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)

关键词:Methods for improving, large language models, language models typically, models typically fall, large language

备注

点击查看摘要

Abstract:Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.

64. 【2608.16975】Margin-Regularized Structured Semantic Alignment for Brain-Language Correspondence

链接https://arxiv.org/abs/2608.16975

作者:Jiaqi Wang,Huawen Hu,Shu Zhang

类目:Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:achieved remarkable progress, remarkable progress, large language models, rapid advancement, advancement of large

备注

点击查看摘要

Abstract:With the rapid advancement of large language models, brain-language decoding has achieved remarkable progress. However, it remains unclear whether decoded content genuinely reflects neural representations or is largely reconstructed by the language model itself. This ambiguity limits interpretability and hinders the investigation of intrinsic brain-language correspondence. To address this challenge, we propose MD-SigLIP. This margin-regularized structured semantic alignment framework directly aligns brain embeddings with text embeddings in a shared semantic space, enabling retrieval-based decoding. This formulation enables explicit modeling of the correspondence between neural representations and language semantics. Building upon duplicate-aware sigmoid contrastive learning, we introduce a listwise margin-regularized term that enforces structured ranking constraints between positive semantic clusters and negative samples. By modeling multi-positive semantic structure and margin-based ordering simultaneously, the method captures the manifold organization of language embeddings reflected in neural signals. Experiments demonstrate state-of-the-art retrieval performance under both full-vocabulary and subset evaluation settings.

65. 【2608.16956】he Price of Thinking: Reasoning Effort as a Model-Specific API Contract

链接https://arxiv.org/abs/2608.16956

作者:Yeabin Moon

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)

关键词:API buyers purchase, API buyers, output rail, service product, reasoning-effort term

备注: 15 pages, 3 figures, 2 tables

点击查看摘要

Abstract:API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. Mean delivered cost was \$0.01031 per call higher under the explicit-high contract than under the omitted contract [+\$0.00204, +\$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. Cost per correct answer was \$0.08665 under the high-effort contract and \$0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied.

66. 【2608.16934】SeqFeed: Improving Agentic RTL Code Generation with Sequential Behavior Feedback

链接https://arxiv.org/abs/2608.16934

作者:Yuxin Du,Juxin Niu,Tao Hu,Xi Wang,Zhe Jiang,Nan Guan

类目:Hardware Architecture (cs.AR); Computation and Language (cs.CL)

关键词:agentic systems offers, RTL code generation, RTL code, correct RTL code, hardware design

备注

点击查看摘要

Abstract:RTL code generation is a critical stage in hardware design, and the emergence of agentic systems offers new opportunities to automate this process. To generate correct RTL code, agents must understand sequential behavior, including how signals evolve and propagate over multiple clock cycles. However, effectively conveying such temporal information to agents remains a significant challenge. RTL code does not expose cycle-level signal behavior for a specific execution, whereas full simulation waveforms are too voluminous and noisy for effective LLM analysis. To address these limitations, we study how human engineers reason about sequential behavior and identify three requirements for effective feedback: it should be event-addressable, dependency-traceable, and iteratively-queryable. Guided by these requirements, we propose \textit{SeqFeed}, which comprises two complementary mechanisms: (1) \textit{SeQuery}, an SQL-like waveform query language that enables agents to anchor queries to semantic events and sample signal values at relative time points; and (2) \textit{SeGraph}, a dependency graph that tracks signal propagation across clock cycles. Experimental results across multiple LLMs demonstrate the effectiveness of SeqFeed in improving pass rates. SeQuery and SeGraph are each effective independently and provide complementary benefits when used together.

67. 【2608.16909】When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice

链接https://arxiv.org/abs/2608.16909

作者:Muhammad Salar Khan,Hamza Umer,Hasan Mahmud,Sandra Rothenberg

类目:Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:bias remains underexamined, Large language models, financial advisory systems, Large language, advisory systems

备注: 50 pages

点击查看摘要

Abstract:Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religious framing, and non-religious clients often received advisor-centered religious appeals. Qualitative findings show that bias is linguistically manifested through religious anchoring, uneven cultural signaling, and tone modulation, varying by model and financial scenario. Stock investment prompts produced more financially technical responses, whereas life insurance advice triggered stronger religious language. The study develops a dual-dimensional framework linking structural bias rooted in model training and design with discursive bias expressed through language, advancing understanding of algorithmic bias in LLM-generated financial advice. It also shows that such advice adapts linguistically to identity cues, revealing a managerial dilemma between personalization and neutrality. Finally, it highlights implications for businesses, financial institutions, and regulators seeking to ensure neutrality, cultural sensitivity, and trust in AI-mediated advice.

68. 【2608.16905】he politics of postmortem privacy

链接https://arxiv.org/abs/2608.16905

作者:Mauricio Figueroa

类目:Computers and Society (cs.CY); Computation and Language (cs.CL); Social and Information Networks (cs.SI)

关键词:postmortem privacy, postmortem privacy evokes, justificatory foundations, increasingly acknowledged, digital spaces

备注

点击查看摘要

Abstract:While the existence of postmortem privacy is increasingly acknowledged (such as the protection of the presence of deceased within digital spaces), far less attention has been paid to its internal instability: its scope (the extent of its application), justificatory foundations (why do we protect the deceased in the first place), and uneven articulation across jurisdictions (for example, some jurisdictions may tolerate or endorse practices that may be contestable in a different jurisdiction). This piece unearths the internal diversity of the concept by illuminating specific points of tension and conflict that the notion of postmortem privacy evokes. These points of tension are collectively refer to as the politics of postmortem privacy. To do so, this paper organises existing contributions of legal scholarship, placing them in dialogue with broader cultural, social, historical and political observations to illustrate the politics of postmortem privacy through three different loci of analysis: the transatlantic divide between European and American approaches, intra-European tensions within data protection governance, and postcolonial and post-authoritarian contexts in the Global South. While existing literature has glimpsed toward the former two, this piece contends that the latter deserves greater attention and inclusion in the debates around privacy and the dead. The piece explains, in continuity with existing scholarship, how postmortem privacy is assembled differently as a productive register through which societies negotiate memory and dignity, which play a great role in the governance of data of the dead and information flows.

69. 【2608.16894】An Investigation of the NeurIPS and ICML 2025 Position Tracks

链接https://arxiv.org/abs/2608.16894

作者:Fan Yang,Wenkai Li,Jun Liu

类目:Computers and Society (cs.CY); Computation and Language (cs.CL)

关键词:ICML Position Paper, Position Paper Tracks, count as rigorous, venues shape, research claims

备注

点击查看摘要

Abstract:ML venues shape what kinds of research claims become legible to reviewers and what forms of evidence count as rigorous. The NeurIPS and ICML Position Paper Tracks were created for agenda-setting work, making their early composition worth auditing. \textbf{This paper argues that the publicly accessible 2025 reviewed pool is dominated by reformist critique, and that the track should explicitly solicit direction-setting work alongside, not in place of, the reformist critiques it already hosts well.} We audit every accessible submission to the NeurIPS 2025 and ICML 2025 Position Tracks under a pre-specified rubric, and compare the resulting pattern with a reference class of widely recognized agenda-shifting ML papers. Three-quarters of audited submissions critique an existing benchmark, evaluation, or methodology; these papers score highly on our artifact-coupling rubric, but evidentiary depth does not predict reviewer rating. The reference class (AlexNet, the Transformer, Concrete Problems in AI Safety, and others) differs from the accessible reviewed pool in \emph{artifact kind}: agenda-shifting papers typically gave the field something new to build on, test against, or contest, such as a measurement protocol, benchmark proposal, toy implementation, dataset card, audit template, or falsifiable experimental program. We close with four CFP-level interventions aimed at broadening the submission mix without displacing the critiques the track already hosts well.

70. 【2608.15382】Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

链接https://arxiv.org/abs/2608.15382

作者:Ummara Mumtaz,Aimen Noor,Awais Ahmed

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Quantitative Methods (q-bio.QM)

关键词:Large language models, healthcare decision support, Large language, reward single-answer accuracy, decision support

备注

点击查看摘要

Abstract:Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.

71. 【2602.14784】Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

链接https://arxiv.org/abs/2602.14784

作者:Christos Koutsiaris

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Breaking long, Breaking long documents, smaller segments, fundamental challenge, Breaking

备注: 8 pages, 4 figures. Code available at [this https URL](https://github.com/unseen1980/IDC)

点击查看摘要

Abstract:Breaking long documents into smaller segments is a fundamental challenge in information retrieval. Whether for search engines, question-answering systems, or retrieval-augmented generation (RAG), effective segmentation determines how well systems can locate and return relevant information. However, traditional methods, such as fixed-length or coherence-based segmentation, ignore user intent, leading to chunks that split answers or contain irrelevant noise. We introduce Intent-Driven Dynamic Chunking (IDC), a novel approach that uses predicted user queries to guide document segmentation. IDC leverages a Large Language Model to generate likely user intents for a document and then employs a dynamic programming algorithm to find the globally optimal chunk boundaries. This represents a novel application of DP to intent-aware segmentation that avoids greedy pitfalls. We evaluated IDC on six diverse question-answering datasets, including news articles, Wikipedia, academic papers, and technical documentation. IDC outperformed traditional chunking strategies on five datasets, improving top-1 retrieval accuracy by 5% to 67%, and matched the best baseline on the sixth. Additionally, IDC produced 40-60% fewer chunks than baseline methods while achieving 93-100% answer coverage. These results demonstrate that aligning document structure with anticipated information needs significantly boosts retrieval performance, particularly for long and heterogeneous documents.

72. 【2311.06273】Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis

链接https://arxiv.org/abs/2311.06273

作者:Ummara Mumtaz,Summaya Mumtaz

类目:atistical Finance (q-fin.ST); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)

关键词:exceptional conversational skills, grasp of language, brought a notable, notable shift, exceptional conversational

备注: total 11 pages including references, 4 figures and one table

点击查看摘要

Abstract:The rise of ChatGPT has brought a notable shift to the AI sector, with its exceptional conversational skills and deep grasp of language. Recognizing its value across different areas, our study investigates ChatGPT's capacity to predict stock market movements using only social media tweets and sentiment analysis. We aim to see if ChatGPT can tap into the vast sentiment data on platforms like Twitter to offer insightful predictions about stock trends. We focus on determining if a tweet has a positive, negative, or neutral effect on two big tech giants Microsoft and Google's stock value. Our findings highlight a positive link between ChatGPT's evaluations and the following days stock results for both tech companies. This research enriches our view on ChatGPT's adaptability and emphasizes the growing importance of AI in shaping financial market forecasts.

信息检索

1. 【2608.17889】VisDocAgentBench: Benchmarking Agents for Visually Rich Document Retrieval

链接https://arxiv.org/abs/2608.17889

作者:Lexiang Hu,Yanzhao Zhang,Mingxin Li,Dingkun Long,Yikang Li,Fuwei Zhang,Yisen Wang,Zhouchen Lin

类目:Information Retrieval (cs.IR)

关键词:Visually rich documents, Visually rich, rich documents encode, structured visual elements, documents encode relevance

备注

点击查看摘要

Abstract:Visually rich documents encode relevance through language, layout, structured visual elements, and corpus context, yet retrieval is typically evaluated by one-shot query--page matching. Agentic-search benchmarks usually score downstream question answering or report generation, leaving document ranking under iterative evidence acquisition underexplored. We introduce VisDocAgentBench, a closed-corpus benchmark comparing static and agentic retrieval under a shared ranked-output contract. It contains 2,375 pages from 100 documents and 120 unique-target queries balanced across direct, one-bridge, and two-bridge evidence structures. Relation-preserving construction yields semantic, relational, and visual queries, followed by full-document review and hard-negative validation. A strong late-interaction visual retriever reaches 97.50% Recall@1 on direct items but 2.50% on two-bridge items, exposing the limits of query--target matching when relevance depends on corpus context. Agents recover much of this loss, but planner choice and retrieval representation remain decisive. Every planner performs better with visual retrieval, whose best R@1 reaches 67.50% versus 37.50% for OCR-text. Ablations identify iterative search and page inspection as consequential capabilities, and providing the complete support context improves ranking on both routes. Trace analysis localizes the remaining losses to target discovery, candidate examination, and evidence-role integration. These findings motivate retrieval agents that combine modality-preserving discovery with evidence-directed verification.

2. 【2608.17632】DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

链接https://arxiv.org/abs/2608.17632

作者:Jingyuan Wang,Richong Zhang,Zhijie Nie,Mingxin Li,Yanzhao Zhang

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Large language models, expand underspecified queries, Large language, dense representations, expand underspecified

备注

点击查看摘要

Abstract:Large language models (LLMs) can both expand underspecified queries and encode text as dense representations, suggesting a unified model for query expansion and retrieval. Existing systems usually rely on prompted expansions, independently trained modules, or staged optimization, leaving generated expansions only indirectly aligned with the retrieval loss that judges them. We train a single decoder-only LLM end to end, where the same model generates the expansion and encodes both the expanded query and candidate documents. This unified setting creates a moving-target problem: retrieval supervision should improve query-side expansion, but the same update also shifts the document embeddings that serve as retrieval targets. We introduce Document Embedding Preservation Tuning (DEPT), which keeps tuned document embeddings close to cached initial embeddings while allowing retrieval gradients to pass through straight-through decoding into the generator. DEPT converts joint query--document movement into query-side adaptation against approximately stable, whitened document embeddings that support index reuse and online hard-negative mining. Experiments with Qwen3-4B-Instruct-2507 and LLaMA-3.2-3B-Instruct on five datasets in BEIR benchmark show that DEPT improves average retrieval quality over training-free, independently trained, and staged unified baselines, while ablations isolate the effects of preservation, whitening, end-to-end expansion training, and online negatives. Code is available at this https URL.

3. 【2608.17618】From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support

链接https://arxiv.org/abs/2608.17618

作者:Ngoc Luyen Le,Marie-Hélène Abel,Bertrand Laforge

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Learning analytics models, Learning analytics, risk of poor, Learning, poor performance

备注

点击查看摘要

Abstract:Learning analytics models can identify students at risk of poor performance, but they do not directly indicate which interventions are feasible, actionable, and compatible with educational constraints. This paper introduces SC2R, a semantics-constrained counterfactual recourse framework for educational decision support. SC2R combines a calibrated predictive model, integer-programming-based recourse generation over discrete action variables, a lightweight RDF vocabulary for intervention-plan representation, and SHACL validation for enforcing timing, budget, immutability, and availability constraints. The framework is evaluated offline on the OULAD dataset using snapshots constructed relative to each assessment at two decision horizons. Results show that the predictive component provides strong performance, that compact intervention plans can be generated at scale, and that semantic validation reveals infeasible plans that lighter optimization-only settings would otherwise accept. Rather than claiming causal improvement in student outcomes, this work shows that counterfactual recourse becomes more operationally meaningful in education when recommendations are not only model-valid, but also semantically feasible and machine-checkable.

4. 【2608.17613】Once Generated, Ranked: End-to-End Generative Slate Recommendation with Unified Semantic-Collaborative IDs

链接https://arxiv.org/abs/2608.17613

作者:Yang Hu,Jiayi Guo,Jingui Ma,Ning Li,Jiangling Qin,Yanming Li,Yang Deng,Xiaoshuang Chen,Kaiqiao Zhan

类目:Information Retrieval (cs.IR); Social and Information Networks (cs.SI)

关键词:requiring joint optimization, individual item, item interactions, requiring joint, Slate recommendation treats

备注: 18 pages, 3 figures

点击查看摘要

Abstract:Slate recommendation treats a slate rather than an individual item as the recommendation unit, requiring joint optimization of item interactions and slate utility. Existing approaches typically separate candidate generation from ranking and restrict optimization to retrieved candidates. Generative recommendation with Semantic IDs (SIDs) offers a path to end-to-end recommendation, but existing SID construction often lacks recommendation-aware semantics and effective local collaborative signals, while next-token prediction is misaligned with slate-level objectives. We propose OGR, an end-to-end framework that directly generates ordered slates-"Once Generated, Ranked." OGR first introduces TUSID, which adaptively fuses item-specific semantic and local collaborative information into hierarchical SIDs. It then uses list-wise preference planning and pipelined position-wise SID decoding to model global preferences and inter-item dependencies while generating ordered slates. We further propose SPA, a reward-guided conservative policy optimization method that aligns generated slates with user preferences beyond likelihood imitation. Offline experiments show that OGR outperforms representative baselines, with 48.2% and 27.2% relative NDCG@5 gains on industrial and public datasets, respectively. Online A/B testing on Kuaishou further yields a 1.120% improvement in Effective Views.

5. 【2608.17316】Empowering Compact LLMs with Fusion of Layer-wise Exits for Recommendation

链接https://arxiv.org/abs/2608.17316

作者:Xurong Liang,Tong Chen,Quoc Viet Hung Nguyen,Jianxin Li,Xiangliang Zhang,Hongzhi Yin

类目:Information Retrieval (cs.IR)

关键词:Large language model-based, model-based recommender systems, demonstrated remarkable capabilities, language model-based recommender, Large language

备注: Accepted by ICDM'26

点击查看摘要

Abstract:Large language model-based recommender systems (LLM-RSs) have demonstrated remarkable capabilities, but are computationally unsustainable for many real-world applications. Compact LLMs offer a practical alternative, yet their reduced capacity often requires reasoning or knowledge distillation methods that increase latency or depend on larger models. Combined with autoregressive generation, these approaches face severe scalability bottlenecks. In contrast, discriminative LLM-RSs enable efficient full-corpus ranking through embedding similarity, but compact backbones remain limited in expressiveness and structural adaptivity. We propose the Fusion of Layer-wise Exits for Sequential Recommendation (FLEXRec), a discriminative framework that enhances compact LLMs while retaining scalable full-corpus ranking. FLEXRec inserts prediction heads (i.e., exits) at multiple transformer layers and adaptively fuses their score distributions. An adaptive continuous router (AC-Router) dynamically selects both the number and identity of exits for each user sequence, while a novel target-k hinge loss regulates routing sparsity. Experiments on three real-world datasets with Qwen 3 1.7B and Llama 3.2 3B show that FLEXRec achieves state-of-the-art accuracy among compact-backbone methods while remaining highly efficient. Code: this https URL

6. 【2608.17138】Overview of the TREC 2025 Product Search and Recommendation Track

链接https://arxiv.org/abs/2608.17138

作者:Dean E. Alvarez,Surya Kallumadi,Daniel Campos,ChengXiang Zhai,Alessandro Magnani,Rikiya Takehi,Michael D. Ekstrand

类目:Information Retrieval (cs.IR)

关键词:online seeking speed, purchasing efforts online, efforts online seeking, past few years, consumers have moved

备注

点击查看摘要

Abstract:In the past few years, consumers have moved the bulk of their product exploration and purchasing efforts online seeking speed, convenience, and price comparison with ease unimaginable for in-person shopping. As product catalogs have grown in diversity and size product search and recommendation have become a cornerstone for e-commerce sites. Despite the widespread usage of search engines in e-commerce, there is no high-quality dataset designed to evaluate end-to-end retrieval quality. In 2025, we ran a revised and continued version of the Product Search track previously run at TREC 2023 and TREC 2024. The 2025 product search track had two tasks: query expansion and related-product recommendation. The related-product recommendation task is particularly novel, providing an annotated data set of product relationships that distinguishes between complementary and related products. We anticipate the data from this track will enable better recommendation and search applications that reflect user needs, as a building block for conversational product discovery experiences.

Subjects:

Information Retrieval (cs.IR)

Cite as:
arXiv:2608.17138 [cs.IR]

(or
arXiv:2608.17138v1 [cs.IR] for this version)

https://doi.org/10.48550/arXiv.2608.17138

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
7. 【2608.16922】owards welfare-oriented recommendations in activity-travel behavior

链接https://arxiv.org/abs/2608.16922

作者:Ekin Ugurel,Takahiro Yabe

类目:Information Retrieval (cs.IR); Computers and Society (cs.CY)

关键词:mainstream recommender systems, rely on diverse, mainstream recommender, diverse heuristics, heuristics to rank

备注

点击查看摘要

Abstract:While mainstream recommender systems (RS) rely on diverse heuristics to rank alternatives, they generally lack a principled account of user welfare (i.e., whether accepting the recommendation will leave the user better off than other alternatives). The problem is particularly acute in activity-based travel behavior, where users incur costs they cannot recoup (i.e., energy, time) regardless of eventual satisfaction. As a result, existing systems may recommend options based on popularity or collaborative filtering, but may still leave users worse off than nearby or self-selected alternatives. We address this gap by introducing a welfare-oriented framework for activity recommendation that evaluates suggestions in terms of net utility, defined as experienced benefit minus travel costs. Specifically, we formalize two operational decision criteria: Positive Utility Probability (PUP) recommends only when the probability of non-negative net utility exceeds a threshold, while Regret Minimization (RM) recommends only when expected regret relative to the user's best organic alternative falls below a tolerance level. To evaluate these criteria, we develop an agent-based simulation in which heterogeneous synthetic travelers interact with multiple RS over time in a spatial environment with realistic travel costs, congestion, and behavioral feedback loops. This framework enables controlled counterfactual evaluations, and offers a practical foundation for designing RS that treat user welfare as a primary objective rather than an incidental byproduct.

8. 【2608.16921】MITRE-SAGE: A Multi-Agent Cybersecurity Question-Answering model

链接https://arxiv.org/abs/2608.16921

作者:Ali Habibzadeh,Farid Feyzi,Reza Ebrahimi Atani

类目:Information Retrieval (cs.IR); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Multiagent Systems (cs.MA)

关键词:Effective cybersecurity operations, heterogeneous security information, analysts increasingly struggle, operations require timely, large-scale heterogeneous security

备注

点击查看摘要

Abstract:Effective cybersecurity operations require timely and accurate analysis of large-scale heterogeneous security information; however, analysts increasingly struggle with information overload, alert fatigue, and time-constrained decision-making. Although large language models (LLMs) have demonstrated promising capabilities for question answering (QA), their effectiveness in cybersecurity remains limited by insufficient domain knowledge, a tendency to hallucinate, and difficulties in capturing both semantic and structural relationships. This work proposes MITRE-SAGE, a multi-agent retrieval-augmented generation framework that integrates semantic and structural cybersecurity knowledge to improve the reliability and interpretability of LLM-based QA systems. By decomposing complex tasks into query interpretation, evidence retrieval, and answer synthesis, MITRE-SAGE effectively supports cybersecurity tasks such as vulnerability assessment, threat profiling, and relationship extraction. Furthermore, we propose MITRE-QA, a comprehensive benchmark comprising 3,000 question-answer pairs for evaluating LLMs across diverse cybersecurity knowledge tasks, and use it to systematically evaluate MITRE-SAGE against representative baseline methods. Extensive experiments demonstrate that MITRE-SAGE consistently outperforms standalone LLMs and conventional RAG approaches. Notably, a lightweight configuration comprising Qwen2.5-7B sub-agents and a Qwen2.5-14B orchestrator achieves superior performance on five of the eight benchmark tasks, indicating the effectiveness of the proposed multi-agent framework. The results highlight the potential of MITRE-SAGE as a scalable and interpretable approach for reliable cybersecurity QA, while MITRE-QA provides a standardized benchmark for future research.

9. 【2608.16919】CARA: Cognitive Adaptive Recommendation Agent

链接https://arxiv.org/abs/2608.16919

作者:Weijun Gao,Jinyang Dong,Chuanru Ren,Hengxiao Li

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:Recent advances, large language models, advances in large, large language, introduced new opportunities

备注

点击查看摘要

Abstract:Recent advances in large language models and agent-based recommendation frameworks have introduced new opportunities for more flexible and context-aware recommendation. However, existing methods still largely rely on semantic matching, end-to-end generation, or loosely structured agent workflows, without explicitly modeling how user preferences are processed and translated into final decisions. To address this limitation, we propose CARA, a cognitively inspired recommendation framework that formulates recommendation as a structured decision-making process. The core intuition of CARA is that user decisions are jointly shaped by two complementary mechanisms: intuitive affective preference and deliberate rational evaluation. Accordingly, CARA organizes recommendation into two coordinated stages: candidate filtering, which narrows the search space based on coarse-grained preference constraints, and dual-perspective decision modeling, which captures recommendation decisions through affective and rational judgment. We further introduce a boundary-aware KTO strategy that prioritizes instructions the model can solve occasionally but not consistently, thereby increasing the density of informative preference signals. Extensive experiments on three Amazon Reviews domains show that CARA achieves the best performance on most evaluation metrics, with relative improvements of up to 10.15% over the baseline.

10. 【2608.16918】Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

链接https://arxiv.org/abs/2608.16918

作者:You Zuo(ALMAnaCH),Kim Gerdes(LISN, Qatent, STL),Éric de la Clergerie(ALMAnaCH),Benoît Sagot(ALMAnaCH)

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)

关键词:structured technical documents, highly structured technical, recall-oriented search task, Sparse Coverage, task over long

备注

点击查看摘要

Abstract:Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but single-vector representations may compress multiple technical components, functions, and constraints into a single embedding. We propose Sparse Coverage, an unsupervised semantic retrieval framework that maps local span embeddings to a sparse vocabulary of embedding-space centers. The centers are selected with a coverage-oriented k-center objective, and spans activate nearby centers to produce sparse representations compatible with inverted-index retrieval. Experiments on CLEF-IP 2013 show that Sparse Coverage matches or exceeds the document-level recall of strong dense patent encoders in several configurations, while remaining competitive for passage-level retrieval. By combining local semantic evidence with sparse inverted-index search, Sparse Coverage provides an effective first-stage retrieval approach for patent search.

11. 【2608.15382】Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot

链接https://arxiv.org/abs/2608.15382

作者:Ummara Mumtaz,Aimen Noor,Awais Ahmed

类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); Quantitative Methods (q-bio.QM)

关键词:Large language models, healthcare decision support, Large language, reward single-answer accuracy, decision support

备注

点击查看摘要

Abstract:Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.

12. 【2602.14784】Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

链接https://arxiv.org/abs/2602.14784

作者:Christos Koutsiaris

类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)

关键词:Breaking long, Breaking long documents, smaller segments, fundamental challenge, Breaking

备注: 8 pages, 4 figures. Code available at [this https URL](https://github.com/unseen1980/IDC)

点击查看摘要

Abstract:Breaking long documents into smaller segments is a fundamental challenge in information retrieval. Whether for search engines, question-answering systems, or retrieval-augmented generation (RAG), effective segmentation determines how well systems can locate and return relevant information. However, traditional methods, such as fixed-length or coherence-based segmentation, ignore user intent, leading to chunks that split answers or contain irrelevant noise. We introduce Intent-Driven Dynamic Chunking (IDC), a novel approach that uses predicted user queries to guide document segmentation. IDC leverages a Large Language Model to generate likely user intents for a document and then employs a dynamic programming algorithm to find the globally optimal chunk boundaries. This represents a novel application of DP to intent-aware segmentation that avoids greedy pitfalls. We evaluated IDC on six diverse question-answering datasets, including news articles, Wikipedia, academic papers, and technical documentation. IDC outperformed traditional chunking strategies on five datasets, improving top-1 retrieval accuracy by 5% to 67%, and matched the best baseline on the sixth. Additionally, IDC produced 40-60% fewer chunks than baseline methods while achieving 93-100% answer coverage. These results demonstrate that aligning document structure with anticipated information needs significantly boosts retrieval performance, particularly for long and heterogeneous documents.

计算机视觉

1. 【2608.18076】From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

链接https://arxiv.org/abs/2608.18076

作者:Xingjian Wang,Zhao Wang,Taihang Hu,Jun Zheng,Qing Jin,Qinye Zhou,Zhengtao Wu,Yongchao Du,Zuan Gao,Chao Lin,Yefeng Shen,Xiaoli Xu,Zhengze Xu,Hao Yan,Yuhang Yu,Mingzhou Zhang,Mengting Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Large-scale image generation, conventional pipelines typically, pipelines typically optimize, optimize task-specific datasets, typically optimize task-specific

备注: 19 pages, 10 figures

点击查看摘要

Abstract:Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

2. 【2608.18063】EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

链接https://arxiv.org/abs/2608.18063

作者:Jiayi Song,Shijie Huang,Fangtai Wu,Yubo Huang,Zhenxiong Tan,Songhua Liu,Jiaming Liu,Ruihua Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:prohibitive memory requirements, existing diffusion-based models, diffusion-based models remain, models remain constrained, quadratic attention complexity

备注

点击查看摘要

Abstract:High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requirements. A prevalent workaround employs a two-stage pipeline: editing at low resolution followed by independent super-resolution. However, this approach suffers from two critical issues: information divergence, where hallucinated details contradict the original high-resolution (HR) source, and texture degradation, manifesting as over-smoothed or over-sharpened artifacts. We propose EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing. Unlike conventional diffusion that regenerates from noise, we formulate refinement as structured data-to-data translation from the low-resolution (LR) edited result to its HR counterpart, explicitly conditioned on the original HR source to preserve authentic details. To efficiently incorporate HR source guidance, we introduce a prior-guided block-wise sparse attention mechanism that exploits semantic correspondence from first-stage editing to constrain cross-image interactions to spatially aligned regions, significantly reducing computational overhead. Extensive experiments demonstrate that EditBridge achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.

3. 【2608.18040】Optimize Your Sampling: Tuned Diffusion Sampling with Bayesian Optimization

链接https://arxiv.org/abs/2608.18040

作者:Travis Zhang,Christian Belardi,Justin Lovelace,Jin Peng Zhou,Saebyeol Shin,Carla P. Gomes,Kilian Q. Weinberger

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:large neural network, making generation computationally, generation computationally expensive, diffusion model typically, neural network

备注

点击查看摘要

Abstract:Sampling from a diffusion model typically requires many forward passes through a large neural network, making generation computationally expensive. While much work has focused on efficient solvers and samplers, comparatively little attention has been paid to selecting the sampling timesteps themselves. A recent line of work optimizes theoretically derived surrogates for sample quality rather than the quality metric itself. We propose Optimizing Your Sampling (OYS), which instead treats timestep selection as a black-box optimization problem, optimizing the target metric directly with Bayesian optimization. OYS outperforms both the default schedules and those of Align Your Steps on text-to-image generation, and improves over the default schedules on inpainting and other image tasks, in both quantitative and human evaluations. OYS requires no additional training, is applicable even to distilled models, and improves both simple and sophisticated samplers such as Euler and DPM-Solver++. A 5-step OYS schedule retains 89%-94% of the quality of a 50-step schedule while reducing inference cost by 10x.

4. 【2608.18035】Plug-and-Play Traffic Element Awareness for End-to-End Autonomous Driving

链接https://arxiv.org/abs/2608.18035

作者:Zongzheng Zhang,Jijun Wang,Saining Zhang,Shuo Wang,Yiru Wang,Hai Yang,Yang Chen,Yuwen Heng,Hao Sun,Anqing Jiang,Hao Zhao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:road signs play, human driving decisions, traffic element awareness, naturally influence, signs play

备注: Accepted by ECCV 2026; Project Page: [this https URL](https://zzongzheng0918.github.io/TE-Aware-E2E-AD/)

点击查看摘要

Abstract:Traffic elements such as traffic lights and road signs play a fundamental role in human driving decisions and should naturally influence end-to-end driving performance. However, existing end-to-end driving research predominantly focuses on dynamic road participants (e.g., vehicles and pedestrians), while the role of traffic elements remains largely unexplored. The community still lacks a systematic study quantifying their impact, largely because public datasets rarely provide structured traffic-element annotations and modern driving systems vary widely in architecture and training paradigm. In this work, we present the first systematic investigation of traffic element awareness for end-to-end autonomous driving. We construct a unified research infrastructure by augmenting multiple public driving datasets with comprehensive traffic-element annotations. To support diverse model families, we adopt a minimal and universal integration design that incorporates traffic-element signals into existing pipelines in a plug-and-play manner with negligible architectural modification. We evaluate this design across modern paradigms, including perception-prediction-planning pipelines, vision-language-action models (VLA), regression-based planners, diffusion-based policies, and trajectory-scoring frameworks, on nuScenes, NAVSIM-v1, NAVSIM-v2, and Bench2Drive. Across all paradigms and datasets, this simple integration consistently improves driving performance, demonstrating that traffic element awareness provides a robust and generalizable signal for end-to-end driving systems. Notably, on the challenging NAVSIM-v2 benchmark, our approach significantly improves state-of-the-art architectures and data pipelines, establishing a new state of the art.

5. 【2608.18034】Deep Academic Survey: Stateful Agentic Closed-Loop Paradigm for Academic Survey Automation

链接https://arxiv.org/abs/2608.18034

作者:Zhikai Xu,Zhucun Xue,Teng Hu,Yabiao Wang,Yong Liu,Jiangning Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:organizing rapidly expanding, coherent knowledge organization, Academic surveys play, requires extensive paper, fine-grained citation support

备注: Project page: [this https URL](https://zhikaixu24.github.io/projects/DAS/) | Code: [this https URL](https://github.com/ZhikaiXu24/DAS) | Data: [this https URL](https://huggingface.co/datasets/ZhikaiXu24/DAS-2M)

点击查看摘要

Abstract:Academic surveys play a central role in organizing rapidly expanding scholarly literature, yet their construction requires extensive paper analysis, coherent knowledge organization, fine-grained citation support, and reliable manuscript assembly. Existing Deep Research and automated survey generation systems address parts of this process, but typically do not coordinate paper understanding, literature organization, evidence-grounded drafting, and manuscript validation through a shared, revisable state. We introduce DAS, a stateful agentic framework for generating publication-oriented academic surveys. Its key idea is to separate reusable paper analysis from topic-specific manuscript construction. DAS builds on DAS-2M, a dynamically updated metadata lake containing survey-oriented representations of approximately two million papers. Its agents maintain explicit literature, organization, writing, and finalization states through candidate-grounded taxonomy planning, reverse paper-to-section routing, and hierarchical claim and citation planning. Semantic review reactivates only the affected writing states for repair and reevaluation, forming a scoped closed loop with deterministic validation. We further introduce DAS-Bench, a 30-topic benchmark, together with DAS-Eval, which assesses scholarly citation quality, taxonomic synthesis, hierarchical discourse, and manuscript assembly reliability through 16 criteria. Among systems evaluated on all 30 topics, DAS achieves the highest average in all four dimensions, with an overall score of 4.34 compared with 4.03 for the strongest competitor, and the same ordering is preserved on the matched 21-topic CS subset. Blinded expert evaluation further prefers DAS to Naive RAG on 27 of 30 topics and to AutoSurvey on 19 of 21 shared CS topics. The project page is available at this https URL.

6. 【2608.18028】Initialization-Free Bundle Adjustment Revisited: A Controlled Experimental Study

链接https://arxiv.org/abs/2608.18028

作者:Simon Weber,Mateo de Mayo,Je Hyeong Hong,Carl Olsson,Daniel Cremers,Ronald Clark

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:scene structure directly, recover camera poses, geometric initialization stages, aims to recover, avoiding the geometric

备注

点击查看摘要

Abstract:Initialization-free bundle adjustment (InitFree BA) aims to recover camera poses and scene structure directly from image observations, avoiding the geometric initialization stages of conventional structure-from-motion pipelines. Recent methods based on Object-Space Error (OSE) formulations and Variable Projection (VarPro) show encouraging optimization behavior from random camera configurations. However, existing evaluations primarily measure optimization success, leaving unclear whether a low OSE objective yields a valid metric 3D reconstruction. We revisit InitFree BA experimentally through a unified evaluation framework combining a C++ implementation of existing OSE formulations with a Blender-based dataset generator providing exact ground truth and controlled camera configurations and observation densities. Our experiments reveal a previously overlooked optimization--reconstruction gap: projective solutions with similarly low OSE values can lead to substantially different Euclidean reconstructions after metric upgrade. We identify initialization priors, landmark observation density, and metric-upgrade stability as key factors governing reconstruction success. Overall, our results suggest that the main challenge of InitFree BA is not merely minimizing OSE objectives, but obtaining projective reconstructions that admit reliable metric upgrade. We believe that the proposed benchmark, implementation, and analysis establish stronger experimental foundations for future research on initialization-free bundle adjustment, a problem largely unexplored within the computer vision community. Project page is available at this https URL.

7. 【2608.18012】Automated ACL Footprint Identification Using 3D Deep Learning

链接https://arxiv.org/abs/2608.18012

作者:Ruida Cheng,Ali Uneri,Gabriel Gibson,Frances T. Sheehan,Barry Boden

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:anterior cruciate ligament, ACL, ACL femoral footprint, ACL footprint, cruciate ligament

备注: 10 pages, 5 figures

点击查看摘要

Abstract:One of the most common reasons for anterior cruciate ligament (ACL) reconstruction failure is femoral tunnel malpositioning (ACL footprint center and tunnel orientation). Such failures may lead to the development of meniscal pathology and osteoarthritis. Accurate ACL femoral footprint identification is therefore essential for precise tunnel placement, restoration of the native knee joint mechanics, post-surgical knee joint health and prevention of graft failure. Recent advances in artificial intelligence (AI) bring new opportunities to improve image-guided orthopedic surgery. However, at present, existing AI research focuses primarily on ACL segmentation and rupture classification based on pre- and post-operative magnetic resonance (MR) images. Identification of the ACL footprint center using deep learning methods has not been thoroughly researched. Thus, the purpose of this study is to explore 3D deep learning models for ACL femoral footprint identification directly from 3D MR images. Two comprehensive 3D deep learning architectures were developed: a 3D graph convolutional neural network-based geometric model applied to 3D femoral meshes; and a 3D landmark-enhanced identification model based on 3D MR images. A total of 4883 right and 3087 left knee image sets were used from a publicly available database. Eighty percent (80%) were applied to model generation, and twenty percent (20%) were preserved for model testing. Both models achieved excellent performance; however, the image-based method outperformed the model-based method (average error of 2.1mm vs 2.8 mm). Thus, 3D deep learning provides a feasible clinical approach for ACL footprint localization and has the potential to improve ACL reconstruction footprint accuracy.

8. 【2608.18009】Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering

链接https://arxiv.org/abs/2608.18009

作者:Hsiang-Wei Huang,Fu-Chen Chen,Li-Wu Tsao,Cheng-Han Lee,Che-Chun Su,Lu Xia,Ronghui Peng,Jenq-Neng Hwang,Min Sun,Cheng-Hao Kuo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision Language Model, presents significant challenges, significant challenges due, resources for Vision, Vision Language

备注: ECCV 2026

点击查看摘要

Abstract:Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VLM) inference. Existing methods adopt visual search key frame retrieval method to select critical question-related key frames for VLM input. However, visual search methods are inefficient because they require visual search among thousands of video frames for each individual user query. In this work, we propose a memory tree guided key frame selection paradigm for efficient 3D question answering in embodied scenarios. Our method leverages a compact and reusable 3D scene representation, termed MemTree3D, which supports real-time online construction leveraging camera 6-DoF poses. MemTree3D captures multi-level 3D scene information, enabling a Large Language Model to efficiently query and retrieve question-relevant key frames through our scoring-based frame selection without reprocessing the entire video stream. On OpenEQA, our method improves the LLM-Match of GPT-4o by 17.4%, LLaVA-OneVision-7B by 5.8%, outperforms existing visual search methods. Our code is available at this https URL

9. 【2608.17995】AViTS: Adaptive Spatiotemporal Token Selection for Efficient Dynamic-Resolution Generation

链接https://arxiv.org/abs/2608.17995

作者:Haoran Qin,Zhengan Yan,Shikang Zheng,Xiaobing Tu,Jiacheng Liu,Yuqi Lin,Chang Zou,JinShan Liu,Peiliang Cai,Xiantao Zhang,Jinkui Ren,Linfeng Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieve high-quality generation, Diffusion Transformers, high-quality generation, costly due, due to iterative

备注: Accepted to ECCV 2026. 20 pages including appendix. Code: [this https URL](https://github.com/QHR69/AViTS)

点击查看摘要

Abstract:Diffusion Transformers (DiTs) achieve high-quality generation but are costly due to iterative sampling. Dynamic-resolution sampling reduces early-stage cost by denoising at low resolution; however, uniformly upsampling all latent tokens at resolution transitions incurs redundant computation and may degrade fine-detail consistency. Existing partial upsampling strategies typically rely on local latent structure cues or single-step statistics, making it difficult to jointly capture token-text semantic relevance and token-wise representation dynamics across diffusion steps. We propose AViTS, an adaptive spatiotemporal token selection framework for dynamic-resolution DiTs. AViTS models spatial importance via latent-text attention and temporal importance via token-level feature variation across diffusion timesteps, and fuses them to enable spatiotemporal importance-aware selective upsampling: it prioritizes resolution refinement for critical tokens while deferring less important ones, thereby reducing redundant high-resolution computation and improving the quality-efficiency trade-off. AViTS achieves up to 6.34x on FLUX and nearly 9x FLOPs reduction on Qwen-Image-Edit and FLUX.1-Kontext-dev, orthogonal to distillation, quantization, and feature caching, and reaching 14.76x with distilled models. Code: this https URL

10. 【2608.17988】GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation

链接https://arxiv.org/abs/2608.17988

作者:Ming Qian,Zijian Wang,Minchao Sun,Jincheng Xiong,Hang Zhang,Mu Xu,Chi Wang,Baoquan Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, spatially irregular, generators operate, vary widely, Splatting

备注

点击查看摘要

Abstract:Many scalable latent 3D generators operate on structured tensors, whereas pre-optimized 3D Gaussian Splatting (3DGS) reconstructions are unordered, spatially irregular, and vary widely in primitive count. We present GS-Voxel, a fitting-free structured latent framework, and evaluate it for large-scale aerial 3D Gaussian scene generation. GS-Voxel deterministically converts a compatible pre-optimized 3DGS reconstruction into sparse active voxels without additional per-scene optimization, retaining the sub-voxel positions and rendering attributes of the selected primitives. A GS-specific factorized VAE then separately encodes voxel geometry and local Gaussian attributes into sparse 3D latents whose size grows with the number of occupied voxels rather than being limited by a fixed scene-wide primitive count. We train image-conditioned flow models in the GS-Voxel latent space to generate aerial 3DGS scenes. A key application enabled by GS-Voxel is large-area scene generation: overlap-aware tiled inference extends synthesis beyond a single training crop conditioned on satellite-view images. Our results show that GS-Voxel provides structured latents for pre-optimized aerial 3DGS reconstructions, with latent capacity that grows with the number of occupied voxels.

11. 【2608.17983】Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity

链接https://arxiv.org/abs/2608.17983

作者:Alisher Myrgyyassov,Zhen Song,Bruce Xiao Wang,Yu Sun,Min Ney Wong,Yihao Zhou,Yongping Zheng

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:cross-dataset domain shift, segmentation remains challenging, degrade model generalization, probe variability, Ultrasound tongue

备注

点击查看摘要

Abstract:Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotations, probe variability, and acquisition noise often degrade model generalization. We present a source-free domain adaptation framework for robust ultrasound tongue segmentation built on a lightweight UltraUNet backbone. Starting from a checkpoint pretrained on only five labeled source images, simulating an underfitted constrained source model, the proposed method adapts to a fully-unlabeled target domain by iteratively refining pseudo-labels, filtering unreliable masks with a contour-based quality-control module, and generating target-style synthetic image-mask pairs through a segmentation-guided conditional GAN. The student model is then trained on a mixture of clean pseudo-labeled target images, noisy pseudo-labels with consistency regularization, and synthetic samples, enabling closed-loop adaptation without access to source data. We evaluate the method on 12 source-target transfer pairs across eight ultrasound tongue imaging datasets, and conduct source-size scaling experiments and ablation studies. Across all comparisons, the proposed framework improves segmentation overlap and contour accuracy over the baselines, including supervised ones. These results suggest that task-specific pseudo-label refinement and synthetic target-style augmentation can substantially improve source-free adaptation for ultrasound tongue imaging.

12. 【2608.17975】aDSL: Agentic 3D Creation via Joint Agent-Program Design

链接https://arxiv.org/abs/2608.17975

作者:Rui-Huan Wang,Si-Tong Wei,Jia-Qi He,Heng-Yi Wei,Baoquan Chen,Peng-Shuai Wang

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:enabling fine-grained edits, Programmatic representations provide, explicit structural control, enabling fine-grained, fine-grained edits

备注

点击查看摘要

Abstract:Programmatic representations provide a compelling paradigm for 3D content creation, enabling fine-grained edits, interpretability, and explicit structural control. Yet, agentic workflows that rely on large language models (LLMs) to author 3D programs remain brittle, often failing to translate high-level intent into consistent low-level geometry. We attribute this fragility to a mismatch between existing programmatic interfaces and the reasoning strengths of LLMs, which favor semantic structure and spatial relations over fragile numeric choices. In this paper, we jointly design an Agent-centric Domain-Specific Language (aDSL) and a role-specialized multi-agent system to close this gap. aDSL bridges semantic logic and geometric constraints by emphasizing composability and spatial reasoning; it enables agents to manipulate geometry through relational operators instead of brittle absolute coordinates. Building on aDSL, our training-free multi-agent system follows a Plan-Execute-Critic loop to decompose requests, synthesize code, and iteratively repair errors and constraint violations using execution feedback. Experiments show that this co-design improves robustness, controllability, and faithfulness to user intent. Our method outperforms prior LLM-based baselines on text-to-shape and image-to-shape tasks while preserving explicit structure, editability, and interpretability. It also enables downstream applications such as articulated object creation and structured scene composition. Our code is available at this https URL.

13. 【2608.17973】LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching

链接https://arxiv.org/abs/2608.17973

作者:Jinshan Liu,Haoran Qin,Xiaobing Tu,Jiacheng Liu,Jiahui Hu,Zhengan Yan,Yukun Xie,Kerui Shen,Jinkui Ren,Yuqi Lin,Xiantao Zhang,Linfeng Zhang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:achieved remarkable success, iterative sampling remains, high computational cost, video generation, practical deployment

备注: Accepted to ECCV 2026. 28 pages including appendix. Code: [this https URL](https://github.com/QHR69/LinCa)

点击查看摘要

Abstract:Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical deployment. Feature caching has emerged as a promising acceleration paradigm by reusing or predicting intermediate features across timesteps. However, existing training-free methods apply uniform prediction strategies that cannot adapt to the heterogeneous feature dynamics, causing significant quality degradation under high acceleration ratios. We propose LinCa, a feature caching framework based on learnable invertible networks. LinCa decomposes cached features into sub-components with distinct continuity properties via a lightweight invertible network and applies differentiated prediction orders matched to each component. The strict invertibility guarantees lossless reconstruction back to the original feature space, forming a unified Decompose-Predict-Reconstruct pipeline. By training separate predictors for different models and timestep segments, LinCa adapts to heterogeneous feature dynamics. Experiments on FLUX, Qwen-Image, and HunyuanVideo demonstrate that LinCa, with less than 0.2% additional parameters, significantly outperforms existing methods and maintains near-lossless quality at 5-7x speedup. Code: this https URL

14. 【2608.17966】SFMformer: A Spatial-Frequency Modulation Transformer for Lightweight Image Super-Resolution

链接https://arxiv.org/abs/2608.17966

作者:Chih-Hsiang Yang,Chia-Min Lin,Ching-Yu Tsai,Yung-Che Wang,Jen-Shiun Chiang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:lightweight image super-resolution, efficient Transformers, Transformers for lightweight, Sparse attention mechanisms, image super-resolution

备注: 20 pages, 13 figures, 5 tables

点击查看摘要

Abstract:Sparse attention mechanisms, which score all token pairs but propagate only the strongest, now underpin the most efficient Transformers for lightweight image super-resolution. This paper observes that sparsification changes what it means to improve such a network. A dense attention layer has one place where representation quality matters: the aggregation of attended features. A sparse layer has two, because the top-k operator first decides which tokens survive and only then decides what to do with them, and a token discarded at the selection stage cannot be recovered downstream. Selection quality and aggregation quality are therefore separable targets, addressed by modules placed before and after the attention respectively. We test this by pairing a dual-branch spatial enhancement on the input of a progressive focused attention with a wavelet-domain modulation on its output, forming SFMformer. Measuring each module alone and jointly over all fifteen benchmark-scale pairs, we find their gains are not additive: the joint gain exceeds the sum of the individual gains on nine pairs, and the sign of the discrepancy is predicted by how much the weaker module contributes on its own (r = -0.72), so the two compound when they relieve different constraints and overlap when they relieve the same one. Enabling spectral modulation once per block rather than once per layer retains the effect at roughly one-sixth of its cost, keeping the model below one million parameters at every scale. SFMformer ranks first on 28 of 30 PSNR/SSIM entries across five benchmarks and three upscaling factors. We report the cases where the pairing does not help, and deploy the model on a Raspberry Pi 5 to confirm the design is practical under tight resource budgets.

15. 【2608.17942】Cross-Domain Generalization in Machine Unlearning via Label-Conditioned Energy Magnitude Regularization

链接https://arxiv.org/abs/2608.17942

作者:Syed Ali Ahmed(1),Syed Bilal Ahsan(1),Muhammad Zaigham Zaheer(2) ((1) National University of Computer and Emerging Sciences, Karachi, Pakistan, (2) Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE)

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Machine unlearning removes, Machine unlearning, unlearning removes, removes the influence, influence of specific

备注: 17 pages, 3 figures, accepted at the ECCV 2026 Workshop on Unlearning and Model Editing (UMe)

点击查看摘要

Abstract:Machine unlearning removes the influence of specific data from a trained model. However, most methods treat the forgotten concept as isolated. In this paper, we study what happens to the rest of the model when a class is forgotten, using a label-conditioned energy-based model (EBM) that assigns per-class energies, making the effect directly observable. We forget a class by raising the energy of its image-label pairs, training with a forget term, a retain anchor to the pretrained model, a global margin, and an energy regularizer that stops the energy magnitudes from growing without limit. A propagation term applies the same forget signal to retain samples, weighted by each sample's DINOv2 similarity to the forget class, so forgetting reaches images that resemble it and leaves the rest untouched. We evaluate on two benchmark datasets: 1) On a subset of DomainNet across four visual domains, we forget tiger, lion, and scissors one at a time. Forgetting a class in the sketch domain also erases it from real, clipart, and painting, with forgetting error reaching 98% and 99% for lion and scissors, and the effect carrying over to the most similar class. 2) On CIFAR-10, we turn off the propagation term and forget each of the ten classes on its own. Forgetting is complete (100%), while the other nine classes retain 98.5% of their pre-unlearning accuracy on average.

16. 【2608.17935】Beyond Instrument Motion: Recognizing Tissue Tension Toward Surgical Skill Assessment

链接https://arxiv.org/abs/2608.17935

作者:Marko Haralovi,Zhiqi Miao,Alexander Machiel Bont,Jiapan Guo,Frans van Workum,Estefania Talavera

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:manual expert review, invasive surgery largely, surgery largely relies, minimally invasive surgery, making it time-consuming

备注: The paper is accepted by ECCV 2026 Workshop On Medical Video Understanding and submitted the camera-ready version to the ECCV organization

点击查看摘要

Abstract:Surgical performance assessment in minimally invasive surgery largely relies on manual expert review, making it time-consuming, subjective, and difficult to scale. While existing surgical video understanding methods address tasks such as instrument segmentation, surgical phase recognition, and action recognition, they do not explicitly capture fine-grained tissue handling, a key indicator of surgical quality. To address this gap, we introduce tissue tension recognition, a new clinically motivated video understanding task for laparoscopic and robot-assisted rectal cancer surgery. To support this task, we construct SurgTension, the first expert-annotated tissue tension dataset, providing a benchmark for objective tissue tension recognition. We further propose TensionTRAC, a lightweight trajectory-based framework that models tissue tension from sparse point trajectories. Using a compact trajectory encoder, TensionTRAC achieves competitive performance against strong pretrained video backbones.

17. 【2608.17926】PerFact: Perception-Derived Fact Prompting for 3D Brain MRI Report Generation

链接https://arxiv.org/abs/2608.17926

作者:Jianyu Sun,Zhenxuan Zhang,Guang Yang,Peter J. Lally

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:brain MRI, Radiology report generation, multi-sequence brain MRI, generation has matured, default route

备注

点击查看摘要

Abstract:Radiology report generation has matured almost entirely on 2D chest radiographs, where the default route to better reports is a larger backbone or a pre-training one on medical data. We revisit that assumption on 3D multi-sequence brain MRI, a volumetric multi-disease regime, and find that the model is not the lever. Zero-shot medical and radiology vision-language models transfer poorly to brain MRI, with chest radiograph specialists failing most conspicuously, and five backbones fine-tuned identically across three model families and an order of magnitude in scale differ only marginally. What determines the quality of the report is the information injected into the prompt. We delegate perception to upstream 3D segmentation and classification, serialize their outputs into a structured fact sentence, and prompt a LoRA-adapted vision-language model with it; we call this \textbf{PerFact}. In a controlled study that fixes the backbone, data split, target reports, and adaptation while varying only the injected grounding, perception-derived facts outperform retrieved prior reports, retrieval becomes redundant once facts are present, and end-to-end predicted facts remain effective without any ground-truth annotation at inference. The residual gap between predicted and oracle facts is explained by the granularity of the facts rather than by the generator. Closed-ended visual question answering comes at no measurable cost to report quality, though the grounding source has little effect on it. On 3D brain MRI, grounding information, not model choice, is the dominant controllable factor in report quality.

18. 【2608.17923】AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM

链接https://arxiv.org/abs/2608.17923

作者:Fahad Ahammed,Omar Faruq Shikdar,Navid Zaman,Md Tahsin,Md. Nawab Yousuf Ali,Golam Sorwar

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:prevent life-threatening conditions, common abdominal emergencies, abdominal emergencies worldwide, requires prompt diagnosis, life-threatening conditions

备注

点击查看摘要

Abstract:Appendicitis is one of the most common abdominal emergencies worldwide and requires prompt diagnosis and treatment to prevent life-threatening conditions. However, accurately differentiating complicated cases, such as perforation or abscess formation, from uncomplicated appendicitis remains a significant clinical challenge. Among other methods, ultrasound is a safer and more cost-efficient diagnostic technique because of the lack of radiation exposure. In this research, an advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed. A dataset consisting of 4679 ultrasound images with 5 classes, namely perforated, abscess, acute, appendicolith, and normal, was used for the proposed model training and testing. Four pretrained deep learning models, DenseNet201, InceptionV3, ConvNextTiny, and VGG19, have been employed for detecting and classifying complicated appendicitis. In the initial configuration, InceptionV3 achieved the second highest accuracy, with a value of 69.21%. Owing to suboptimal performance with raw images, further optimization techniques, including image preprocessing, hyperparameter tuning, model fine-tuning, and image sharpening, were applied. These enhancements significantly improved the model's performance, with an accuracy of 95.58% for InceptionV3. The model performance is then explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas. This could make crosschecking with experts much easier.

19. 【2608.17917】Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

链接https://arxiv.org/abs/2608.17917

作者:Alma M. Liezenga,Lotte Nijskens,Henrik R. Baumann,Stefan Becker,Simon Bensberg,Niccolò Camarlinghi,Håvard R. Eiring,Alexander W. Johnsgaard,Tanel Liiv,Giuseppe Martino,Matteo Marturini,Matthias Rapp,Jan Erik van Woerden,Alexander Wolpert,Hugo J. Kuijf

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Automatic Target Detection, Detection and Recognition, Automatic Target, military decision support, autonomous operations

备注: This paper was originally presented at the International Conference on Military Communication and Information Systems, organized by the Information Systems Technology Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026

点击查看摘要

Abstract:Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advances in object detection and artificial intelligence (AI) significantly boosted the potential performance of ATD/R. However, the scarcity of publicly available military datasets limits the application of these systems. As a solution, this paper explores the use of publicly available models and civilian datasets to achieve reasonable performance in military contexts. We benchmark several state-of-the-art models, including six iterations of the YOLO series and two variations on the DETR framework, on a newly acquired military relevant dataset. This dataset features military vehicles and challenging circumstances, including various degrees of occlusions and small targets. The out-of-the-box version of each model is validated alongside a version finetuned on the VisDrone dataset. This dataset features small objects, an Air-to-Ground (A2G) perspective and relevant classes, potentially generalizing to our military ATD/R task. We compare the performance of the models using mAP@0.5 and mAP@0.5:0.95, across A2G and Ground-to-Ground (G2G) perspective, target size and model size, giving insight into the real-time capabilities of models. Our main findings are: (1) bigger models outperform smaller models, (2) DETR-based models show promising results compared to the YOLO series,(3) fine-tuning models on an out-of-domain A2G dataset, improves their A2G performance and slightly improves their performance on small objects, but (4) all models still struggle with detecting small objects in an A2G scenario. We conclude that, despite recent advances in object detection, in-domain training is still crucial for creating capable ATD/R systems.

20. 【2608.17884】CFB-GBM v2.0: An Augmented Longitudinal Dataset for Multi-Modal Glioblastoma Segmentation, Radiomics, and RANO Progression Tracking

链接https://arxiv.org/abs/2608.17884

作者:Alexandre G. Leclercq,Noémie N. Moreau,Hugo Audebert,Andros Nassar,Thomas Cochin,Thomas Leleu,Loïc Le Henaff,Alexis Desmonts,Yoann Poirier,Aurélie Dubru,Laura Guillemette,Pascal Lecoeur,Kévin Lemasson,Cyril Jaudet,Sébastien Bougleux,Romain Hérault,Carole Brunaud,Samuel Valable,Dinu Stefan,Charlotte Raboutet,Alain Batalla,Joëlle Lacroix,Roman Rouzier,Aurélien Corroyer-Dulmont

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:tumor in adults, median overall survival, Gross Tumour Volume, aggressive primary brain, primary brain tumor

备注: 9 pages, 2 figures,

点击查看摘要

Abstract:Glioblastoma (GBM) is the most aggressive primary brain tumor in adults, with a median overall survival of 15 months. Longitudinal, multi-modal imaging datasets with comprehensive clinical and treatment data are essential to support the development of reproducible computational methods for treatment response prediction, disease progression modelling, and personalized medicine. We present CFB-GBM v2.0, an extension of our previously released CFB-GBM dataset comprising 264 GBM patients treated according to the standard Stupp protocol. The primary contribution of this release is the completion of Gross Tumour Volume (GTV) delineations across all available timepoints ($t_0$, $t_1$ and $t_2$), increasing the overall GTV completion rate from 35% to 97%. This was achieved using a nnU-Net model pre-trained on BraTS 2021 and fine-tuned on CFB-GBM ground-truth contours, with the generated segmentations validated by five radiation oncologists. From these longitudinal GTV annotations, volumetric RANO 2.0 response category labels were derived for all available temporality pairs ($t_0 \rightarrow t_1$, $t_0 \rightarrow t_2$ and $t_1 \rightarrow t_2$). To further ease dataset usability and reproducibility, brain masks computed with HD-BET and pre-computed radiomic features extracted with PyRadiomics are provided for each patient timepoint and MRI modality. Additionally, the WHO classification guideline (2016 vs. 2021) applicable to each patient's diagnosis is now explicitly documented. CFB-GBM v2.0 is publicly available on The Cancer Imaging Archive (TCIA) at this https URL .

21. 【2608.17883】Improving Complex Moiré Removal with Generative Supervision

链接https://arxiv.org/abs/2608.17883

作者:Xinyang Gu,Zhilu Zhang,Honglei Xu,Yanting Mei,Yukang Ding,Wangmeng Zuo

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:high-quality paired data, learning-based image demoiréing, complex moiré patterns, complex moiré, moiré patterns

备注: 14 pages, 5 figures. Project page: [this https URL](https://xinygu-pavo.github.io/WildMoire/)

点击查看摘要

Abstract:The availability of high-quality paired data is essential for training learning-based image demoiréing models. However, it remains challenging for existing datasets to encompass the complex moiré patterns captured in uncontrolled real-world scenarios. Such degradations typically manifest as large-scale, multicolored moiré patterns. Moreover, these patterns frequently occur in images for which clean counterparts are difficult to obtain, such as photographs acquired from public displays or existing online resources. In this work, we propose a novel data engine designed to improve the removal of complex moiré patterns by generating training supervision. Specifically, we initially collect real-world images containing complex moiré patterns and localize the corresponding screen regions. Multiple image-conditioned generative foundation models are subsequently deployed to produce candidate references. To establish reliable supervision, these candidates are subjected to patch-level quality control to filter and select the optimal results. Based on this systematic paradigm, we construct the WildMoiré dataset, which contains 6.8K moiré-GT training pairs. For evaluation, we additionally build an independent test set comprising $\sim$250 pairs with captured clean ground truth. Extensive experiments on ESDNet, SDXL, and Qwen-Image-Edit demonstrate that the proposed generative supervision consistently improves the performance of complex moiré removal.

22. 【2608.17872】DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Model Performance

链接https://arxiv.org/abs/2608.17872

作者:Ramon Kaspar,Andrey Ignatov,Valentina Boeva

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:high-performing pathology tile, hundreds of millions, pathology tile encoders, high-performing pathology, released pathology encoders

备注: 26 pages, 5 figures. Accepted at the ECCV 2026 Workshop on Medical Foundation Models and Benchmarks (MedFM-Bench)

点击查看摘要

Abstract:Many high-performing pathology tile encoders are now foundation models with hundreds of millions to over a billion parameters. Encoding and storing the thousands of tiles in each whole-slide image with such models is costly on commodity hardware, so compact encoders that retain useful downstream performance are a valuable alternative. We present DistillPath-KS16, which starts from the existing 22M kaiko ViT-S/16 encoder and improves it by distilling from released pathology encoders used as frozen teachers. The recipe reads only the teachers' final class and patch tokens and trains on 6,000 public slides, needing neither their DINO nor iBOT pretraining heads nor a billion-tile corpus, so it applies to any released encoder that exposes backbone tokens. We distill four teachers spanning 86M to 1.1B parameters into the same student. Every variant improves the kaiko baseline on all three benchmarks we use, EVA, HEST, and PLISM, and the strongest teacher is task-dependent. On the seven-task EVA mean, DistillPath-KS16-Virchow2 reaches $0.795$, within $0.015$ points of Virchow2, the top-scoring model in our evaluation, at about $29\times$ fewer parameters; it also scores above H0-mini and GPFM on this aggregate metric, though that advantage is task-concentrated rather than uniform. Because it remains a 22M ViT-S/16 with 384-dimensional features, DistillPath-KS16 runs more than $25\times$ faster than Virchow2. Code is available at this https URL, and released model weights are available at this https URL.

23. 【2608.17832】GenRec: Knowing Where to Reconstruct and Where to Generate

链接https://arxiv.org/abs/2608.17832

作者:Ata Çelen,Jaewoo Jung,Federico Tombari,Marc Pollefeys,Sunghwan Hong,Michael Niemeyer,Daniel Barath

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:sparse input images, captured volume admit, view-dependent shading, plausible completions, synthesis from sparse

备注

点击查看摘要

Abstract:Generative novel view synthesis from sparse input images is rarely all reconstruction or all generation: pixels visible in some source view have a unique correct value modulated only by view-dependent shading, while pixels in disocclusions or beyond the captured volume admit a distribution of plausible completions. Existing generative novel-view-synthesis methods conflate these regimes under a single uniform loss, blurring the line between geometric fidelity and creative hallucinations even when scene geometry is injected through warped point clouds or projected depth. We introduce GenRec, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow. Guided by an observation mask derived from the source cameras and a monocular depth estimator, a flow matching backbone jointly denoises RGB and scene-coordinate maps across all target views, while a pixel-space refinement stage restores high-frequency detail on observed pixels; the same mask gates supervision so regression signals do not contaminate the generative prior. Across RealEstate10K, DL3DV-10K, and Mip-NeRF~360, in both single-view extrapolation and two-view interpolation, GenRec attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones, showing the effectiveness of our approach.

24. 【2608.17803】Scale Matters: Adaptive Granularity Selection for Cross-Species 3D Plant Organ Segmentation

链接https://arxiv.org/abs/2608.17803

作者:Carla Salazar,Lazaros Nalpantidis

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:models provide powerful, powerful feature representations, provide powerful feature, point cloud learning, foundation models provide

备注: Project page: [this https URL](https://dtu-pas.github.io/ags-plantseg/)

点击查看摘要

Abstract:Recent 3D foundation models provide powerful feature representations for point cloud learning by controlling spatial granularity. However, relying on a fixed spatial granularity severely limits generalization in applications like plant phenotyping, where organ morphology and size vary substantially across species and growth stages. To address this, we propose AGS-PlantSeg, a few-shot 3D plant organ segmentation method that leverages the frozen Utonia (arXiv:2603.03283) foundation model combined with Adaptive Granularity Selection. By dynamically selecting the best granularity levels for each specific plant model, our method extracts optimized geometric features for a lightweight MLP segmentation head. Extensive experiments across PLANesT-3D (arXiv:2407.21150), Pheno4D , and Crops3D demonstrate that AGS-PlantSeg significantly improves cross-species generalization, achieving 88.9% average mIoU performance and outperforming fixed-granularity baselines by 2.5 mIoU points. Despite requiring minimal annotated data, our approach is highly competitive with fully supervised, plant-specific architectures.

25. 【2608.17799】raining with synthetic data for drone detection in thermal imagery

链接https://arxiv.org/abs/2608.17799

作者:Tanel Liiv,Sander Soodla,Nzamba Bignoumba,Alma M. Liezenga,Toomas Pruuden

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Emerging Technologies (cs.ET); Robotics (cs.RO)

关键词:reduced texture information, weak thermal contrast, sensor noise, long-wave infrared, texture information

备注: To be presented at SPIE: Sensors + Imaging, Artificial Intelligence for Security and Defence Applications IV

点击查看摘要

Abstract:Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information, sensor noise, weak thermal contrast, and the scarcity of annotated data. This work investigates a synthetic-first training strategy that combines synthetic scene generation with fine-tuning on real data. We show that synthetic data provides an effective basis for learning initial object representations, while real in-domain thermal imagery is still essential for reliable deployment. Even small amounts of real IR data substantially reduce domain gaps. Our experiments indicate that dataset alignment has a stronger impact on performance than model scale. Finally, our analysis of the dataset suggests that semantic alignment in feature space is the strongest predictor of model performance, while radiometric properties such as entropy and dynamic range also contribute to detection robustness. This work provides a foundation for combining synthetic and real IR data for effective G2A drone detection.

26. 【2608.17787】ETHEREAL: A 25.6-$μ$s/inf. Low-latency Event-driven Graph-neural-network Processor for High-resolution Vision at the Edge

链接https://arxiv.org/abs/2608.17787

作者:Adrian Kneip,Martin Lefebvre,Daniel Gehrig,Victoria Catalán Pastor,Davide Scaramuzza,Marian Verhelst,Charlotte Frenkel

类目:Hardware Architecture (cs.AR); Computer Vision and Pattern Recognition (cs.CV)

关键词:Dynamic vision sensors, s-level time resolution, Dynamic vision, vision sensors, reach the low-latency

备注: This work has been submitted to the IEEE JSSC for possible publication

点击查看摘要

Abstract:Dynamic vision sensors (DVS) are enticing candidates to reach the low-latency, sub-ms target of edge-vision applications, as they generate events with a $\mu$s-level time resolution. However, using DVS front ends also calls for novel algorithm/hardware back ends capable of efficiently handling streams of sparse spatiotemporal events. While event-driven graph neural networks (EV-GNNs) have emerged as a solution on the algorithmic side that is both accurate and efficient, there is no dedicated hardware to date capable of efficiently supporting their mixed requirements of dense-regular compute operations and sparse-irregular memory accesses. We therefore introduce ETHEREAL, the first EV-GNN processor chip, capable of bridging this gap by means of a neighbor-parallel spline-convolution engine combined with a split-2D/3D memory hierarchy that introduces a novel spatiotemporal event-caching mechanism. Measurement results demonstrate a 25.6$\mu$s latency and a 1.6$\mu$J energy per end-to-end event-wise inference on the state-of-the art DAGr-GNN workload and VGA-resolution (640x480 pixels) DSEC dataset.

27. 【2608.17747】INA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Consistent Text-Free Inversion

链接https://arxiv.org/abs/2608.17747

作者:Qianlong Xiang,Miao Zhang,Kun Wang,Haoyu Zhang,Junhui Hou,Liqiang Nie

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:preventing harmful content, remarkable generative power, exhibit remarkable generative, harmful content, models exhibit remarkable

备注: The project page is [this https URL](https://qianlong0502.github.io/TINA-Plus-Homepage/)

点击查看摘要

Abstract:Although text-to-image diffusion models exhibit remarkable generative power, concept erasure techniques are essential for preventing harmful content. Existing adversarial probes evaluate these methods by testing whether erased concepts can still be recovered. However, existing erasure and probe methods remain largely text-centric, focusing on whether the text-to-image mapping is severed while overlooking whether the corresponding visual knowledge remains. To investigate this question from a visual perspective, we leverage diffusion inversion to probe whether a generative trajectory can reconstruct visual instances of an erased concept. Under a null-text condition, standard inversion avoids the textual pathway but amplifies approximation errors, hindering faithful trajectory recovery. To address this challenge, we introduce TINA+, a diffusion-consistent Text-free INversion Attack equipped with optimization-based inversion. We also find that unconstrained diffusion inversion may discover spurious trajectories, even allowing a randomly initialized diffusion model to reconstruct the target concept. Such trajectories may falsely indicate residual visual knowledge. TINA+ therefore introduces Diffusion-Consistent Trajectory Regularization to suppress this failure mode. By penalizing trajectories that fall far below the expected marginal energy evolution of diffusion, TINA+ suppresses spurious inversion paths while preserving its ability to recover erased concepts. Experiments across twelve erasure methods, four concept-erasure tasks, and different model architectures demonstrate that TINA+ reliably probes residual visual knowledge through diffusion-consistent visual trajectories. These results provide stronger evidence that current methods often obscure concepts by severing text-image links rather than eliminating the underlying visual knowledge.

28. 【2608.17726】Evaluation of AI-based Visual Crack Detection in Steel Bridges Using Probability of Detection

链接https://arxiv.org/abs/2608.17726

作者:Andrii Kompanets,Finn Michael Sherry,Remco Duits,Davide Leonetti,H.H. Snijder

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reduce maintenance costs, ensure public safety, maintenance costs, regularly inspected, corrosion in order

备注: Submitted

点击查看摘要

Abstract:Bridge structures are regularly inspected for structural damage such as cracks and corrosion in order to ensure public safety and reduce maintenance costs. Much research has been done on automating this process using computer vision methods, which are often evaluated and compared using metrics such as intersection over union, mean average precision, etc. However, predicting the actual effectiveness of an inspection method within the field of structural engineering from these metrics remains challenging. To enable the systematic use of these increasingly popular methods in engineering practice, evaluating the performance of these methods in a way that is compatible with standard engineering approaches is therefore an urgent necessity. We present a new statistical evaluation framework to allow the comparison of computer vision methods with conventional visual inspection for crack detection in steel bridges. The framework is based on probability of detection curves and can account for the influence of image resolution. We apply this evaluation method to the real-world ``Cracks in Steel Bridges'' dataset, which contains annotated images of cracks in bridge structures. The quantification of the probability of detection and its uncertainty enables a practical assessment of the effect of automated methods for damage detection in structural reliability analyses. In turn, this enables the wide-spread use of automated (AI-based) damage detection in safety critical applications. This evaluation method provides evidence that the proposed computer vision approach approach is robust for the crack detection task and can have a high added value as an addition to conventional visual inspection methods.

29. 【2608.17723】Vision-Language Models for Analog Gauge Reading: An Empirical Study of Specialization, Transfer and Reliability

链接https://arxiv.org/abs/2608.17723

作者:Abdul Mueez,Aaditya Baranwal,Junior Chaj-Mejia,Guneet Bhatia,Jason T. Voelker,Shruti Vyas

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Analog gauges remain, Analog gauges, Pressure Gauge dataset, costly or hazardous, gauges remain common

备注: Submitted to Engineering Applications of Artificial Intelligence

点击查看摘要

Abstract:Analog gauges remain common in industrial environments where manual inspection is costly or hazardous. The engineering application addressed here is direct numerical reading of single-target analog-gauge images, while the artificial-intelligence contribution is a systematic evaluation of specialization, transfer, robustness and reliability for a general-purpose vision-language model (VLM) without an explicit pointer-segmentation and geometric-reading pipeline. The Qwen2.5-VL-7B-Instruct model is evaluated using zero-shot prompting, in-context learning (ICL) and parameter-efficient fine-tuning with Quantized Low-Rank Adaptation (QLoRA) on a public synthetic dataset, a video-derived Pressure Gauge dataset and a proprietary industrial dataset. All fine-tuning experiments use a fixed 20-epoch protocol with the final epoch used for analysis; separate models with and without supplied gauge ranges remove prompt-setting confounds. The primary metric is range-normalized mean percentage error (MPE). The best fine-tuned MPE values are 2.39% on the synthetic dataset, with a 95% bootstrap confidence interval (CI) of 1.43-3.90%; 2.61% on the Pressure Gauge dataset, with a CI of 1.66-3.80%; and 4.43% on the proprietary industrial dataset, with a CI of 2.31-7.14%. Leave-one-dataset-out experiments reveal substantial transfer degradation on held-out synthetic and proprietary data, while robustness tests identify Gaussian blur as the strongest tested corruption. Reliability analysis shows that high-confidence errors remain possible, motivating abstention and independent validation in safety-critical use. These results support QLoRA-specialized VLMs for direct single-gauge reading but not yet a deployment-ready plant-monitoring pipeline.

30. 【2608.17707】DynaForcing: Overcoming Dynamic Collapse in Self-Forcing Distillation for Streaming Avatar Generation

链接https://arxiv.org/abs/2608.17707

作者:Yubo Huang,Sirui Zhao,Xinchen Yao,Zhengye Zhang,Jinyang Huang,Fengqi Cui,Shiwei Wu,Enhong Chen

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:Audio-driven avatar generation, generation requires realistic, Distribution Matching Distillation, requires realistic lip-sync, avatar generation requires

备注: Accepted at ACM International Conference on Multimedia (MM '26)

点击查看摘要

Abstract:Audio-driven avatar generation requires realistic lip-sync, expressive motion, and real-time streaming. Recent work achieves the latter via self-forcing with Distribution Matching Distillation (DMD), but this paradigm suffers from a critical failure that has not been systematically characterized: dynamic collapse, where the student model converges to a near-static optimum with high perceptual quality but severely suppressed temporal dynamics. We trace this to two causes: the reverse KL objective in DMD, which biases toward low-motion modes, and unanchored self-conditioning, which creates a feedback loop that amplifies collapse. This is especially harmful for avatars, where even subtle motion loss breaks lip-sync and expression. To address this, we propose DynaForcing, a training framework with three complementary strategies applied at different levels. Specifically, Hybrid Forcing anchors rollouts to ground-truth dynamics at the data level to break the feedback loop. Dynamics-Aware Reward Regularization introduces explicit motion rewards via the RL interpretation of DMD to counteract the reverse KL bias at the loss level. Reference Perturbation perturbs reference images to decouple identity from static details, forcing the model to rely on audio for motion at the conditioning level. We further introduce computation graph pruning and gradient replay, reducing the GPU footprint of self-forcing by over an order of magnitude. Experiments show that DynaForcing recovers dynamics to teacher-comparable levels (Dyn-Deg: 0.31 - 0.73, Sync-C: 7.03 - 7.68) while improving visual quality, resolving the quality-dynamics trade-off throughout training without early stopping.

Comments:
Accepted at ACM International Conference on Multimedia (MM '26)

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

Cite as:
arXiv:2608.17707 [cs.CV]

(or
arXiv:2608.17707v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2608.17707

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
31. 【2608.17704】Monitoring Pasture Restoration from Satellite Image Time Series: Caveats and Opportunities

链接https://arxiv.org/abs/2608.17704

作者:Linnea Sartorius,Isak Randahl,Delia Fano Yela,Georg Andersson,Sadegh Jamali,Aleksis Pirinen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Monitoring nature restoration, difficult ecological problem, Monitoring nature, satellite image time, image time series

备注: Accepted at the 3rd Workshop on Computer Vision for Ecology at ECCV 2026

点击查看摘要

Abstract:Monitoring nature restoration at scale is an important but difficult ecological problem. Deep learning methods to analyze satellite image time series (SITS) have been widely used for land surface monitoring. In semi-natural grasslands - the habitat type in focus in this work - restoration outcomes develop gradually, yet satellite observations are influenced by weather, acquisition conditions, and processing artefacts, making it difficult to distinguish genuine restoration signals from unrelated temporal variation. In this work, we examine - to the best of our knowledge, for the first time - whether restoration status can be detected directly from satellite image time series by formulating pasture restoration as a binary deep learning classification problem. We evaluate two common SITS deep learning architectures on different Sentinel-2 image combinations, across 1,397 restored Swedish pastures and find that explicitly modeling intra-year variability and per-pasture normalization increases separability, reaching 0.88 accuracy for the best model. We further investigate our results and perform a targeted bias analysis finding that reliable deployment requires temporally balanced labels and evaluation protocols that explicitly test for year-related confounding. We therefore frame our contribution not as a solved restoration-monitoring system, but as a realistic case study of what works, what fails, and what future studies should control for. Code and models are available at this https URL.

32. 【2608.17700】Environment-Invariant Subspace Learning for Generalizable Deepfake Detection

链接https://arxiv.org/abs/2608.17700

作者:Shenghao Chen,Hao Jia,Chen Li,Chunjie Ma,Zan Gao,Shengyong Chen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Cross-distribution generalization remains, Cross-distribution generalization, critical bottleneck, semantic priors, Cross-distribution

备注: 12 pages, 4 figures, 11 tables

点击查看摘要

Abstract:Cross-distribution generalization remains a critical bottleneck in deepfake detection. While recent efforts leverage the semantic priors of large-scale visual foundation models (VFMs), a noteworthy yet underexplored challenge remains: the susceptibility of these semantic priors to environmental interference from factors such as lighting and style. Crucially, this interference establishes spurious correlations between forgery cues and environmental patterns that severely limit generalization. To address this fundamental challenge, we propose an innovative Environment-Invariant Subspace Learning (EISL) framework. The core contribution of EISL is that it aims to disentangle features into orthogonal forgery-relevant invariant factors and environment-related residual factors via a learnable low-rank projection. To facilitate robust feature disentanglement, we also design an Environmental Intervention module that generates diverse and challenging intervention pairs, simulating out-of-distribution environmental shifts to guide the model toward discovering truly invariant forgery representations. Experiments across cross-dataset, cross-generator, whole-face synthesis, and corruption settings show consistent gains and competitive or leading performance against strong detectors, demonstrating improved robustness to unseen forgery types and environmental variations. This work provides a new perspective and a valuable exploration for understanding and tackling the generalization barriers of VFMs in deepfake detection.

33. 【2608.17695】Magnitude-Direction Decoupling for Fast Video Generation with Flow Matching Models

链接https://arxiv.org/abs/2608.17695

作者:Haonan Xu,Feiyang Chen,Songkui Chen,Hongpeng Pan,Zhefeng Wang,Xinyu Duan,Baoxing Huai,Yang Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Flow matching models, video generation achieve, generation achieve impressive, computational overhead due, Flow matching

备注

点击查看摘要

Abstract:Flow matching models for video generation achieve impressive performance but suffer from high computational overhead due to iterative denoising. In fact, the original model is not necessary for all denoising steps, allowing some steps to use lightweight alternatives for faster sampling. However, directly using caching or lightweight models can deviate from the original denoising trajectory, resulting in suboptimal performance. Through empirical analysis, we find that lightweight models can robustly capture the magnitude components of the original model's output, while caching provides reliable directional guidance. Building on this insight, we propose the Magnitude-Direction Decoupling (MDD) method, which adaptively employs a direction-calibrated lightweight model as a substitute for the original model to accelerate inference and effectively correct deviations in the denoising trajectory. Moreover, MDD further reduces inference costs by reusing magnitude information under classifier-free guidance (CFG). As a result, MDD offers a more reliable and lightweight solution to accelerate sampling. Experiments show that MDD outperforms existing acceleration methods, delivering promising speedups (e.g., up to 2.95x on Wan2.1) while preserving high visual fidelity and content richness.

34. 【2608.17682】Differentiable Voronoi Ray Tracing Beyond Rasterization Speeds

链接https://arxiv.org/abs/2608.17682

作者:Bernardo Taveira,Carl Lindström,Joakim Johnander,Fredrik Kahl

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:rasterized explicit primitives, explicit primitives, view synthesis, synthesis is dominated, dominated by rasterized

备注

点击查看摘要

Abstract:Real-time novel view synthesis is dominated by rasterized explicit primitives. These projection-based pipelines provide high throughput but require specialized extensions for non-pinhole effects such as distortion, rolling shutter, and depth of field. Ray-based rendering expresses these effects naturally but is generally assumed too slow for competitive real-time rendering. We analyze the factors governing throughput in differentiable Voronoi ray tracing and identify traversal length, per-cell work, and memory locality as principal determinants. Guided by this, we introduce VoroTracing, which co-designs the scene representation, optimization, and GPU execution to reduce these costs. Compact octahedral appearance textures reduce memory traffic, while surface-concentrated opacity promotes early termination. The fixed-budget representation is optimized without pruning or densification and rendered with a GPU implementation designed for coherent traversal. On Mip-NeRF 360, VoroTracing renders at 623 FPS on an RTX 5090, providing $3.2\times$ the throughput of the fastest prior ray-based method and $2.8\times$ that of 3D Gaussian Splatting, while maintaining competitive reconstruction quality. Our renderer supports fisheye, rolling-shutter, motion-blur, and depth-of-field effects through ray generation and sampling, requiring no specialized rasterization. These results show that real-time throughput can be achieved with the flexibility of ray-based rendering. We release our source code, see this https URL

35. 【2608.17662】Is Haar Enough? Exploring Symlets and Coiflets for Wavelet Convolution Layers

链接https://arxiv.org/abs/2608.17662

作者:Md Rifat Ur Rahman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:underlying efficiency frontier, enlarging receptive fields, chosen decomposition depth, Wavelet convolution layers, multiresolution analysis

备注

点击查看摘要

Abstract:Wavelet convolution layers have recently emerged as an efficient mechanism for enlarging receptive fields through multiresolution analysis, but prior work has fixed the wavelet basis to Haar or Daubechies at a chosen decomposition depth, leaving open whether a different basis can shift the underlying efficiency frontier. We identify and characterize a previously unexplored trade-off in this setting: bases with stronger approximation properties (longer filters) can reduce the decomposition depth required for competitive accuracy, yielding a net reduction in parameters and FLOPs despite higher perlevel transform cost. We formalize this as an F-vs.-L tradeoff (filter length vs. decomposition levels) and study it systematically across Haar, Daubechies, Symlets, and Coiflets under controlled architectures and budgets. On image classification (CIFAR-10, ImageNet-1K) and semantic segmentation (Cityscapes), Coiflet-based wavelet convolutions match Haar at deeper levels with approximately 32% fewer additional parameters and 33% fewer additional FLOPs, providing a concrete and actionable design choice for practitioners building wavelet-based architectures.

36. 【2608.17657】Denoised Variance-Based Pruning with Optimal Brain Bias Compensation

链接https://arxiv.org/abs/2608.17657

作者:Geon Tack Lee,Jaegul Choo,Kang Eun Jeon

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision Transformers, restricts edge deployment, carry massive computational, massive computational overhead, edge deployment

备注: Accepted to ECCV 2026

点击查看摘要

Abstract:Vision Transformers (ViTs) achieve state-of-the-art performance but carry massive computational overhead that restricts edge deployment. Although structural pruning has emerged as a key strategy to reduce these costs, existing methods often suffer from severe accuracy degradation or require expensive retraining. Recently, Variance-Based Pruning (VBP) introduced a promising paradigm by selecting neurons based on activation variance; however, it remains limited by statistical noise in finite-sample activation covariance and reliance on bias-only updates that cannot fully account for structural reconstruction error. To address these limitations, we introduce Denoised Variance-Based Pruning with Optimal Brain Bias Compensation (DVBP + OB$^2$C). We leverage random matrix theory to filter noise from the activation covariance spectrum for robust neuron selection and mathematically prove that integrating mean-shift compensation into the Optimal Brain Compression objective reduces the layer-wise Hessian exactly to the activation covariance matrix. This enables an optimal, closed-form update of the remaining weights using the same statistics gathered for selection. Extensive experiments on DeiT, Swin, and ConvNeXt architectures demonstrate that DVBP + OB$^2$C achieves state-of-the-art training-free performance; at 50% MLP pruning, it retains over 90% of the original Top-1 accuracy on Small and Base variants, outperforming VBP by up to 29.46% (ConvNeXt-T) and 7.33% (Swin-S). The code is available at: this https URL.

37. 【2608.17635】MaLViL: Multi-axis Low-rank Vision-LSTM for Medical Image Segmentation

链接https://arxiv.org/abs/2608.17635

作者:Afshin Bozorgpour,Sina Ghorbani Kolahi,Moein Heidari,Ilker Hacihaliloglu,Dorit Merhof

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:enables efficient global, efficient global modeling, existing segmenters confine, lose fine anatomical, segmenters confine ViL

备注: Accepted at the MICCAI Workshop on Machine Learning in Medical Imaging (MLMI), 2026

点击查看摘要

Abstract:Vision-LSTM (ViL) enables efficient global modeling, but its cost still scales with the number of spatial tokens, so existing segmenters confine ViL to a coarse bottleneck and lose fine anatomical detail. Rasterizing 2D features into a 1D sequence further breaks adjacency across the orthogonal scan axis. We propose MaLViL, a Multi-axis Low-rank Vision-LSTM network that extends ViL across decoder resolutions. Bidirectional low-rank ViL (Bi-LRViL) reasons on a compact orthonormal subspace and preserves detail through an orthogonal residual; scale-aware SaLViL restores cross-axis neighbors before serialization; and a Cross-Directional Mixer (CDM) fuses orthogonal horizontal and vertical traversal paths. Statistics-Guided Skip Modulation (SGSM) further retains boundary cues in encoder skips. On skin-lesion, ultrasound, and multi-organ CT benchmarks, MaLViL achieves competitive or state-of-the-art segmentation accuracy, while reducing ViL operator memory by up to $83\times$ at fine decoder resolutions. Code is available at: this https URL.

38. 【2608.17623】RetiWave-Mamba: A Dual-Stream Network for Retinal Disease Detection based on Multi-scale Context and Frequency-Adaptive Mamba Projection

链接https://arxiv.org/abs/2608.17623

作者:Cheng Cheng,Jin Hong

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:irreversible vision impairment, Optical Coherence Tomography, accurate diagnosis essential, vision impairment, making early

备注

点击查看摘要

Abstract:Retinal diseases are a leading cause of irreversible vision impairment, making early and accurate diagnosis essential for effective treatment. Optical Coherence Tomography (OCT) serves as a critical imaging modality for this purpose, yet its automated analysis is hindered by inherent speckle noise, varying lesion scales, and subtle inter-class similarities. To address these challenges, we propose a novel framework, RetiWave-Mamba, which integrates spatial-frequency domain learning with state-of-the-art state space models. The framework utilizes Discrete Wavelet Transform (DWT) to decompose OCT images into low- and high-frequency streams, enabling decoupled processing of structural context and fine-grained details. For the low-frequency branch, we design a Multi-scale Contextual Localization Module (MCLM), which synergizes multi-scale dilation with spatial attention to expand the global receptive field and precisely localize lesion regions. For the high-frequency branch, we introduce an Attention-Guided High-Resolution Network (AG-HRNet) equipped with an intelligent gating mechanism to suppress noise propagation during multi-scale interactions. Furthermore, a Frequency-Adaptive Mamba Projector (FAMP) is incorporated to capture long-range dependencies within disjoint high-frequency textural features. Extensive experiments on the OCT-C8 dataset demonstrate that our approach achieves a state-of-the-art (SOTA) classification accuracy of 98.25%, surpassing existing methods. These results highlight the efficacy of RetiWave-Mamba in robustly identifying retinal pathologies under noisy conditions, offering a promising tool for clinical diagnosis.

39. 【2608.17607】PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts

链接https://arxiv.org/abs/2608.17607

作者:Bowen Liu,Qixiang Zhang,Xiaomeng Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:integrate gigapixel-scale visual, primarily measure final, measure final answer, current question-answering benchmarks, question-answering benchmarks primarily

备注

点击查看摘要

Abstract:Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric vulnerable to linguistic priors and benchmark regularities, and insufficient to establish that predictions are grounded in the supplied tissue. We introduce PathoArgus-Bench, a benchmark and evaluation protocol that explicitly tests the full evidence chain: availability, accessibility, use, and responsiveness. PathoArgus-Bench comprises 22,078 four-choice questions from 4,913 patients across 15 TCGA projects, covering six pathology capabilities across three levels of evidence demand, and operates under a fixed reader budget that retains only a small fraction of the gigapixel context. To further isolate evidence-grounded reasoning, we contribute ESG (Evidence State Quartets), a controlled set of 483 quartets where the question text is fixed while the target WSI set is moved, replaced, or removed, requiring consistent predictions across all states. Evaluating 20 general-purpose, medical, and pathology-specific systems reveals a stark gap: while GPT-5.6 achieves 57.09% overall accuracy and 57.04% on ESG, it correctly completes only 19 of 483 quartets (3.93% QExact), exposing that row-level accuracy does not translate into reliable evidence grounding. We also introduce PathoArgus, a fixed-budget reader that allocates context via question relevance and spatial coverage, attaining 50.39% overall accuracy yet only 1.86% QExact--demonstrating that improved context access alone does not ensure consistent evidence-based prediction. Our benchmark and diagnostics establish that acquiring useful whole-slide context is necessary but far from sufficient, and call for a shift from answer-centric to evidence-grounded evaluation in computational pathology.

40. 【2608.17598】SpurCon: Weighted Supervised Contrastive Learning for Mitigating Spurious Cues in Medical Imaging

链接https://arxiv.org/abs/2608.17598

作者:Shenhav Nadir,Meir Yossef Levi,Eyal Gofer,Guy Gilboa

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:applications remains limited, remains limited due, high-risk medical applications, medical applications remains, deep neural networks

备注

点击查看摘要

Abstract:Despite the rapid progress of deep neural networks in visual recognition, their adoption in high-risk medical applications remains limited due to reliability and robustness concerns. Models may exploit spurious correlations, particularly in medical imaging, where devices or treatment artifacts often co-occur with pathology. In small or imbalanced datasets, such cues further reduce worst-group performance and undermine clinical trust. To solve these issues, two major challenges should be addressed: identifying dataset-specific spurious cues, which typically require domain knowledge, and mitigating reliance on them. To tackle both, we propose SpurCon, a lightweight framework based on a novel supervised contrastive loss formulation that leverages available metadata and predicted spurious labels to enhance robustness. We introduce a fast few-shot procedure, without network training, to estimate spurious labels using a small number of expert-annotated samples. We then propose a weighted supervised contrastive objective, WtSupCon, that reshapes the representation geometry by assigning sample-specific weights that depend on the [pathology, spurious, metadata] combination. For example, the highest weight is assigned to samples that differ only in their spurious label. This yields highly similar representations for images with the same metadata and pathology, differing only in the predicted spurious label. Our method operates on pretrained image encoders (such as BiomedCLIP) and trains only a lightweight projection head. We evaluate SpurCon on a synthetic setting and on Waterbirds, CheXpert, a chest X-ray classification dataset, and ISIC 2020, a skin cancer classification dataset. Our approach delivers the best spurious-mitigation performance, balancing well worst-group and overall accuracy on multiple datasets.

41. 【2608.17566】CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

链接https://arxiv.org/abs/2608.17566

作者:Fuchen Long,Cong Wang,Zitao Gao,Wenhao Zhong,Yu Cheng,Xiaolu Hou,Yan Li,Xiao Cao,Xinlong Sun,Xi Chen,Yu Liu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:compositional instruction-guided video, video editing, instruction-guided video editing, instruction-based video editing, editing

备注: Project page: [this https URL](https://coinve200k.github.io) ; Dataset is available at [this https URL](https://huggingface.co/datasets/FireCRT/CoinVE-200K;) see source codes at [this https URL](https://github.com/coinve200k/CoinVE-200K)

点击查看摘要

Abstract:The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided video editing. In particular, multiple editing intents must be jointly understood and faithfully executed within the same video. To address this issue, we introduce CoinVE-200K, a large-scale, high-quality dataset for Compositional Instruction-Guided Video Editing. CoinVE-200K contains 1080p video-editing pairs of up to 201 frames, covering diverse compositional scenarios where each sample involves 2 to 5 atomic editing operations. The instructions target humans, objects, and backgrounds, and cover edit types such as addition, removal, modification, and stylization. All samples are built through a carefully designed generation and filtering pipeline to ensure instruction faithfulness, visual quality, temporal consistency, and compositional diversity. We also introduce CoinVE-Bench, a benchmark for compositional-instruction video editing across diverse subjects, operation types, and instruction complexities. Furthermore, we present CoinVE-Edit, a 22B compositional video editing model built upon Wan2.1-T2V-14B and Qwen3-VL-8B-Instruct. CoinVE-Edit disentangles region-aware attention for different editing instructions, enabling precise multi-region editing while preserving irrelevant content and temporal coherence. Experiments on CoinVE-Bench show that CoinVE-Edit achieves strong performance in instruction following, compositional editing accuracy, visual quality, and temporal consistency.

42. 【2608.17564】Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

链接https://arxiv.org/abs/2608.17564

作者:Zongyang Qiu,Yihan Wu,Kaixuan Fan,Bo Li,Hui Xiong

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:controlled ablations repeatedly, leaves understanding flat, ablations repeatedly find, Unified multimodal models, controlled ablations

备注: 27 pages, 10 figures

点击查看摘要

Abstract:Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model's behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman $\rho = +0.68$). That objective's alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a $0.1\%$ relative loss of the model's general text-to-image ability, against $41\%$ for the standard generative route. Our code is at this https URL.

43. 【2608.17561】Leveraging existing sparse point annotations for benthic imagery dense segmentation

链接https://arxiv.org/abs/2608.17561

作者:Cesar Borja,Breck A. McCollum,Jarret E. Byrnes,Kenneth Sebens,Ana C. Murillo

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:processing marine imagery, marine imagery severely, global environmental change, imagery severely limit, marine ecosystems

备注

点击查看摘要

Abstract:The health of marine ecosystems is a critical indicator of global environmental change, yet the physical constraints of underwater observation and the intrinsic challenges of processing marine imagery severely limit the scalability of systematic monitoring. While recent visual foundation models such as the Segment Anything Model (SAM) series show great promise, they still struggle with the fine-grained recognition required in these complex scenarios and still require expert supervision. Our work addresses this gap by bridging state-of-the-art foundation models with existing sparse supervision. Because historical benthic surveys are typically annotated with only a few sparse expert points per image, we utilize these legacy point-labels as visual prompts for SAM2. Our primary contribution is a novel mechanism to automatically identify which of these points are suitable, and which are actively harmful, when used for propagation. By filtering out unreliable points, we extract high-quality pseudo-ground-truth masks capable of training more accurate, fine-grained semantic segmentation models. We demonstrate the effectiveness of our approach on public benthic data and introduce a new, challenging benchmark featuring real-world sparse expert annotations, paving the way for scalable ecological analysis.

44. 【2608.17559】MSEditor: Toward Consistent Multi-Shot Video Editing

链接https://arxiv.org/abs/2608.17559

作者:Kunyu Feng,Yue Ma,Bingyuan Wang,Yuefeng Wang,Zhiyuan Qin,Hao Cheng,Hao Li,Qifeng Chen,Zeyu Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multi-shot video sequence, unified modifications, tackle the problem, problem of performing, multi-shot video

备注: ECCV 2026

点击查看摘要

Abstract:In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos consist of discontinuous temporal segments that vary significantly in viewpoint, camera scale, and subject pose, leading to severe identity drift and cumulative error propagation. Achieving coherent edits requires establishing reliable cross-shot semantic awareness to maintain stable subject appearance and visual continuity across these disjointed boundaries. To address this, we propose MSEditor, the first framework designed specifically for consistent multi-shot video editing. To overcome the scarcity of high-quality multi-shot training data, we repurpose existing multi-view video datasets to provide robust cross-shot supervision. Architecturally, we introduce a Supervisory Adapter that injects this cross-shot information into the diffusion backbone, enabling the model to learn identity-consistent representations. Furthermore, to effectively mitigate cumulative errors and ensure long-range temporal coherence, we design a Cross-Shot Packing strategy that dynamically aggregates information from semantically related shots within the self-attention window. Extensive experiments demonstrate that MSEditor significantly outperforms existing methods on our curated multi-shot video editing benchmark in terms of identity preservation, temporal stability, and overall visual quality.

45. 【2608.17550】Code as Representation: A Compilable Parsing Paradigm for Academic Documents

链接https://arxiv.org/abs/2608.17550

作者:Rihui Jin,Jun Wang,chengyuan zhu,Liang Mingyu,Yue Gao,Li Yunxuan,Kuicai Dong,Guilin Qi,Lin Ren,Yongrui Chen,Xinbang Dai,Jiaqi Li,Tongtong Wu,Gholamreza Haffari

类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)

关键词:knowledge remains locked, Multimodal Large Language, Structured Academic Elements, Large Language Models, knowledge remains

备注: Accepted by ACM MM 2026

点击查看摘要

Abstract:Academic papers are a primary carrier of scientific knowledge, yet most of this knowledge remains locked in PDFs that are optimized for human reading rather than machine use. For Multimodal Large Language Models (MLLMs), the core challenge is not only perception, but representation: scientific pages interleave text with Structured Academic Elements (SAEs) such as tables, formulas, charts, and pseudocode, whose structure, data, and logic are poorly preserved by common surrogates like Markdown. We therefore propose Compilable Academic Document Parsing (CADP), a paradigm that reconstructs a full page as contextual \LaTeX{} plus executable Python, so that structure-preserving elements and executable chart representations can be reconstructed, recompiled, and directly verified against the source page. To support this setting, we introduce CADP-Bench, an expert-verified benchmark of full academic pages containing tightly coupled text and multiple SAE types, evaluated through a re-injection compilation protocol. We further study current capabilities using SOTA MLLMs and an exploratory multi-agent baseline that incorporates common agentic techniques. Results show that even frontier models still struggle to produce high-fidelity executable reconstructions, highlighting substantial room for improvement in structure-aware scientific document parsing. CADP-Bench is released for future research.

46. 【2608.17535】GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting

链接https://arxiv.org/abs/2608.17535

作者:Qijian Tian,Zimeng Wu,Xuhong Wang,Lizhuang Ma,Xin Tan

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Simultaneously reconstructing, Gaussian Splatting, reconstructing and understanding, environments is essential, embodied agents

备注

点击查看摘要

Abstract:Simultaneously reconstructing and understanding 3D environments is essential for embodied agents. Toward this goal, feed-forward semantic 3D Gaussian Splatting (3DGS) efficiently constructs semantic scene representations from sparse multi-view observations. However, existing methods lack explicit instance discrimination and mainly support category- or phrase-based semantic queries. To this end, we propose GroupForward, an instance-grouped feed-forward Gaussian splatting model that reconstructs geometry, appearance, instance structure, and semantics from sparse, unposed, and uncalibrated multi-view images. Unlike existing methods that attach high-dimensional semantic features to each Gaussian, GroupForward learns compact instance embeddings that group Gaussians into cross-view consistent 3D instances, reformulating feed-forward semantic 3DGS from per-Gaussian semantic feature rendering to instance-level semantic aggregation and propagation. Building on these instance groups, we further propose a Referential Scene Reasoning Framework (RSRF) for complex 3D referring segmentation. RSRF constructs an instance-grouped 3D scene graph and retrieves candidate instances for a given referring expression. A vision-language model then reasons over structured instance evidence and multi-view observations to identify the referred instance among the candidates. RSRF thereby extends language interaction from simple semantic querying to complex referential scene reasoning. Experiments on semantic reconstruction and referential reasoning demonstrate the effectiveness of our instance-grouped reconstruction and reasoning framework.

47. 【2608.17522】Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

链接https://arxiv.org/abs/2608.17522

作者:Mohammad Javad Ahmadi,Hamid D. Taghirad

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Persistent shortages, traditional training methods, training methods highlight, data-driven approaches, workforce and inherent

备注

点击查看摘要

Abstract:Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.

48. 【2608.17521】BrainNorm: A Foundation Model that knows Normal via Semantic Atlas Pretraining

链接https://arxiv.org/abs/2608.17521

作者:Madhumitha Venkatesh,Shanawaj S Madarkar,Konda Reddy Mopuri

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:normative foundation model, structural MRI, Semantic Atlas Latent, foundation model, trained and tested

备注

点击查看摘要

Abstract:We introduce BrainNorm, a normative foundation model, trained and tested on ~66,000 T1-weighted structural MRI (T1w sMRI) scans. By leveraging language-image style contrastive pretraining on healthy cohorts across ages, BrainNorm learns a Semantic Atlas Latent space (SAL), where each scan is represented as a set of atlas-parcel embeddings. This yields parcel-specific healthy aging template trajectories that support age-consistent template matching and localized deviation scoring relative to a subject's chronological age. Across 6 downstream cohorts, BrainNorm demonstrates generalization evaluated across 25 task-setting combinations spanning age estimation, brain-age gap estimation, parcel identification, and single- multi-disease classification tasks under direct inference, zero-shot, few-shot full-data linear-probe settings. The resulting deviation patterns in SAL space enable zero-shot tasks for disease prediction using parcel-wise abnormalities. Fine-tuning on healthy-only cohorts of downstream datasets further improves the performance of various tasks. Across all classification tasks, linear probing on BrainNorm's frozen embeddings outperforms 9 baselines finetuned under end-to-end supervision. Furthermore, the localized deviations identified by BrainNorm across various neurodegenerative disorders closely align with established neurodegeneration pathology in clinical literature.

49. 【2608.17519】Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

链接https://arxiv.org/abs/2608.17519

作者:Hanna Hoffmann,Felix von Bechtolsheim,Stefanie Speidel,Rebecca Hisey

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:shown strong in-domain, encode dataset-specific visual, question remains unasked, strong in-domain results, fundamental question remains

备注

点击查看摘要

Abstract:Vision-based surgical skill assessment has shown strong in-domain results, yet a fundamental question remains unasked: do these models learn transferable representations of surgical proficiency, or do they merely encode dataset-specific visual patterns? This paper systematically analyzes what limits cross-domain skill transfer between the GOALS and OSATS assessment scales using the LASANA and JIGSAWS datasets. Each evaluated method serves a targeted diagnostic purpose: end-to-end training to test whether supervised skill learning transfers directly, Adaptive Sharpness-Aware Minimization (ASAM) to probe whether flatter loss landscapes improve generalization, and augmentation-based self-supervised and contrastive learning to assess whether domain-invariant pretraining decouples skill from visual context. Transfer is evaluated in both directions using a disjoint-participant held-out test set for JIGSAWS. Results reveal an asymmetry: backbones pretrained on JIGSAWS achieve CCC values of 0.77 to 0.80 on LASANA, closely matching the end-to-end baseline, showing cross-rubric transfer is feasible when the target domain provides consistent supervision. Transfer to JIGSAWS fails across all methods, likely due to annotation inconsistencies. Control experiments with a Kinetics-pretrained backbone suggest task-specific heads carry the majority of the skill prediction burden, while the backbone need only provide adequate spatiotemporal features. These findings offer a new perspective on vision-based skill assessment: the central question of whether skill representations transfer across scoring systems has not been previously investigated. Results indicate the visual component is dominant but not solely responsible for skill prediction; further work is needed to conclusively disentangle transferable skill features from those bound to a specific visual domain.

Subjects:

Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

Cite as:
arXiv:2608.17519 [cs.CV]

(or
arXiv:2608.17519v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2608.17519

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)

Submission history From: Hanna Hoffmann [view email] [v1]
Tue, 18 Aug 2026 08:44:31 UTC (6,125 KB)

50. 【2608.17514】SE-MoLoRA: Shared-Expert LoRA Adapters for Domain-Specific Photographic Assessment

链接https://arxiv.org/abs/2608.17514

作者:Bishwash Khanal,Anlan Zhang,Sasu Tarkoma,Tommi Mikkonen,Abhishek Kumar

类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)

关键词:describe images fluently, judgment remain entangled, aesthetic judgment remain, provide actionable photographic, Vision-language models

备注

点击查看摘要

Abstract:Vision-language models can describe images fluently, but they often fail to provide actionable photographic critique because semantic content and aesthetic judgment remain entangled. We propose SE-MoLoRA, a modular parameter-efficient adaptation framework for domain-specific photographic assessment. The method separates general photographic knowledge from specialist residual judgments using an always-active shared LoRA expert and routed adapters for composition, lighting, and technical quality. A lightweight query router selects the relevant specialist, enabling targeted critique without training separate full models. A rank-64 shared adapter captures broad photographic vocabulary, while rank-32 specialists learn domain-specific residuals with an orthogonal regularization penalty that encourages disentangled representations. Training data is obtained by distilling the Reddit Photo Critique Dataset into domain-labeled critique samples. On held-out critique generation, SE-MoLoRA improves BERTScore-F1 from 0.2317 to 0.4215 over monolithic LoRA and is preferred in 84.6\% of pairwise comparisons, while using fewer active parameters than separate specialist models. SVD-based ablation study shows that shared-specialist decomposition and orthogonal regularization reduce expert overlap. These results demonstrate that modular adaptation improves controllability and specificity in multimodal photographic critique.

51. 【2608.17490】When More Foundation Models Means Less: Diagnosing and Addressing Multi-View Fusion Failure

链接https://arxiv.org/abs/2608.17490

作者:Yibo Liu,Bowen Jiang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Foundation-model hubs turn, hubs turn multi-view, heterogeneous encoder pool, Foundation-model hubs, large heterogeneous encoder

备注: 26 pages, 4 figures. Code and results: [this https URL](https://github.com/yibol9768-alt/Quantifying-Representation-Reliability)

点击查看摘要

Abstract:Foundation-model hubs turn multi-view fusion into a selection problem: from a large heterogeneous encoder pool, which views should be fused, and how many? We show that downstream performance is non-monotonic in the number of fused encoders; later views can be redundant or task-misaligned, causing accuracy to saturate or decline. We formalise this setting as view-set composition and propose KAGES (Kernel-Alignment Greedy Encoder Selector), a label-aware method that orders frozen encoders by their marginal gain in centred kernel-target alignment. KAGES requires no downstream classifier training during selection, evaluates each candidate in $\mathcal{O}(n^2)$ time independent of encoder dimension, and admits a conditional $(1-e^{-\gamma})$ prefix-wise guarantee under monotonicity and a positive submodularity ratio. Across five recognition regimes and low-shot, larger-pool, and full-data protocols, KAGES improves average AULC over full fusion by 3.9, 5.8, and 3.3 points, respectively, and exceeds DPP and facility-location selection in average AULC. Image retrieval exhibits later, task-dependent saturation along the KAGES ordering, while peak-then-decline reproduces in frozen-LLM fusion. These results show that effective large-pool fusion depends on selecting a compact, task-aligned set of views rather than indiscriminately fusing more encoders.

52. 【2608.17487】NeuroPath: Brain-Inspired Dual-Pathway Graph Convolutional Networks for Skeleton-Based Action Recognition

链接https://arxiv.org/abs/2608.17487

作者:Kanglei Zhou,Ruizhi Cai,Hubert P. H. Shum,Frederick W. B. Li,Xiaohui Liang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Graph Convolutional Networks, Convolutional Networks, Graph Convolutional, aims to recognize, Spatial-Temporal Graph Convolutional

备注: Accepted to Pattern Recognition

点击查看摘要

Abstract:Skeleton-based action recognition aims to recognize human actions from sequences of human joint coordinates. Most existing Spatial-Temporal Graph Convolutional Networks (STGCNs) have achieved promising results by modeling skeletal structures with implicit spatial-temporal representations. However, our empirical study reveals a clear performance imbalance across different skeletal modalities, indicating that implicitly coupling spatial and temporal information limits the full exploitation of complementary structural and motion cues. Inspired by the ventral and dorsal pathways in human perception, we propose Dual-Pathway Graph Convolutional Networks (NeuroPath), which adopt a dual-pathway architecture for separate yet collaborative modeling of spatial and temporal information. Specifically, transformation units first convert the input into pathway-specific skeletal representations, allowing each pathway to focus on complementary aspects of human motion. To further capture coordinated joint behaviors and their interrelationships, we introduce a group graph convolution block that dynamically identifies key body parts and models their spatial-temporal dependencies. In addition, inter-pathway dynamic fusion modules integrate complementary inter-modal information across pathways, facilitating higher-level semantic interpretation of actions. Extensive experiments on Kinetics Skeleton 400, NTU RGB+D 60, and NTU RGB+D 120 demonstrate consistent performance improvements, validating the effectiveness of dual-pathway spatial-temporal modeling for skeleton-based action recognition.

53. 【2608.17475】S$^3$AM: A Single-Stream SAM with Reliability-Calibrated Frequency Adapter for Multi-modal Salient Object Detection

链接https://arxiv.org/abs/2608.17475

作者:Ruichao Hou,Boyue Xu,Tongwei Ren,Dongming Zhou,Gangshan Wu,Jinde Cao

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision foundation models, recently advanced multi-modal, advanced multi-modal salient, Vision foundation, recently advanced

备注

点击查看摘要

Abstract:Vision foundation models have recently advanced multi-modal salient object detection (MSOD) through parameter-efficient tuning and prompt learning. However, existing Segment Anything Model (SAM)-adapted MSOD methods often rely on dual-stream encoders or auxiliary prompt generators, leading to redundant computation. Although a single-stream alternative can reduce this cost, early fusion may also propagate noisy or misaligned auxiliary high-frequency cues through the backbone. In this paper, we propose a novel single-stream framework that integrates reliability-calibrated frequency adaptation into the adopted SAM backbone for MSOD. It avoids duplicated foundation backbones while explicitly controlling auxiliary frequency injection. Specifically, we design a mixture of frequency experts module, which uses the stationary wavelet transform to decompose each modality and aggregate cross-modal frequency information. We further introduce a reliability-calibrated frequency adapter with a dual-gate calibration mechanism, which selectively propagates the calibrated residual across transformer stages while jointly controlling its injection strength and cross-modal reliability. A hypernetwork-guided semantic-structural decoder then combines semantic mask features from the adopted backbone with Mamba-based structural detail recovery. Comprehensive experiments on RGB-D, RGB-T, and RGB-NIR salient object detection benchmarks validate that the proposed framework achieves competitive performance with only 12.20M trainable parameters, accounting for 5.4\% of the total parameters. The code will be available at this https URL.

54. 【2608.17447】NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting

链接https://arxiv.org/abs/2608.17447

作者:Hao Qin,Yukai Sun,Luyuan Chen,Mengxu Lu,Feng Zhang,Ming Kong,Zhenhong Du,Qiang Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, effective copyright protection, increasingly critical, rapid development, development and adoption

备注

点击查看摘要

Abstract:With the rapid development and adoption of 3D Gaussian Splatting (3DGS), the need for effective copyright protection has become increasingly critical. Existing watermarking techniques for 3DGS mainly focus on protecting rendered images via pre-trained decoders, leaving the underlying 3D Gaussian primitives vulnerable to misuse. In particular, they are ineffective against Partial Infringement, where an adversary extracts and reuses only a subset of Gaussians. In this paper, we propose NGS-Marker, a novel native watermarking framework for 3DGS. It integrates a jointly trained watermark injector and message decoder, and employs a gradientbased progressive injection strategy to ensure full-scene coverage. This enables robust ownership decoding from any local region. We further extend NGS-Marker with hybrid protection (combining native and indirect watermarks) and support for multimodal watermarking. Extensive experiments demonstrate that NGS-Marker effectively defends against partial infringement while offering practical flexibility for real-world deployment.

55. 【2608.17427】Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs

链接https://arxiv.org/abs/2608.17427

作者:Yifan Lu,Adinath Dukre,Abhijit Das,Ziyun Zou,Haolin Yang,Yutong Xie,Imran Razzak

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:visual question answering, generating clinically unsupported, clinically unsupported statements, Medical vision-language models, medical visual question

备注: Accepted by MICCAI 2026

点击查看摘要

Abstract:Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy-guided Spatial-Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier-free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC-CXR datasets across three Med-VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at this https URL.

56. 【2608.17426】SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

链接https://arxiv.org/abs/2608.17426

作者:Keyu Tu,Zhuowei Chen,Mengqi Huang,Yuxin Wang,Jiahao Zhu,Zhendong Mao,Yongdong Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Semantic Task Completion, Task Completion Video, Completion Video Generation, Task Completion, Completion Video

备注

点击查看摘要

Abstract:We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding. Semantic grounding characterizes the correspondence between the reference image and the generated outcome in terms of high-level semantics relevant to the task. Evaluation focuses on the generated outcome and requires neither the presentation of a complete sequence of intermediate task steps nor conventional appearance consistency with the reference image. To support systematic evaluation, we construct SemComp-Data, an evaluation dataset covering six domains. Each instance comprises a reference image, a detailed instruction, a brief instruction, and an outcome-centric video clip. A scalable four-stage curation pipeline converts raw videos into standardized SemComp-Data instances. We further introduce SemComp-Bench, an evaluation protocol that uses a vision-language model (VLM) to answer structured binary questions. SemComp-Bench reports the OA Score and the GR Score for Outcome Achievement and Generation Reliability, respectively. Experiments on representative video generation models show that achieving intended outcomes while maintaining task-relevant semantic grounding in reference images remains challenging.

57. 【2608.17425】GSToken: Geometry-Structured Gaussian Tokens for Compact 3D Medical Image Representation

链接https://arxiv.org/abs/2608.17425

作者:Xiaoduo Li,Quan Gu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:improving neural network, neural network accuracy, MRI is central, Effective segmentation, multi-modal MRI

备注

点击查看摘要

Abstract:Effective segmentation of multi-modal MRI is central to improving neural network accuracy in brain tumor recognition. Existing methods typically compress 3D volumes into token sequences via fixed patch encoding or learned attention pooling (e.g., TokenLearner). However, these compression schemes discard explicit spatial shape information; the resulting tokens convey no notion of lesion morphology or spatial extent. Meanwhile, end-to-end evaluation entangles a tokenizer's information retention with the reconstruction capacity of the downstream decoder, and the lack of a unified capacity contract across methods makes performance differences difficult to attribute. In this paper, we introduce Gaussian tokens to multi-modal brain tumor segmentation for the first time: each token carries not only a semantic feature but also a learned 3D center, anisotropic scale, and orientation, endowing the representation with explicit geometric support at negligible parameter cost. We further propose a frozen-token utility evaluation protocol: the trained tokenizer is frozen, its output is cast into a fixed-capacity serialized contract, and a shared lightweight Transformer probe independently measures each tokenizer's retained information under strictly matched conditions. Multi-seed paired statistical testing shows that GSToken consistently and substantially outperforms capacity-matched adaptive baselines under frozen probing, with uniform advantages across all tumor sub-regions, surface, and distance metrics. These results demonstrate that explicitly encoding spatial geometry within tokens significantly improves the information density of volumetric representations, offering a new design principle for compact 3D medical image representation and downstream reading.

58. 【2608.17422】F-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

链接https://arxiv.org/abs/2608.17422

作者:Yearang Lee,Ho-Joong Kim,Seong-Whan Lee

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Temporal Action Detection, Zero-Shot Temporal Action, recognize action instances, Action Detection, calize and recognize

备注

点击查看摘要

Abstract:Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing se- mantic distinctions between action classes, resulting in text- irrelevant predictions. To address this issue, we propose a Text-Foreground Concentrated Alignment for zero-shot temporal action DEtector (TF-CADE) that explicitly aligns textual information with action-relevant foreground regions. Specifically, we introduce Action Concentrate Aggregation (ACA), which extracts action concentrate scores to aggregate temporally informative video segments into a foreground- weighted video embedding. This foreground concentrated alignment enhances the semantic consistency between text and video features and improves inter-class discriminabil- ity. In addition, a Certainty-based Confidence Re-weighting (CCR) strategy refines per-snippet confidence scores by lever- aging foreground-aware similarity, effectively suppressing irrelevant action classes during inference. Extensive evalua- tions show that our TF-CADE not only achieves state-of-the- art performance under in-distribution settings but also excels in cross-dataset generalization to unseen action classes.

59. 【2608.17421】EAMS: Text-prompted spatiotEmporal dual-heAd Mamba Snake

链接https://arxiv.org/abs/2608.17421

作者:Ruicheng Zhang,Jianhui Lei,Kaiwen Shen,Haowei Guo,Jun Zhou,Bin Chen,Mengtang Li,Shen Zhao,Shuo Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:overcoming common pixel-level, common pixel-level misclassification, pixel-level misclassification issues, accurately predicts object-level, predicts object-level contours

备注: Medical Image Analysis (MedIA), 2026, In Press, Online Early Access Available

点击查看摘要

Abstract:Deep snake is a promising family of instance segmentation methods that accurately predicts object-level contours, thereby overcoming common pixel-level misclassification issues such as mask cavities and jagged edges in semantic segmentation approaches. However, existing deep snake methods face challenges in handling complex morphological variations, accurately capturing fine-grained organ details, and correcting base detection errors. To mitigate these limitations, we propose a cohesive Text-prompted spatiotEmporal dual-heAd Mamba Snake (TEAMS), a novel vision-language Mamba snake framework with three key innovations: (1) A Spatiotemporal Snake Evolution Strategy (SSES) is introduced to tackle complex morphological variations by capturing bidirectional spatial dependencies along the snake contour and temporal dynamics across evolution steps in a state space model. (2) A Contour Morphology-Aware Mamba (CMAM) is proposed to quantify local contour morphologies to modulate the structured attention mask in the Mamba2 SSD dual form, which extends Mamba's capability to perceive the relative importance of its input sequence elements for better delineation of fine-grained organ details. (3) A Text-prompted Collaborative Dual-Head Snake (TCDHS) is designed to incorporate cues from textual prompts and transfer the evolved contour information to the base detection head, which enhances the deep snake workflow and mitigates wrong detections. Comprehensive evaluations on five datasets covering different organs and imaging modalities demonstrate that TEAMS outperforms existing semantic and deep snake segmentation methods (e.g., relative mDice/mBF improvements of 6.9%/9.1% in a spinal dataset), underscoring its potential as a reliable tool across diverse medical image segmentation scenarios.

60. 【2608.17420】SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering

链接https://arxiv.org/abs/2608.17420

作者:Gen Li,Shu Han,Yun Xi Qiao,Hua Chen,Xuyang Dai,Bohan Li,Hao Zhao,Chaojian Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Gaussian Splatting, autonomous driving simulation, important component, component of autonomous, Driving scene reconstruction

备注: Project page: [this https URL](https://li00147.github.io/SPVC-Project-Page/)

点击查看摘要

Abstract:Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground-background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image-level novel-view repair or object-editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross-dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross-dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset-specific or scene-specific fixers. Concretely, we construct paired degraded-clean training data by simulating under-constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two-stage controllable video diffusion model that first addresses video-level appearance and then refines scene layout with structured controls.

61. 【2608.17415】Spectral Gradient Orthogonalization Improves Differentially Private Training at Scale

链接https://arxiv.org/abs/2608.17415

作者:Sabari Shanmugam,Nick Barnes,Kerry Taylor

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:adds isotropic Gaussian, isotropic Gaussian noise, Differentially private training, Differentially private, isotropic Gaussian

备注: Accepted at ECCV 2026

点击查看摘要

Abstract:Differentially private training adds isotropic Gaussian noise to clipped gradients, corrupting every singular direction equally. In vision models, where spatial correlation concentrates gradient energy into a low-rank subspace, most of this noise falls in directions that carry little signal. Spectral gradient orthogonalization via polar decomposition is introduced as a post-processing step that recovers directional signal from the noisy gradient's low-rank structure at zero additional privacy cost. A phase transition governs the utility of this approach: orthogonalization improves accuracy only when the per-direction spectral signal-to-noise ratio (SNR) suffices for singular vector recovery; in low-SNR regimes, the directional bias of the gradient is replaced by a nearly random orthogonal update, and the transformation is harmful. The recovery threshold is determined by the spectral gap of the gradient and is surpassed at large batch sizes. Empirically, the benefit scales with model capacity: spectral orthogonalization achieves a +20.9% improvement over DP-SGD on WRN-28-10 (B = 4096) and +14.9% on ResNet-18, while reducing inter-run variance by a factor of two to three. In the fine-tuning regime, spectral orthogonalization matches the stability of DP-Adam while maintaining a first-order memory footprint. Combining spectral with temporal denoising yields 50.3% on CIFAR-10 (epsilon = 4), the highest accuracy in any tested configuration. These gains are specific to moderate-to-high-SNR regimes such as large-batch training of higher-capacity models. Small-batch or low-SNR settings are better served by DP-SGD or temporal denoising.

62. 【2608.17414】REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

链接https://arxiv.org/abs/2608.17414

作者:Yuanbang Liu,Chenxi Ruan,Yihan Hou,Qiong Luo,Wei Zeng

类目:Computer Vision and Pattern Recognition (cs.CV); Programming Languages (cs.PL)

关键词:reference chart image, chart image based, challenging fine-grained visual, editing requires inferring, modifying visualization code

备注

点击查看摘要

Abstract:Chart editing requires inferring and modifying visualization code from a reference chart image based on an editing instruction, challenging fine-grained visual reasoning, instruction following, and executable code synthesis capabilities of MLLMs. Large reasoning models (LRMs) with extended Chain-of-Thought (CoT) reasoning are suitable for tackling such complex multimodal tasks. However, our preliminary study reveals an ``inverted-U'' relationship between reasoning length and chart-editing performance: Excessive reasoning often leads to ``overthinking,'' where models drift toward hallucinated visual details or get stuck in redundant reasoning loops. To address the gap, we introduce REChart, a two-stage training framework that provides process-level supervision over intermediate reasoning steps, improving both editing fidelity and reasoning efficiency. First, we synthesize 200k high-quality reasoning trajectories for supervised fine-tuning from a large image-instruction-code pool, using a role-specialized agentic Reason-Score-Refine workflow that iteratively refine the chart code toward higher quality. Second, we optimize the model via reinforcement learning with two complementary rewards: a \emph{fidelity} reward evaluating code correctness, visual fidelity, and structural consistency, and an \emph{efficiency} reward that assigns each rollout a random thinking budget, truncates the reasoning process, and credits the final reasoning segment according to its contribution to the output. On the ChartEdit and ChartMIMIC benchmarks, our model achieves state-of-the-art chart-editing performance among open-source models of comparable scale, while mitigating overthinking and reducing average reasoning token usage by 79.0\% under a maximum thinking budget of 16,384 tokens compared with the base model.

63. 【2608.17402】MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

链接https://arxiv.org/abs/2608.17402

作者:Bonan Zhang,Shiyu Dong,Quan Hung Tran,Katharina Gschwind,Shuqi Yang,Sijia Chen,Adel Ahmadyan,Seungwhan Moon,Lu Zhang,Ahmed Kirmani,Babak Damavandi,Anuj Kumar

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:capacity effectively improves, effectively improves performance, critical component, component of vision-language, capacity effectively

备注: Accepted to ECCV 2026

点击查看摘要

Abstract:Vision encoders are a critical component of vision-language models, and scaling their capacity effectively improves performance. However, dense scaling increases compute cost and inference latency. Mixture-of-Experts (MoE) architectures offer a compelling alternative, having enabled efficient scaling in LLMs, yet the MoE design space for CLIP-style vision encoders remains underexplored at State-of-the-Art (SOTA) levels. In this work, we systematically study MoE designs for vision encoder scaling and find that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts. We further propose an auxiliary-loss-free balancing variant for better expert utilization, and design a specialized MoE kernel to mitigate inference latency overhead. To enhance video capabilities while preserving image knowledge, we introduce frame-level distillation paired with a novel freezing mechanism. We pretrain a series of Mixture-of-Experts Vision Encoders (MoE-ViE) across a range of sizes, all consistently outperforming their dense counterparts. Our largest model matches the zero-shot performance of a SOTA encoder 1.7x its size at 76% of its latency. When aligned with an LLM, MoE-ViE surpasses all compared encoders on image and video benchmarks, including those with up to 5x more activated parameters. Code is available at this https URL.

64. 【2608.17398】o Remove or Not to Remove Clouds: A Comparative Analysis and Fusion of Raw SAR and Synthetic NDWI for Overcast Water Segmentation

链接https://arxiv.org/abs/2608.17398

作者:Saleh Sakib Ahmed,Sara Nowreen,M. Sohel Rahman

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Persistent clouds blind, Persistent clouds, blind optical satellites, satellites during floods, raw SAR

备注

点击查看摘要

Abstract:Persistent clouds blind optical satellites during floods. While Synthetic Aperture Radar (SAR) penetrates clouds, its raw data is noisy and lacks clear contrast. To mitigate this, recent studies utilize deep learning models to translate SAR into cloud-free synthetic optical imagery for downstream tasks like water body segmentation. However, because raw SAR is the original source for both of these operations, a critical methodological dilemma arises: during complete overcast should segmentation models process the raw SAR directly, or rely on a translated synthetic Normalized Difference Water Index (NDWI) proxy? This study resolves the debate by demonstrating that synthetic NDWI yields better results, as the translation process acts as a powerful filter against radar noise. This raises a natural second question: what if we utilize both? Building on our findings, we introduce a Combined Framework that integrates both raw SAR and synthetic NDWI into a unified model. By fusing the sharp physical boundaries of raw SAR with the high contrast of synthetic NDWI, this hybrid approach consistently outperforms all standalone methods.

65. 【2608.17394】Noisy group neurons with synchronous resetting for high-performance spiking neural networks

链接https://arxiv.org/abs/2608.17394

作者:Yajie Zhai,Yanmei Kang,Meng Li,Zigang Huang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Spiking neural networks, attained significant progress, bio-inspired neuronal dynamics, Spiking neural, characterized by bio-inspired

备注

点击查看摘要

Abstract:Spiking neural networks (SNNs), characterized by bio-inspired neuronal dynamics and event-driven communication, have attained significant progress in recent years. Nevertheless, training deep SNNs remains challenging due to spatiotemporal information loss and gradient mismatching. To simultaneously address these issues, we propose a noisy group neuron (NGN) model, which incorporates population-level synchronous resetting and neural stochasticity as fundamental computational mechanisms. We then develop the NGN method as a framework that combines the NGN model with backpropagation learning based on mean-field dynamics. We demonstrate the advantages of the NGN method through theoretical analysis and experimental validation on CIFAR-10, CIFAR-100, Tiny-ImageNet, DVS-Gesture, N-Caltech101, and CIFAR10-DVS. The proposed approach achieves an accuracy of 87.35% on CIFAR10-DVS within 10 inference time steps. These results support NGN as a practical approach to high-performance neuromorphic computing.

66. 【2608.17389】GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly

链接https://arxiv.org/abs/2608.17389

作者:Tinghao Jiang,Sheng Tang,Shengzhe Wei,Juntong Fang,Weiqi Zhang,Junsheng Zhou,Zesong Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:RGB videos requires, reconstruction from RGB, RGB videos, consistent camera motion, globally consistent camera

备注: 15 pages, including supplementary material; 7 figures and 10 tables. Project page: [this https URL](https://kosmoresearch.github.io/GeoWeaver/)

点击查看摘要

Abstract:Long-sequence 3D reconstruction from RGB videos requires both accurate local geometry and globally consistent camera motion. Feed-forward models provide strong depth and pose predictions, but their memory cost prevents joint inference over long sequences. Chunk-wise processing improves scalability, yet independently predicted chunks often exhibit scale drift, pose errors, and point-cloud misalignment. We present GeoWeaver, a unified framework comprising a Geometric Prior Model (GPM) and Test-Time Adaptation (TTA). The GPM predicts chunk-wise depth, confidence, and camera parameters as adjustable geometric priors. TTA then performs sequential initialization, global chunk-level Sim(3) alignment, and coarse-to-fine refinement of camera poses, affine depth corrections, and intrinsics. Dense correspondences provide adjacent, cross-chunk, and long-range constraints, while a robust CDF-style objective jointly optimizes weighted 2D reprojection and 3D consistency residuals. This design preserves local geometric accuracy while correcting accumulated pose, scale, depth, and calibration errors. Experiments across diverse long-sequence benchmarks demonstrate improved camera accuracy, global consistency, and point-cloud quality. Ablations verify the contribution of each adaptation stage, and applying the same TTA procedure to different geometric prior models consistently improves their trajectory estimates, demonstrating that GeoWeaver is not tied to a specific GPM.

67. 【2608.17362】Continuity-Driven Representation Learning for Industrial Defect Detection

链接https://arxiv.org/abs/2608.17362

作者:Minjong Kim,Hyun Jun Kim,Jeongrae Kim,Heeseung Shin,Changwon Lim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:percentage points, natural-image object detection, defect detection differs, large normal-dominant regions, repetitive structures

备注: Accepted at the British Machine Vision Conference (BMVC) 2026

点击查看摘要

Abstract:Industrial defect detection differs from natural-image object detection because inspection images are captured under controlled conditions and contain large normal-dominant regions with repetitive structures. Defects therefore appear as localized disruptions of otherwise predictable patterns, while conventional detectors rely mainly on sparse bounding-box supervision, resulting in weakly constrained normal-region representations. We propose a continuity-driven representation regularization framework that exploits normal-dominant regions as dense auxiliary supervision. The framework introduces two detector-agnostic objectives: Multi-Continuity Loss, which combines 1D patch-sequence prediction and 2D masked spatial prediction, and Differencing Loss, which regularizes first-order feature variation and second-order curvature between neighboring patch embeddings. Both objectives are applied with box-derived region weighting to stabilize normal-region representations while preserving defect-related discontinuities. Experiments on two real-world industrial datasets and the public NEU-DET benchmark, using six detector architectures including YOLO-family models, MambaYOLO, and DETR, demonstrate consistent improvements over native detector baselines. In the full-data setting, the proposed regularizers improve average mAP@0.5:0.95 by up to 3.49 percentage points on Industrial Metal, 5.38 percentage points on MEA, and 5.03 percentage points on NEU-DET. Under limited-data conditions, the gains become more pronounced, with Differencing Loss achieving improvements of up to 21.07 percentage points in mAP@0.5 and 8.23 percentage points in mAP@0.5:0.95 on NEU-DET using only 25% of the training data. These results suggest that continuity-driven regularization provides an effective prior for improving industrial defect detection, particularly when annotated data are scarce.

Comments:
Accepted at the British Machine Vision Conference (BMVC) 2026

Subjects:

Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2608.17362 [cs.CV]

(or
arXiv:2608.17362v1 [cs.CV] for this version)

https://doi.org/10.48550/arXiv.2608.17362

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
68. 【2608.17351】Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

链接https://arxiv.org/abs/2608.17351

作者:Fangling Jiang,Qi Li,Bing Liu,Weining Wang,Quilin Huang,Zhenan Sun,Ming-Hsuan Yang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:target domains differ, target domains, diverse attack types, attack types absent, Open-world face anti-spoofing

备注

点击查看摘要

Abstract:Open-world face anti-spoofing must address both covariate and semantic shifts: source and target domains differ in imaging conditions, while target domains contain diverse attack types absent from training. Existing prompt-based approaches often express spoofing through category semantics or language guidance, which is effective for modeling high-level concepts but is less suited to explicitly capturing the evolving fine-grained and spatially heterogeneous forensic evidence of unseen attacks. Motivated by the hypothesis that many unseen attacks can be characterized by new combinations of recurring visual cues, we propose a compositional forensic visual prompt learning framework that operates entirely in the visual feature this http URL on a frozen ViT-based vision foundation model, the framework employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches. Class-specific global contextual prompts then provide input-dependent routing weights that adaptively select and compose these primitives into compositional forensic visual prompts for real/spoof discrimination. The primitives are not assigned predefined semantic meanings; instead, their specialization and reuse emerge from shared parameterization and joint optimization across this http URL experiments on nine open-world protocols demonstrate state-of-the-art performance, strong cross-domain generalization, and robust adaptation to unseen attacks.

69. 【2608.17337】Learning latent progression states from spatial heterogeneity in uterine histopathology

链接https://arxiv.org/abs/2608.17337

作者:Qiming He,Yan Liu,Shuang Ge,Fan Yang,Yuxiang Wang,Ieng Man Zhang,Jing Yang,Zihao Jia,Ajin Hu,Yexing Zhang,Zixiu Song,Qiang Huang,Xiaoya Zhao,Zihan Wang,Xianjing Zheng,Yijun Zheng,Liling Lin,Shuxing Liu,Bin Bao,Yue Xie,Tian Guan,Yonghong He,Congrong Liu

类目:Computer Vision and Pattern Recognition (cs.CV); Emerging Technologies (cs.ET)

关键词:static diagnostic categories, progression is accompanied, compressed into static, progression-associated tumor states, Tumor progression

备注

点击查看摘要

Abstract:Tumor progression is accompanied by changes in architecture, morphology and microenvironmental organization, yet progression-associated heterogeneity is usually compressed into static diagnostic categories in histopathology. Here we present SpaTIE, a uterus-specific computational pathology framework that learns morphology-aware representations and organizes spatial histopathological heterogeneity into progression-associated tumor states. SpaTIE was developed using 10,426 uterine hematoxylin and eosin whole-slide images and evaluated in TCGA-UCEC and TCGA-UCS cohorts. The learned representations formed morphology manifolds, supported diagnostic, molecular and survival-related prediction tasks, and localized attention to informative tumor regions. Beyond supervised prediction, SpaTIE inferred tumor-state axes from cross-sectional morphology without temporal or molecular supervision. These morphology-derived states were spatially coherent and showed associations with clinicopathological variables and survival outcomes, while not simply recapitulating staging or diagnostic labels. Integrative multi-omics analyses linked the inferred states to DNA methylation, somatic copy-number variation, mutation, RNA-seq and RPPA profiles, highlighting molecular programs related to chromatin regulation, copy-number-associated structural variation, receptor tyrosine kinase signaling, cell adhesion, extracellular-matrix remodeling and metabolic adaptation. Progression-guided virtual perturbation further prioritized molecular features coupled to the morphology-derived state organization. Together, these findings suggest that uterine histopathology contains recoverable progression-associated tumor-state information and establish SpaTIE as a framework for connecting spatial morphology with multi-omics-informed tumor-state discovery.

70. 【2608.17328】MS-MFAD : Multimodal large language models for Face Anti-spoofing Detection

链接https://arxiv.org/abs/2608.17328

作者:Xiaoyong Yu,Rongzhen Li,Shuming Shi,Xinge You

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Facial biometric recognition, high-fidelity physical spoofing, Multimodal Large Language, compound threats intertwining, threats intertwining generative

备注

点击查看摘要

Abstract:Facial biometric recognition systems currently face compound threats intertwining generative AI and high-fidelity physical spoofing. Existing defenses suffer from systemic bottlenecks, including poor generalization, non-auditable reasoning, and reliance on massive, low-quality datasets. To address these challenges, we propose Multimodal Large Language Models (MFAD) for face anti-spoofing detection, an explainable reasoning system for Unified Face Anti-Spoofing Detection (UFAD), accompanied by a semantic-level annotation benchmark. Unlike methods relying on external tools or coarse alignment, MFAD activates the intrinsic reasoning capabilities of Multimodal Large Language Models (MLLMs) via a fine-grained pixel-semantic anchoring mechanism. This eliminates localization hallucinations and ensures auditable reasoning paths. We introduce a cross-attack semantic-level unified annotation paradigm: by annotating only 1,000 precise masks per attack category, we generate reasoning evidence chains strictly corresponding to spoofed regions. Supervised fine-tuning on the Qwen-VL foundation model demonstrates that, using limited high-quality samples, the system achieves a 40-50% relative reduction in in-domain ACER and restricts cross-domain performance degradation to within 11.62%/5.23%, significantly outperforming existing frameworks. Furthermore, under white-box adversarial attacks, detection accuracy drops by only 3.2%, validating the robustness of semantic anchoring compared to models trained on massive short-text data. Domain practitioners rated the evidence reliability of reasoning paths at 4.57/5, with inference latency satisfying real-time deployment requirements. These results confirm that a few-shot, high-quality semantic annotation paradigm is effective for building trustworthy, explainable, and cost-efficient UFAD systems.

71. 【2608.17318】If, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation

链接https://arxiv.org/abs/2608.17318

作者:Seoyoung Lee,Neel P. Bhatt,Pranay Samineni,Cong Liu,S P Sharan,Timothy Barclay,Gregory M. Wagner,Daniel Milan,Sandeep Chinchali,Ufuk Topcu,Atlas Wang

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:fixed goal, navigation, branch, agents, instructions

备注: 11 pages, 1 figure, 3 tables. Project page: [this https URL](https://condvln.github.io/)

点击查看摘要

Abstract:Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on observed states of the environment: if a condition holds, then follow one path, otherwise take another. Such instructions require an agent to evaluate scene evidence, select the correct logical branch, and execute the corresponding navigation behavior. Existing evaluations provide limited control over conditional branch execution, making it difficult to determine whether agents fail because of perception, grounding, navigation, or logical decision-making. We introduce CondVLN, a scene-graph-grounded benchmark for diagnosing conditional branching in vision-language navigation. CondVLN programmatically generates instructions whose branch conditions are grounded in verifiable 3D scene-graph predicates, with controlled variation in branch depth, dependency chain length, spatial composition, evidence observability, and instruction horizon. CondVLN contains over 11,500 generated conditional instructions across AI2-THOR, Matterport3D, Gibson, and ReplicaCAD, and evaluates agents using standard VLN metrics and branch-specific diagnostics: Branch Selection Accuracy and Conditional Success Rate. Evaluating four state-of-the-art VLN agents (VLN-Zero, NaVid, NaVILA, and Open-Nav) shows that conditional branching exposes failures that are not captured by standard success rate or path length alone: agents can navigate plausibly while committing to a branch inconsistent with the observed scene condition. We also present a lightweight neurosymbolic branch-selection model that separates condition grounding from navigation execution, improving performance by 2x. CondVLN provides a reusable testbed for measuring whether embodied agents can not only follow instructions, but follow the right instruction under the right condition.

72. 【2608.17314】Scanline-Aware Animatable Gaussian Avatars from Rolling-Shutter Videos

链接https://arxiv.org/abs/2608.17314

作者:Youxiang Wang

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Animatable human avatars, silent assumption, routinely reconstructed, Animatable human, frame observes

备注

点击查看摘要

Abstract:Animatable human avatars are routinely reconstructed from multi-view video under a silent assumption: that every pixel of a frame observes the same instant of the body's motion. Rolling-shutter (RS) sensors expose image rows sequentially, so within one frame the head and the feet of a moving person are separated by tens of milliseconds of articulated motion, and every scanline sees a different pose. Feeding such video to a state-of-the-art avatar bakes the distortion into the canonical representation, where it survives as shear and wobble under novel views and novel poses. Worse, every camera in a rig follows its own readout schedule, so the multi-view consistency that drives the reconstruction is violated even when the geometry is correct. We present RS-Avatar, which reconstructs a sharp, undistorted, animatable 3D Gaussian avatar directly from RS video. The formulation is minimal: a motion-aware avatar already renders the body at several sub-frame instants, and where a blur model averages those renderings, a rolling-shutter model composites them scanline by scanline. Changing that operator is the only modification required. On RS-ZJU, a benchmark we build from ZJU-MoCap, this improves novel-view synthesis over training as if the frames were instantaneous, on every subject. A motion-aware blur model built on the same sub-frame machinery does not transfer, and in fact falls below the shutter-oblivious baseline: the machinery is reusable, the operator is not.

73. 【2608.17306】Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

链接https://arxiv.org/abs/2608.17306

作者:Yang Chen,Zhan Zhuang,Yanbin Wei,Zebin Chen,Hua Liu,Yu Zhang

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:existing methods aggravate, vision-language models efficiently, methods aggravate robust, adversarial prompt tuning, Disentangled Prompt Tuning

备注

点击查看摘要

Abstract:While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robust generalization overfitting on seen classes, leading to a rapid degradation in performance against adversarial examples of unseen classes as training progresses. We empirically identify that this degradation stems from the tendency of the model to learn pseudo-robust features (i.e., non-generalizable shortcuts). To mitigate this, we propose ADAPT (Adversarial Disentangled Prompt Tuning), a robust prompt tuning framework following the philosophy of ``Learning What Not to Learn''. Specifically, ADAPT uses a dual-prompt mechanism with a target prompt and a pool of decoy prompts. During training, the decoy prompts are guided to entrap diverse pseudo-robust features, while the target prompt is constrained to be orthogonal to the decoys in the embedding space to learn robust features. By disentangling the robust features from the pseudo-robust features, ADAPT effectively prevents robust generalization overfitting. We further provide an analysis showing that the orthogonal loss bounds the effect of shifts in pseudo-robust features on unseen classes, yielding a testing error guarantee. Empirically, extensive experiments demonstrate that ADAPT substantially improves the robustness of the target prompt on unseen classes. The code is available at this https URL.

74. 【2608.17298】3D Gaussian Accelerated Ray Tracing: Fast training through particle-based backward propagation

链接https://arxiv.org/abs/2608.17298

作者:Laurent Vit,Oliver Batchelor,Richard Green

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:highly efficient representation, limit accurate view-dependent, accurate view-dependent ordering, Splatting has made, secondary ray effects

备注

点击查看摘要

Abstract:3D Gaussian Splatting has made Gaussian primitives a highly efficient representation for real-time novel view synthesis, but its rasterisation-based formulation relies on screen-space approximations that limit accurate view-dependent ordering and the integration of secondary ray effects such as reflections, refractions, and shadows. Gaussian ray tracing addresses these limitations by evaluating explicit ray-primitive intersections, yet it remains costly to train. We observe that the main bottleneck is not ray traversal alone, but the pixel-centric backward propagation, where many threads concurrently accumulate gradients into the same primitive parameters, causing severe atomic contention and thread serialisation. We present 3DGART, a practical training framework for ray-traced Gaussian rendering. Our key idea is to reorganise backward propagation around primitives rather than pixels. Using conservative perspective-correct screen-space bounds, we build a compact intermediate buffer and a tile-primitive mapping that allows each thread to accumulate the contribution of one primitive over its covered pixels within a tile. This transforms gradient computation from a contention-heavy scatter operation into a structured gather-like process. On Mip-NeRF 360, 3DGART achieves an $\approx 3-3.5\times$ raw training speedup over per-pixel baseline and $\approx4 \times$ over 3DGRT on Mip-NeRF 360 while improving quality. More importantly, 3DGART makes fully ray-traced Gaussian training practical, reaching runtimes competitive with rasterisation-based pipelines while preserving benefits of ray tracing.

Subjects:

Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

Cite as:
arXiv:2608.17298 [cs.GR]

(or
arXiv:2608.17298v1 [cs.GR] for this version)

https://doi.org/10.48550/arXiv.2608.17298

Focus to learn more

              arXiv-issued DOI via DataCite (pending registration)</p>
75. 【2608.17291】B-Spline Embedded Structure Learning for 3D Tooth Segmentation

链接https://arxiv.org/abs/2608.17291

作者:Xianghan Wei,Jianwen Lou,Zhiguo Lu,Hairong Jin,Haihua Zhu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:formidable challenge due, high morphological similarity, tooth segmentation forms, Embedded Structure Learning, B-Spline Embedded Structure

备注

点击查看摘要

Abstract:Accurate 3D tooth segmentation forms the cornerstone of digital dentistry, yet it remains a formidable challenge due to the inherent intricacy of real-world dentitions, such as crowding, misaligned teeth and high morphological similarity between adjacent teeth. To resolve this, we present B-Spline Embedded Structure Learning, a novel framework that distills the inherent sequential arrangement of teeth into a continuous structural constraint to regularize representation space. Our approach parameterizes the global dental topology by fitting a parametric B-spline trajectory to tooth centers, assigning each point a continuous structural embedding that forces the shared backbone to capture global arch organization. To fully exploit these embedded priors, we introduce a Structure-Aware Dynamic Classifier (SADC) to substitute rigid static templates with adaptive, case-calibrated decision boundaries. SADC regularizes dynamic prototype pooling via a localized Gaussian proximity gate and contextually co-evolves them through an attention block modeling spatial relations and bilateral symmetries across teeth. Extensive evaluations on the 3DTeethSeg22 benchmark demonstrate that our method establishes a new state-of-the-art accuracy with exceptional structural robustness and efficiency in computational overhead, markedly enhancing the model's capacity to handle complex dental configurations.

76. 【2608.17283】UniQuery4R: Unified 4D Scene Reconstruction from a Single Query

链接https://arxiv.org/abs/2608.17283

作者:Tiancheng Chen,Sheng Tang,Wenhua Jin,Weiqi Zhang,Juntong Fang,Junsheng Zhou,Zesong Li

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:Reconstructing dynamic, requires jointly estimating, object motion, jointly estimating correspondence, scenes requires jointly

备注: 16 pages, 8 figures. Project page: [this https URL](https://kosmoresearch.github.io/UniQuery4R/)

点击查看摘要

Abstract:Reconstructing dynamic 4D scenes requires jointly estimating correspondence, geometry, object motion, and camera motion. Existing feed-forward methods typically predict dense task-specific maps or independently process source-target pairs, leading to unnecessary computation for sparse queries and limited feature reuse across different frame pairs. We present UniQuery4R, a query-conditioned framework that encodes a multi-frame clip once and selects the source view, target view, and continuous source-image coordinate only at decoding time via source-to-target cross-attention. Each query jointly predicts target correspondence, target-time 3D position, and scene flow, along with source depth, while camera parameters are estimated per view. This design allows the encoded clip to be reused across arbitrary source-target selections and supports both sparse inference and dense reconstruction through batched queries, without learned temporal embeddings tied to a fixed clip length. We further introduce a direction-magnitude parameterization of scene flow with separate supervision for moving and static points. Among the evaluated methods, UniQuery4R achieves the best macro-average results on WorldTrack for both scene-flow estimation and dynamic-point reconstruction.

77. 【2608.17279】Key-Frame Reasoning with SAM3: Third Place Solution for the MeViS-Text Track of the 8th LSVOS Challenge

链接https://arxiv.org/abs/2608.17279

作者:Ce Bian,Xusheng He,Jinrong Zhang,Canyang Wu,Xianjing Han,Jianlong Wu

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:presents a two-stage, training-free solution, LSVOS Challenge, report presents, MeViS-Text track

备注

点击查看摘要

Abstract:This report presents a two-stage, training-free solution for the MeViS-Text track of the 8th LSVOS Challenge. The task requires a model to localize and segment the object specified by a natural-language expression throughout a video. Such expressions often depend on temporal cues, including actions, interactions, directions, and relative positions. Our first stage uses Gemini-3.1 Pro via API to decompose a video-level event into instance-level targets, select a key frame for each target, and generate a discriminative description aligned with that frame. In the second stage, SAM3-agent produces a pixel-level seed mask on the selected frame, and the SAM3 video tracker propagates the mask bidirectionally through the video. Valid instances are grounded and propagated independently before their frame-wise masks are merged. All local SAM3 processing runs on a single NVIDIA GeForce RTX 4090 without task-specific training or model ensembling. Our method ranked third on the challenge test set, obtaining JF, J, F, N-acc., T-acc., and Final scores of 0.761, 0.7367, 0.7852, 0.8333, 0.9755, and 0.856593, respectively.

78. 【2608.17255】Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

链接https://arxiv.org/abs/2608.17255

作者:Yifei Wu,Yicheng Wu,Qiang Ma,Qi Chen,Renyang Gu,Xinyu Liu,Yongsheng Pan,Yong Xia

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:volumetric attenuation field, underlying volumetric attenuation, attenuation field, accumulated attenuation, ray path

备注

点击查看摘要

Abstract:X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumulated attenuation along a corresponding ray path. Reconstructing a CT volume from only a few X-ray views is therefore severely ill-posed, as the projections collapse depth information and leave 3D locations of anatomical regions and their corresponding intensity distributions highly entangled and ambiguous. We observe that once the spatial organization of anatomical regions is established, estimating their CT intensities becomes substantially more tractable. Motivated by this, we propose LiftXR, an interleaved, geometry-guided framework that explicitly incorporates spatial layout recovery into CT reconstruction. Specifically, a layout lifter first generates a 3D anatomical layout from bi-planar X-rays, providing spatial guidance for an intensity renderer to reconstruct a CT volume. An anatomical parser then performs volumetric perception on the reconstruction, exploiting its spatially resolved boundary and intensity cues to recover a refined anatomical layout. This transition from projection-conditioned layout generation to reconstruction-conditioned anatomical perception allows the parsed layout to provide feedback for region-specific intensity calibration. Extensive experiments on two public datasets demonstrate that LiftXR consistently outperforms recent X-ray-to-CT reconstruction methods, establishing a new state of the art. Moreover, the reconstructed CT achieves superior performance in external downstream segmentation, indicating improved anatomical fidelity. Code will be released.

79. 【2608.17254】Heterogeneity-Aware Deep Learning for Tumour Classification from Multiparametric MRI

链接https://arxiv.org/abs/2608.17254

作者:Yue Xia,Euijoon Ahn,Tian Xia,Yuan Yuan,Michael Fulham,Jinman Kim

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:reflects spatial variation, deep learning, Intra-tumoural heterogeneity, Deep Learning Classification, reflects spatial

备注

点击查看摘要

Abstract:Intra-tumoural heterogeneity (ITH) reflects spatial variation in tumour biology and is an important determinant of tumour behaviour, prognosis, and treatment response. Radiomics and deep learning have shown promise for tumour classification from multiparametric MRI (mp-MRI), but radiomics relies on handcrafted features, while most deep learning methods use whole-tumour representations or manually defined sub-regions, limiting scalable modelling of tumour heterogeneity. We propose a Heterogeneity-Aware Deep Learning Classification (HA-DLC) framework that explicitly models imaging-derived tumour sub-regions for lesion-type diagnosis and molecular-status prediction. HA-DLC consists of: (1) a Heterogeneous Sub-region Generation (HSG) module that produces initial pseudo-labelled sub-regions via unsupervised clustering, followed by Cross-Patient Sub-region Alignment (CPSA), which maps cluster-derived regions to a shared label space using soft assignments; and (2) a Dual-Stream Feature Extraction (DSFE) module that integrates local heterogeneity-aware features with global tumour representations. Given the initial clustering masks, CPSA, segmentation, feature extraction, and classification are jointly optimized end-to-end using soft-target segmentation and classification objectives. We evaluate HA-DLC on the LLD-MMRI2023 liver lesion dataset and the RSNA-ASNR-MICCAI 2021 Radiogenomic Brain Tumour dataset. HA-DLC consistently outperforms state-of-the-art radiomics and deep learning baselines, demonstrating the value of cross-patient sub-region alignment and dual-stream heterogeneity modelling for tumour classification from mp-MRI.

80. 【2608.17253】Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

链接https://arxiv.org/abs/2608.17253

作者:Yunhao Yang,Yuexin Bian,Yunjie Tian,Di Fu,Tianjin Huang,Yuanyuan Shi,Ziang Xiao,Nuno Vasconcelos,Yijiang Li

类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)

关键词:Reinforcement learning, powerful approach, approach for improving, language and vision-language, strongest successes

备注: 30 pages, 5 figures, 11 tables

点击查看摘要

Abstract:Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward). Such annotations are costly to obtain and become increasingly scarce as reasoning capabilities advance beyond what humans can reliably evaluate. Self-rewarding RL reduces this dependence by enabling models to derive reward signals from their own completions. However, training solely on self-generated feedback can reinforce existing biases and suboptimal behaviors, reduce response diversity, and ultimately lead to homogenized responses and training collapse. In this work, we show that unsupervised reasoning can emerge through cooperative multi-agent training. We introduce Co-RL, a framework in which multiple decoupled models, sharing no parameters, are simultaneously optimized through RL using rewards derived from their peers. We further show that increasing cohort diversity, through heterogeneous model families, sizes, and rephrased training samples, reduces the correlated errors that drive self-reinforcing feedback loops. This diversity consistently improves reasoning performance, maintains behavioral diversity, and mitigates training collapse. Across text-only and multimodal domains, Co-RL consistently outperforms the base models and prior label-free approaches, while matching or surpassing supervised methods, without access to any ground-truth labels. Concretely, Co-RL yields average gains of 3.0-8.6% across seven text-only benchmarks for LLMs and 2.3-7.2% across four multimodal benchmarks for VLMs. Code is available at this https URL.

81. 【2608.17237】Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement

链接https://arxiv.org/abs/2608.17237

作者:Mohammad Talebi-Kalaleh,Qipei Mei

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:drafts remains labor-intensive, finite-element model drafts, model drafts remains, Converting structural framing, Converting structural

备注

点击查看摘要

Abstract:Converting structural framing plans into editable finite-element model drafts remains labor-intensive and prone to transcription error. Existing drawing-understanding systems for building components rely on task-specific trained neural detectors, and language-model agents in structural engineering operate on text or model data rather than the drawing itself. This paper presents, to the authors' knowledge, the first framework applying an agentic vision-language layer to structural component detection and model drafting from framing-plan PDFs, without task-specific detector training or fine-tuning. A deterministic stage extracts primitives, estimates scale by dimension-ratio consensus, recognizes five entity classes with a drafting grammar, and assembles an editable layout. The agentic stage proposes typed corrections constrained by deterministic candidates, operation-specific admission tests, change-level review, and fail-closed transactions. Evaluation used an author-generated benchmark of 100 plans: a development half that informed every rule revision, and a seed-disjoint held-out half generated after the rules froze, evaluated once. All reported scores are end-to-end results of the complete framework on the held-out half. Scale was estimated within 0.1% of the generator reference for every drawing. Recall and precision were 0.922/0.997 for columns, 0.886/0.990 for beams, 1.000/1.000 for walls, 1.000/1.000 for braces, and 1.000/0.964 for openings. A controlled study repeated two corruptions three times on three development drawings. Calibration passed all nine trials; member repair met every strict end-state predicate in five of nine. Guarded review corrected missed framing and false marks within explicit bounds. The held-out half shares the development generator, so the study excludes independently drafted plans, raster evaluation, analytical connectivity, and solver validation.

82. 【2608.17224】Probing Association Instability with Track-State Perturbations for Clip-Level Active Learning in Query-Propagation Multi-Object Tracking

链接https://arxiv.org/abs/2608.17224

作者:Riku Inoue,Shogo Sato,Kazuhiko Murasaki,Tomoyasu Shimada,Toshihiko Nishimura,Ryuichi Tanida

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:dataset construction expensive, models requires dense, making dataset construction, requires dense bounding-box, Training query-propagation

备注: Accepted at the 37th British Machine Vision Conference (BMVC 2026)

点击查看摘要

Abstract:Training query-propagation end-to-end multi-object tracking (MOT) models requires dense bounding-box and identity annotations across video sequences, making dataset construction expensive. Clip-level active learning reduces this cost by selecting video clips for annotation, but prior acquisition criteria based on output-level temporal uncertainty may miss clips whose informativeness comes from association instability in propagated track states. We propose QPID (Query-Propagation Instability and Diversity), a clip acquisition method for query-propagation MOT that targets association instability in propagated track states. QPID estimates this instability by applying two-sided perturbations to internal track states and measuring prediction differences from a clean reference branch. The key idea is that, in stable clips, each propagated track should continue to follow the same target under small perturbations, whereas in ambiguous clips, small changes in the track state can alter which target the track follows, leading to changes in localization or confidence. QPID measures these perturbation-induced prediction differences with two metrics: Localization Drift and Entropy-Weighted Confidence Discrepancy. These metrics are aggregated into a clip-level association-instability score. To avoid redundant uncertainty-only selection, QPID selects a representative annotation batch from high-instability clips using Uncertainty-Weighted Visual Coverage with track-level visual prototypes. Experiments on DanceTrack and SportsMOT with MeMOTR and SambaMOTR show that QPID achieves strong performance compared with active learning baselines under the same annotation budget.

83. 【2608.17209】ach and Grow: An Agent-Centered Architecture for General Robot Learning

链接https://arxiv.org/abs/2608.17209

作者:Chang Nie,Zhe Liu,Hesheng Wang

类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:world-action models offer, general-purpose robotics, world-action models, models offer, offer an elegant

备注

点击查看摘要

Abstract:End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability is bounded by validated physical coverage. When an unfamiliar object, sensor, embodiment, or contact falls outside that coverage and no validated fallback exists, correcting the failure requires new robot data, a policy update, and regression testing. This recurring burden is the retraining tax. Unlike text, embodied data must often be created by operating machines. We present Teach-and-Grow Learning (TGL), an agent-centered architecture for general robot learning. In its general form, a multimodal agent turns a few successful demonstrations into reusable Skill Blocks: closed-loop behaviors for meaningful subgoals. In a new scene, the agent grounds and composes these blocks, selects learned or geometric tools, observes the physical outcome, and revises the route when execution departs from intent. A Skill Library stores executable behavior, while structured Experience Memory carries forward success, failure, and repair. New tasks are acquired without task-specific policy retraining. Our LIBERO evaluation attains state-of-the-art performance; controlled studies expose skill induction, persistent reuse, and agent-directed adaptation. Finally, we propose the Teach-and-Grow scaling-law hypothesis: if X denotes effective reusable experience, future-task error and teaching demand should approach irreducible floors as power laws in X. The architecture therefore treats deployment as a period of continued learning, in which one task can make the next easier.

84. 【2608.17205】Which Source Wins? Task-Dependent Reliance in Vision-Language Models

链接https://arxiv.org/abs/2608.17205

作者:Rodela Ghosh,Aviral Gupta,Guangjing Wang

类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)

关键词:Vision-language models, harder to read, Vision-language, combine images, text

备注: 20 pages. Under review

点击查看摘要

Abstract:Vision-language models (VLMs) combine images and text, but when the two conflict and one becomes harder to read, it is unclear how a model shifts its reliance between them. We study this modality reallocation with a controlled setup: we degrade either the image or the text across four levels of legibility while keeping the other clean, and track how the model's preference changes. We build conflicts from GSM8K and SVAMP by pairing the rendered image of one arithmetic problem with the text of another, so the two sources support different answers. We also introduce ChartQA-Conflict, a manually reviewed benchmark of 229 chart-report conflicts with matched chart and table-image representations. We evaluate six open-weight VLMs using both generated answers and a length-normalized conditional log-likelihood margin. On GSM8K and SVAMP, five of six models shift more strongly away from degraded text than from degraded images. On ChartQA-Conflict, all six likelihood-scored models exhibit the opposite pattern, shifting more strongly away from the degraded visual source. This reversal persists after calibrating for unimodal accuracy loss and after replacing charts with plain table images. Two frontier API models, GPT-5.6-Luna and Gemini-3.5-Flash, behaviorally replicate the ChartQA-Conflict reversal, with GPT-5.6-Luna also matching the arithmetic direction. These results show that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings. The source code is available at this https URL.

85. 【2608.17190】How smoothing the affinity matrix affects neighborhood preservation in t-SNE

链接https://arxiv.org/abs/2608.17190

作者:Shirin Mohebi,Guillaume Bied,Jefrey Lijffijt

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:visualize high-dimensional data, Dimensionality reduction methods, Dimensionality reduction, high-dimensional data, instrumental to visualize

备注: Accepted at the 29th International Conference on Discovery Science (DS 2026). 15 pages, 7 figures

点击查看摘要

Abstract:Dimensionality reduction methods are instrumental to visualize high-dimensional data, and t-SNE stands as one of the most widely used methods due to its emphasis on local neighborhood preservation. A central component of t-SNE is the affinity matrix, which expresses pairwise similarities in the form of symmetrized probabilities, over which the optimization problem of t-SNE is defined. We study how the sharpness of this probability distribution affects neighborhood preservation at different scales. We introduce a row-wise power transform controlled by a parameter gamma that can smooth or sharpen each row of the affinity matrix while preserving sparsity and rank order. We show that this transform is equivalent to rescaling the Gaussian bandwidth and thus to changing the perplexity. However, as the sharpness of the probability distribution varies per point, a fixed gamma leads to point-dependent effective perplexities, making it distinct from changing the global perplexity. Empirically, we find that sharpening improves preservation of the very nearest neighbors, while smoothing improves preservation of broader local neighborhoods, outperforming alternative affinity constructions including multiscale methods in the mid-local range.

86. 【2608.17182】RADmesh: Remesh-Aware Mesh Deformation

链接https://arxiv.org/abs/2608.17182

作者:Nam Anh Dinh,Itai Lang,Oded Stein,Rana Hanocka

类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)

关键词:generatively deforming shapes, generatively deforming, visual losses, propose a remeshing-enhanced, deformation

备注: ECCV 2026 (Oral). Our project page is at [this https URL](https://threedle.github.io/radmesh)

点击查看摘要

Abstract:We propose a remeshing-enhanced method for generatively deforming shapes with visual losses. It is intuitive that sufficiently drastic deformations of a mesh without changing its triangulation can easily compromise element quality, even if such large geometry changes may be semantically desired. Shape deformation methods could thus benefit from changing the triangulation; however, this is not done by most generative, text-based, visually-supervised mesh deformation methods. Remeshing is a discrete operation, proven to be especially challenging to couple with the notoriously noisy supervision signal provided by visual losses. We propose a vertex-based deformation optimization quantity capable of large deformations and robustness to such noise; we periodically remesh using an isotropic remesher that interpolates and carries forward the deformation optimization state. This enables continuous, geometry-informed progress in coarse-to-fine addition of resolution. The resulting shapes' triangulations fit their optimized geometry and have neat isotropic elements. Further, our method is localizable, able to grow new features on a base shape with expressive detail, leaving the rest unchanged. We showcase the effectiveness of our method on a variety of shapes and prompts, both local and global deformations, and demonstrate its superior visual quality and triangle efficiency. Our project page is at this https URL.

87. 【2608.17178】Mask What Matters: Saliency-Guided Video Self-Supervised Learning for Autonomous Driving

链接https://arxiv.org/abs/2608.17178

作者:Christopher Lang,Alexander Braun,Abhinav Valada

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:unlabeled data, promising paradigm, Video self-supervised learning, Video self-supervised, masked spatiotemporal prediction

备注: Accepted at GCPR 2026. The final publication will be available through Springer

点击查看摘要

Abstract:Video self-supervised learning through masked spatiotemporal prediction has emerged as a promising paradigm for learning feature representations from unlabeled data. However, existing methods typically rely on random masking, which indiscriminately removes regions irrespective of their semantic or temporal relevance. In ego-centric driving videos, this can weaken the pretext signal since safety-critical cues such as pedestrians, vehicles, lane boundaries, and dynamic interactions often occupy only a small portion of the frame, yet are central to downstream perception. We introduce V-JEPA4A, a domain-specialized variant of V-JEPA for autonomous driving that is pre-trained on publicly available driving videos with a novel saliency-driven masking policy. It accounts for semantically and temporally relevant context. The proposed policy preserves and predicts context according to semantic importance and temporal relevance, yielding more informative representation learning while retaining the efficiency of masked prediction. We evaluate the resulting encoders on four driving benchmarks spanning tracking, semantic segmentation, and depth estimation. The results demonstrate that V-JEPA4A reduces identity switches on BDD100k MOT by 25% over V-JEPA with random masking, achieves 73.2 mIoU on Cityscapes, and 3.75 RMSE on KITTI-2015 depth, while incurring only ~14% additional pre-training iteration overhead.

88. 【2608.17165】Rapid Debris-Volume Estimation from Post-Hurricane Aerial Imagery

链接https://arxiv.org/abs/2608.17165

作者:Kooshan Amini,Jamie Ellen Padgett,Guha Balakrishnan

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:Hurricane debris removal, removal is planned, documented over-estimation, hauling begins, federally reimbursed

备注

点击查看摘要

Abstract:Hurricane debris removal is planned, contracted, and federally reimbursed on the basis of volume estimates, yet operational practice still relies on parametric forecasts with 41-90% documented over-estimation or on truck-load tallies that arrive only after hauling begins. We present DebrisHeightNet, a segmentation-conditioned monocular debris-height network that estimates spatially explicit debris volume from a single pass of post-event aerial RGB imagery, the kind of survey routinely flown within days of a hurricane landfall. We train only a lightweight 1.08 M-parameter head on top of two frozen vision foundation models. This head regresses height from a Depth Anything V2 backbone, conditioned on the debris segmentation of CLIPSeg-debris from our prior work. Because no post-hurricane debris-height ground truth exists, we synthesize the training target by confidence-weighted LiDAR-monocular fusion (CW-LMF), designed to suppress non-debris LiDAR returns. This fused target is a constructed supervision signal rather than ground truth, so we corroborate it against external references rather than claiming it as truth. A region-level power-law calibration, driven by each region's low-density debris fraction, converts model volume into an estimate of the reported hauled debris with quantified uncertainty. Across ten regions spanning five hurricanes and three states, the uncalibrated model agrees with an independent uncrewed-aerial-vehicle (UAV) survey of the training region at Spearman $\rho = 0.87$ and lands within 30% of the reported record where the Hazus and FEMA-hybrid parametric forecasts over-predict it by 2.7-4.8$\times$. Deployment requires no LiDAR, no ground access, and no second flight, so the method can produce spatially explicit volume estimates wherever single-pass post-event imagery is flown.

89. 【2608.17151】Lymphocyte Mimicry Correction via Region-Level Tissue Reasoning and Unbalanced Optimal Transport

链接https://arxiv.org/abs/2608.17151

作者:Xiang Li,Yuqi Wang,Casey C. Heirman,Jihye Heo,Kyle J. Lafata

类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:morphologically similar, types appear morphologically, Unbalanced Optimal Transport, cell foundation model, Cell

备注: 13 pages, 3 figures. Accepted to the MICCAI 2026 COMPAYL Workshop

点击查看摘要

Abstract:Cell mimicry arises when different cell types appear morphologically similar. Human pathologists resolve this ambiguity using surrounding tissue context, whereas current vision models either lack contextual reasoning (cell foundation models) or cannot operate at the cell level (pathology MLLMs). We present Loki-OT, which propagates region-level tissue reasoning to individual cell predictions via Unbalanced Optimal Transport, using MLLM-derived density priors as soft guidance for ambiguous cell reassignment. Loki-OT is motivated by the observation that pretrained cell foundation model features already encode discriminative information, including tissue context, but standard cell-level supervision fails to use tissue context effectively. The resulting transport plan is distilled into a lightweight student MLP classifier that learns context-aware decision boundaries within the pretrained feature space. On the independent TCGA-BRCA cohort, Loki-OT achieved lower patient-level MAE than the fully supervised in-domain PanopTILs classifier and improved F1 in epithelium-rich mimicry tissues, using 278 weak region-level MLLM estimates built on a general-domain cell foundation model. Code: this https URL

90. 【2608.17129】PROBE: Manipulation-Grounded Visual Question Answering with VLM Agents

链接https://arxiv.org/abs/2608.17129

作者:Vineet Bhat,Siyi Chen,Alex Zook,Xuning Yang,Stan Birchfield,Valts Blukis,Jonathan Tremblay

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:Vision-language Models, planning in static, Vision-language, agentic tool-based planning, spatial reasoning

备注

点击查看摘要

Abstract:Vision-language Models (VLMs) excel at 2D grounding, spatial reasoning and agentic tool-based planning in static scenes. However, consider asking a home robot "Is my medication still in the cabinet?" The answer may be physically hidden behind a row of containers that must first be moved aside. Answering such questions in real-world cluttered environments requires reasoning in dynamic scenes: distractors must be manipulated to reveal occluded objects, and each action changes the scene the model must reason over. We formalize this setting as Manipulation-Grounded Visual Question Answering (MG-VQA) and introduce PROBE, a framework for benchmarking and finetuning VLM agents on such tasks. We first develop PROBE-Sim, a high-fidelity tabletop simulator with everyday objects and a robot manipulator equipped with grasping and pushing tools. PROBE-Sim is used to create PROBE-Bench: an evaluation suite of 150 tasks across 6 question types on cluttered tabletop scenes, where a VLM perceives, picks up or pushes objects before answering. We observe consistent trend across all frontier VLMs: agentic tool-based methods outperform their perception-only baselines (8.0% on average) across all task types. We further design PROBE-Agent, a finetuning recipe to distill successful trajectories from a powerful teacher foundation model to a smaller open-weight model using a mixed data recipe that encourages manipulation-efficient question answering. PROBE Agent finetuned models outperform their off-the-shelf agent baseline (11.5% on average) and demonstrate positive transfer to unseen objects and a held-out task. We validate sim-to-real transfer by deploying PROBE-Agent finetuned policies in real-world tabletop environments.

91. 【2608.17110】OV3D-Bench: A Diagnostic Benchmark for Open-Vocabulary Monocular 3D Detection

链接https://arxiv.org/abs/2608.17110

作者:Mariia Gladkova,Neehar Peri,Ishan Khatri,Deva Ramanan,Daniel Cremers

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:report strong in-domain, detectors report strong, strong in-domain performance, unavailable at deployment, single AP metric

备注: Accepted to OpenSUN3D workshop at ECCV'26; benchmark is released on [this https URL](https://github.com/mgladkova/ov3d-bench)

点击查看摘要

Abstract:Open-vocabulary monocular 3D detectors report strong in-domain performance, but each evaluates under a different protocol, several rely on per-image category oracles unavailable at deployment, and all collapse geometry and semantics into a single AP metric. To address this, we introduce OV3D-Bench, a diagnostic benchmark that compares open-vocabulary monocular 3D detectors under deployment-realistic conditions across seven indoor and outdoor datasets. Our benchmark replaces the per-image class name oracle with test-time dataset-level class name prompts, and decouples detection accuracy along three axes: localization, semantic robustness, and cross-domain transfer. We evaluate seven representative detectors and find that (i) they localize objects well yet often mislabel a correctly localized box as a semantically adjacent category; (ii) accuracy is highly sensitive to prompt phrasing (e.g. WildDet3D's performance collapses from 18.6 to 5.4 AP when prompted with "a detailed high-resolution photo of a car" rather than "car"); and (iii) the widely adopted target-aware protocol hides these errors (e.g. inflating DetAny3D's AP by 1.9 $\times$ on ScanNet). Lastly, we demonstrate that simply remapping a frozen closed-vocabulary detector's predictions using a contrastive vision-language encoder such as SigLIPv2 performs competitively against recent purpose-built open-vocabulary methods. This indicates that geometric localization is more mature, while open-vocabulary semantics remains the primary bottleneck.

92. 【2608.17095】Inference-Time Attention Steering for Vision-Language-Action Driving Models

链接https://arxiv.org/abs/2608.17095

作者:Darshan Nagendra Prasad,Lars Ullrich,Knut Graichen

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:driving models couple, diffusion-based trajectory decoder, time without retraining, give a direct, inference time

备注: Attention Steering, Vision-Language-Action, AutonomousDriving, Inference-Time Intervention

点击查看摘要

Abstract:Vision-language-action (VLA) driving models couple a reasoning stage with a diffusion-based trajectory decoder, but do not give a direct way to redirect attention toward safety-critical actors at inference time without retraining. We studied a bounded additive pre-softmax attention bias on the visual tokens of detector localized traffic actors on Alpamayo-R1's Qwen3-VL backbone. It is applied as a fail open forward pre-hook with no weight changes. On 50 lane-change scenarios from the Physical AI World Model Synthetic dataset. The trajectory decoder shows a monotonic dose response in the bias magnitude, separate from a paired zero bias control at every tested magnitude. It reaches $\approx 17$\,cm mean displacement with lateral shifts up to $\sim 140$\ cm at the clamp. A layer ablation places the action-relevant signal in late layers, where the effect increases with the number of hooked layers (2.0cm for the first 8 layers; 67.6cm for all 36). A per call injection audit explains why the Chain-of-Causation text never changes. The mask based bias never reaches the reasoning pathway in this serving stack, so the invariance is verified exposure, not robustness. Steered trajectories tend to shift toward the attended actor, suggesting the bias governs where the model looks rather than encoding a target behavior.

93. 【2608.17060】CAS-FD: Contact-Aware Temporal Sampling for Single-View Foul vs Dive Recognition

链接https://arxiv.org/abs/2608.17060

作者:Md. Jahidul Islam,Mahfujul Alam,Md. Nazmul Islam Seyam,Md. Tamim Hossain

类目:Computer Vision and Pattern Recognition (cs.CV)

关键词:multi-view camera angle, Distinguishing a genuine, single broadcast view, camera angle, contested fine-grained recognition

备注

点击查看摘要

Abstract:Distinguishing a genuine foul from a simulated dive in football remains one of the sport's most contested fine-grained recognition problems, especially when such decisions have to be from a single broadcast view without multi-view camera angle. We introduce a balanced 600-clip single-view Foul/Dive dataset and show that contact-aware sampling concentrating the model's attention around the moment of physical contact rather than treating all frames equally yields substantially improved recognition of this contact- specific problem. The proposed approach achieves 86.0% accuracy and macro-F1 0.860 on the held-out test split, a 12 percentage- point gain over contact-unaware alternatives that grows further on unseen data. We also evaluate each pipeline component against human annotations, establishing where and why the system suc- ceeds and fails. The result is a documented dataset, a reproducible single-view pipeline, and a grounded evaluation framework for fine-grained contact-event recognition in broadcast football footage. The dataset and code are available at this https URL tamim/contact-aware-dive.

94. 【2608.17044】he 10th AI City Challenge

链接https://arxiv.org/abs/2608.17044

作者:Zheng Tang,Shuo Wang,David C. Anastasiu,Ming-Ching Chang,Anuj Sharma,Quan Kong,Munkhjargal Gochoo,Jun-Wei Hsieh,Tomasz Kornuta,Zhedong Zheng,Renran Tian,Judah Goldfeder,Fulgencio Navarro,Yuxing Wang,Yizhou Wang,Sameer Satish Pusegaonkar,Anqi Li,Nalin Dadhich,Ridham Kachhadiya,Dhanishtha Patil,Haoquan Liang,Jiajun Li,Han Zhang,Yilin Zhao,Zaid Pervaiz Bhat,Shuyu Yang,Ashutosh Kumar,Rong Wang,Rafael Martin Nieto,Peter Christiansen,Ahmed Abduljawad,Mohanrasu Shanmugam,Nadeem Shaik,Sujit Biswas,Xunlei Wu,Vidya Murali,Rama Chellappa

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:held with ECCV, City Challenge, smart cities, marks a decade, decade of community

备注: Summary of the 10th AI City Challenge Workshop in conjunction with ECCV 2026

点击查看摘要

Abstract:The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference.

95. 【2608.17033】YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition

链接https://arxiv.org/abs/2608.17033

作者:Serdar Yildiz,Abbas Memiş,Songül Varli

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)

关键词:Yildiz Technical University, aims to recognize, query image, geo-referenced images, Visual Place Recognition

备注

点击查看摘要

Abstract:Visual Place Recognition (VPR) aims to recognize the location of a query image by comparing it with a set of geo-referenced images. Although many datasets have been proposed for VPR, collecting dense and diverse visual data from pedestrian-level viewpoints is still an important need. In this paper, we introduce YILDIZ-VPR, a visual geo-localization dataset collected through repeated walking traversals on the Davutpasa campus of Yildiz Technical University. The dataset includes outdoor scenes captured at different times of day, seasons, and weather conditions. It contains a wide range of visual content, including historical buildings, modern structures, roads, green areas, and wooded regions. Each video was recorded with a GoPro 9 camera and synchronized with GPS sensor data to provide location labels for the extracted frames. In addition to GPS coordinates, the dataset also includes auxiliary sensor information such as gyroscope, speed, and temperature data. With its dense coverage and long-term visual variability, YILDIZ-VPR provides a useful resource for studying image-based and temporal visual place recognition under realistic outdoor conditions.

96. 【2608.16984】PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

链接https://arxiv.org/abs/2608.16984

作者:Zhiyuan Yuan,Guanying Chen,Lingteng Qiu,Ruimao Zhang,Shuguang Cui,Xiaochun Cao

类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Graphics (cs.GR)

关键词:estimators achieve strong, Recent monocular depth, strong zero-shot generalization, Recent monocular, depth estimators achieve

备注: Project Page: [this https URL](https://yuanzhy29.github.io/PXDepth-Page/)

点击查看摘要

Abstract:Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel-level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-level depth prediction. Specifically, a large-patch ViT captures global scene context, while a pixel-space predictor composed of Context-Modulated Pixel Transformer blocks maintains high-resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at this https URL.

97. 【2608.16973】AerialYield-B2D: A Greenhouse Blueberry Dataset with Five-Stage Ripeness Masks and Fruit Counts

链接https://arxiv.org/abs/2608.16973

作者:Iyyakutti Iyappan Ganapathi,Afeefa Azam,Muhammad Owais,Irfan Hussain,Yusra Abdulrahman

类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)

关键词:public green house, masks remain limited, dense ripeness-stage masks, ripeness-stage masks remain, cluster composition

备注

点击查看摘要

Abstract:Blueberry ripeness is judged by berry colour, cluster composition, and the distribution of maturity stages within a plant, however, public green house image resources with dense ripeness-stage masks remain limited. We present AerialYield-B2D, where B2D denotes BlueBerry Dataset, acurated real-image resource containing 514 RGB images and 30,195 annotated blueberry instances across five ripeness stages: green immature, pale pink, pink-turns-purple, fully ripe and over-ripe. The release provides class-specific binary masks, overall berry masks, semantic label maps, image-level count tables, SHA-256 hashes, source metadata, recommended train/validation/test splits and technical validations. AerialYield is the broader project name; this release does not provide harvest weight, fruit mass or per-area yield measurements, and the count labels should therefore be interpreted as image-level berry counts rather than yield estimates. The images include 424 smartphone greenhouse images, 67 video-derived frames, and 23 DJI Fly video-frame samples, providing a reproducible dataset for ripeness segmentation, berry counting, and class-imbalance analysis in controlled-environment blueberry production.

98. 【2608.16966】Multi-Observer Vehicle Localization Case Study with Roadside Radar and Connected Vehicle Sensing

链接https://arxiv.org/abs/2608.16966

作者:Aleksi Pippuri,Nilusha Jayawickrama,Risto Ojala

类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)

关键词:intelligent transportation systems, modern intelligent transportation, estimate vehicle positions, conventional vehicles coexist, accurately estimate vehicle

备注: 12 pages, 7 figures and 8 tables

点击查看摘要

Abstract:In modern intelligent transportation systems, it is essential to accurately estimate vehicle positions, especially in mixed traffic conditions where both connected and conventional vehicles coexist. Roadside infrastructure and connected vehicles can provide complementary observations of the same traffic scene, but real-world evidence on decision-level fusion between these sources remains limited. This paper proposes a multi-observer vehicle localization framework that fuses compact object-level detections from a static roadside radar and a dynamic LiDAR-equipped connected vehicle. We evaluate the framework with real-world data collected at an urban intersection in Helsinki, Finland, with a separately instrumented target vehicle used as the reference trajectory. Two extended Kalman filter based strategies for the localization task were benchmarked. The performance of the radar and LiDAR sensors were evaluated separately, and the two fusion strategies were explored under nominal sensing conditions, reduced LiDAR update rates, simulated LiDAR occlusions, and different target-vehicle motion states. The results show that, under full LiDAR availability, fusion performance is dominated by the LiDAR observations, while the less accurate and less consistent radar observations provide only limited additional improvement. Nevertheless, AEKF achieves small gains over the LiDAR-only baseline, and object-level connected vehicle observations remain useful when shared at reduced update rates. These findings indicate that decision-level fusion provides scenario-dependent benefits rather than automatic improvement over a strong single-sensor baseline. We release the dataset and implementation on Github to support further research: this https URL

99. 【2608.16927】Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

链接https://arxiv.org/abs/2608.16927

作者:Peng Sun,Yi Yang,Antong Zhang,Chunxiao Li,Yanbo Wang,Dianbo Liu,xin chen,Kai Yu,Lu Chen,Tianfan Fu

类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)

关键词:improving model performance, large candidate pools, selecting high-value subsets, supervised fine-tuning data, fine-tuning data continues

备注

点击查看摘要

Abstract:As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise. We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS. MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder. Experiments on Vision Flan and LLaVA-CoT show that MASS consistently outperforms strong data selection baselines across multiple budgets, and in several settings matches or surpasses full data training with only a small subset of data.

100. 【2608.18055】Primitive Representation Learning for Unsupervised Dynamic Contrast Enhanced MRI Reconstruction

链接https://arxiv.org/abs/2608.18055

作者:Veronika Spieker,Wenqi Huang,Cemre Ariyurek,Liam Timms,Daniel Rueckert,Onur Afacan,Julia A. Schnabel,Sila Kurugol

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Signal Processing (eess.SP); Medical Physics (physics.med-ph)

关键词:Reliable quantitative analysis, MRI requires high-quality, requires high-quality spatiotemporal, dynamic contrast-enhanced MRI, contrast-enhanced MRI requires

备注

点击查看摘要

Abstract:Reliable quantitative analysis of dynamic contrast-enhanced MRI requires high-quality spatiotemporal reconstructions at high undersampling rates. Scan-specific reconstructions using Gaussian and Gabor primitives have shown promising results without the need for large training datasets, but have not addressed the additional dimension of dynamic contrast. We propose a multi-dimensional, primitive based framework for dynamic contrast-enhanced MRI reconstruction that disentangles the underlying anatomy, the dynamic contrast enhancement, and residual motion into separate temporal basis functions, thereby enabling a geometrical interpretation of the representation. We show that this architecture achieves performance competitive with conventional reconstruction methods, both in reconstruction quality and in the accuracy of extracted aorta and kidney enhancement curves. The modular tier design extends naturally to additional dynamic factors and higher acceleration rates. Code available at this https URL.

101. 【2608.18036】Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors

链接https://arxiv.org/abs/2608.18036

作者:Mahdi Saberi,Yaşar Utku Alçalar,Merve Gülle,Chetan Shenoy,Mehmet Akçakaya

类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Medical Physics (physics.med-ph)

关键词:data naturally utilize, naturally utilize complex-valued, utilize complex-valued measurements, k-space data naturally, MRI reconstruction

备注

点击查看摘要

Abstract:MRI reconstruction methods for undersampled k-space data naturally utilize complex-valued measurements. Parallel developments in sparse phase retrieval have shown that magnitude-only measurements may provide complementary information for signal recovery. However, their use in MRI reconstruction remains largely unexplored, due to lack of practical settings where informative magnitude measurements can be obtained without additional scan time. In this work, we investigate the use of auxiliary k-space magnitude information for accelerated steady-state dynamic MRI reconstruction, and demonstrate strong consistency of k-space magnitudes across time-frames. Building on this observation, we propose $\mathbb{C}+\text{Mag}$, a magnitude-informed physics-driven deep learning reconstruction method. The proposed method employs an ADMM-based unrolling framework with a novel magnitude-aware data-fidelity formulation, where quadratically smoothed optimization and momentum-based updates are introduced to address the non-differentiability and non-convexity of the magnitude constraints. Experiments on retrospectively undersampled cine MRI and phase-contrast flow MRI datasets, as well as prospectively undersampled real-time cine MRI acquisitions, demonstrate improved artifact suppression, sharper anatomical recovery, and better preservation of phase information compared to conventional PD-DL methods, which is further supported through blinded expert reader evaluations.

102. 【2608.16959】MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology

链接https://arxiv.org/abs/2608.16959

作者:Nabil Ashab,Soumit Kumar Kundu,Saif Mahmud Parvez,Shahadat Hossain Sohag,Bidhan Biswas,Nazmus Subha

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)

关键词:accuracy, common types, image accuracy, Breast cancer, patient accuracy

备注: 13 pages, 5 figures, Accepted for publication in the International Conference on Electrical, Computer and Communication Technologies (ECCT 2026) proceedings by Taylor \ Francis Books. This is the author-produced version

点击查看摘要

Abstract:Breast cancer is one of the most common types of cancer among women around the world. Rapid detection and early treatment can hinder its progress to more complex stages and can impede its spread to other parts of the body. Histopathological image classification is the most common task in cancer detection due to its robustness in analyzing cellular data. Breast histopathology classification requires handling both multi-scale tissue morphology and clinically relevant generalization beyond the source domain. This paper presents MagViT, an interpretable multi-magnification transformer framework with scale-gated fusion and patient-level model selection. The model uses four BreakHis magnifications (40X, 100X, 200X, 400X) and extracts per-scale representations with a ViT backbone, and combines them via a learnable gate that masks missing scales. Patient-level five-fold cross-validation with a fixed seed has been run and compared with three architectural branches. The most accurate branch is then selected as the final model due to the strongest patient-level accuracy while retaining the simplest fusion pathway. On BreakHis, our architecture achieves a mean image accuracy of 0.9191, a mean patient accuracy of 0.9643, and a mean macro-F1 of 0.9042. External transfer experiments provide preliminary evidence of cross-dataset generalization under controlled adaptation settings on BUSI (image accuracy 0.8306, macro-F1 0.7480, patient accuracy 0.8291) and IDC (image accuracy 0.8577, macro-F1 0.8191, patient accuracy 0.8372). Grad-CAM visualization indicates that the model focuses on diagnostically significant and meaningful regions across magnifications. Relative to prior ViT-centered BreakHis work, this study emphasizes patient-level selection and cross-dataset robustness under a reproducible protocol.

103. 【2608.16958】ORViT-DR: Ordinally-Robust Hybrid ViT for Low-Resolution Diabetic Retinopathy Grading

链接https://arxiv.org/abs/2608.16958

作者:Soumit Kumar Kundu,Nabil Ashab,Bidhan Biswas,Shahadat Hossain Sohag,Saif Mahmud Parvez,Souvik Kumar Kundu,Zunayed Ahmed Rafi

类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)

关键词:Diabetic retinopathy, impaired vision, Diabetic, Vision Transformer architecture, reliable automated grading

备注: Accepted at ECCT 2026

点击查看摘要

Abstract:Diabetic retinopathy (DR) is one of the main causes of impaired vision. A good and reliable automated grading system can make the screening process safer and more accurate. Because DR stages progress gradually, the task of grading disease severity naturally follows an ordinal structure in which neighboring classes share similar visual characteristics. In this study, ORViT-DR, a hybrid deep learning framework, is designed to improve DR grading from low-resolution retinal images. The proposed approach combines convolutional feature extraction with transformer-based global context modeling through a pre-trained ViT-Hybrid backbone, which integrates BiT-ResNetv2 with a Vision Transformer architecture. The approach is tested on the RetinaMNIST subset of the MedMNISTv2 dataset, which contains 28x28 retinal fundus images annotated with five levels of disease severity. To promote stable training and better feature learning, the training strategy applies progressive layer unfreezing, layer-wise learning rate decay, exponential moving average (EMA) parameter updates, and ensemble-based prediction during inference. Experimental results on the official RetinaMNIST test set show that the proposed method achieves 57.00% classification accuracy, along with a quadratic weighted kappa score of 0.5963 and a macro-F1 score of 0.4293. These results suggest that hybrid CNN-Transformer architectures can provide effective representations for ordinal retinal image analysis.