本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。
统计
今日共更新1419篇论文,其中:
- 自然语言处理171篇
- 信息检索36篇
- 计算机视觉314篇
自然语言处理
1. 【2608.09930】Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
链接:https://arxiv.org/abs/2608.09930
作者:Oluwanifemi Bamgbose,Simon Rosen,Jash Shah,Lindsay Devon Brin,Hoang H Nguyen,Anke Koelzer,Rachel Hansen,Tara Bogavelli,Fanny Riols
类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large Language Models, Audio Large Language, Opinion Score, Language Models, Audio Large
备注: Work in progress
点击查看摘要
Abstract:Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
2. 【2608.09928】Multimodal Model Diffing for Feature Discovery and Control
链接:https://arxiv.org/abs/2608.09928
作者:Hunar Batra,Lachin Naghashyar,Ashkan Khakzar,Philip Torr,Christian Schroeder de Witt,Constantin Venhoff,Ronald Clark
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large Language Models, Multimodal Large Language, Language Models, Large Language, exhibit strong visual
备注: Preprint. Accepted at ICML 2026 Trustworthy AI for Good Workshop
点击查看摘要
Abstract:Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
3. 【2608.09925】From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
链接:https://arxiv.org/abs/2608.09925
作者:Laurens Samson,Iva Gornishka,Gossa Lô,Yuki M. Asano,Sennay Ghebreab
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large language models, frameworks jointly reflect, Large language, existing evaluation frameworks, evaluation frameworks jointly
备注: Accepted at AIES 2026
点击查看摘要
Abstract:Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the "Grip on LLMs" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.
4. 【2608.09900】Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
链接:https://arxiv.org/abs/2608.09900
作者:Tadanobu Chuyo Kamijo,Ori Rottenstreich,Javier Conde,Gonzalo Martínez,Pedro Reviriego
类目:Computation and Language (cs.CL)
关键词:optimized generation corridor, highly optimized generation, Large language model, evaluations typically focus, model evaluations typically
备注:
点击查看摘要
Abstract:Large language model evaluations typically focus on performance under nominal conditions, creating an illusion of capability where models comfortably walk a narrow, highly optimized generation corridor. In real-world deployments, however, complex system prompts, safety guardrails, and structural constraints continuously force models off this nominal path, driving a divergence between benchmark scores and deployment performance. To address this issue, we introduce Decoding-Level Taboo, a zero-prompt diagnostic stress test that intervenes directly in logit space at runtime, forcing models out of their nominal paths. By dynamically masking primary candidate tokens at word boundaries, Taboo forces machine circumlocution. Evaluating Taboo across several open-weight model families reveals that off-path robustness is heavily influenced by both parameter scale and post-training instruction alignment, with robustness generally improving with model size and alignment. Beyond the results presented in this paper, Taboo provides a novel primitive for generating diverse synthetic datasets, stress-testing runtime safety guardrails, and auditing model reliability prior to real-world deployment.
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2608.09900 [cs.CL]
(or
arXiv:2608.09900v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.09900
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
5. 【2608.09898】Consilience for Verifier-Free Test-Time Scaling
链接:https://arxiv.org/abs/2608.09898
作者:Lecheng Kong,Like Hui,Haitao Mao,Jun Huan
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Test-time scaling, Verifier-free test-time scaling, confidence-based VF-TTS methods, VF-TTS methods, obtain high-quality rollouts
备注:
点击查看摘要
Abstract:Test-time scaling often uses an external verifier, such as compilers and test cases in coding or trained value functions in robotics applications, to obtain high-quality rollouts. Verifier-free test-time scaling (or VF-TTS) is gaining extensive attention as a mechanism to enhance Large Language Model (LLM) reasoning, primarily because we do not have access to such high-quality verifiers in many real-world applications. Among existing VF-TTS methods, confidence-based VF-TTS methods, which compute and rank rollouts solely by confidence, are particularly promising. Such methods introduce near-zero overhead for sample evaluation and require minimal access to internal model states, making the methods highly flexible across models and tasks. In this paper, we demonstrate a critical limitation of existing confidence-based VF-TTS methods by showing that such methods catastrophically break down on complex tasks. We observe a very interesting phenomenon: uniformly high confidence frequently indicates a failure to explore, favoring confidently wrong answers. To address this, our core insight is that robust cognitive search requires a specific confidence trajectory pattern: such methods perform exploratory branching at the beginning, as manifested by low initial confidence, and converge to a high final confidence solution. To implement this insight, we introduce consilience, a novel selection framework that explicitly evaluates the temporal asymmetry of confidence in reasoning. We operationalize this via a combinatorial metric that actively penalizes high initial confidence while strictly demanding final certainty. Extensive experiments covering both graduate-level mathematics problems and free-form code generation demonstrate that consilience effectively outperforms existing baselines, validating our novel perspective on completion confidence.
Subjects:
Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:
arXiv:2608.09898 [cs.CL]
(or
arXiv:2608.09898v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.09898
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
6. 【2608.09893】Fusion Training for Mathematical Generalization in Large Language Models
链接:https://arxiv.org/abs/2608.09893
作者:Congfeng Cao,Pengyu Zhang,Jelke Bloem
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large language models, enables large language, Thinking Mode, language models, single model
备注: ACL SRW 2026
点击查看摘要
Abstract:Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and a thinking mode within a single model. However, its training dynamics, including the \emph{data ratio} and \emph{training schedule} between the two modes, remain underexplored. In this work, we present a systematic study of TMF by analyzing the effects of the training schedule and data ratio between thinking and non-thinking modes. Focusing on mathematical problem solving, we construct a benchmark with multiple thinking-to-non-thinking data ratios and three training schedules. Our results reveal an asymmetric interaction between the two modes: increasing the ratio of non-thinking supervision reduces the accuracy of the thinking mode. We further show that different training schedules modulate this trade-off and that the optimal schedule depends on the data ratio. Finally, we quantify a negative correlation between non-thinking and thinking mode supervision, highlighting an inherent tension between these two modes. These findings provide practical guidance for designing effective TMF training settings. All code and data are released to support further research at: \href{this https URL}{\textbf{Fusion Bench}}.
7. 【2608.09861】owards Expert-level Medical AI for Real-time Video Consultations
链接:https://arxiv.org/abs/2608.09861
作者:Mahvish Nagda,Jihyeon Lee,Matthew Thompson,Chunjong Park,Tim Strother,Valentin Liévin,Roma Ruparel,Akshay Goel,Teya Bergamaschi,Suhana Bedi,Meet Shah,Pavel Dubov,Liviu Panait,Toshiyuki Fukuzawa,Sam Schmidgall,Craig Schiff,Joseph Xu,Aliya Rysbek,Yana Lunts,Jan Freyberg,Rebecca Hemengway,Sunny Virmani,David Racz,Carey Radebaugh,Joëlle Barral,Kavi Goel,Dale R. Webster,Katherine Chou,Avinatan Hassidim,Yossi Matias,James Manyika,Gregory Wayne,Tao Tu,Yun Liu,Ethan Goh,Christina Chen,Ryutaro Tanno,Po-Hsuan Cameron Chen,Mike Schaekermann,Anil Palepu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:enabling natural communication, enabling natural, AMIE, standard for patient-physician, natural communication
备注:
点击查看摘要
Abstract:Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
8. 【2608.09855】Agentic Auto-Research is Fuzz Testing
链接:https://arxiv.org/abs/2608.09855
作者:Yifeng He,Jicheng Wang,Yinzhe Zhao,Jiachen Liu,Hao Chen
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Autonomous research agents, generate experiments faster, Autonomous research, Autonomous, generate experiments
备注:
点击查看摘要
Abstract:Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the proposer and ranking more samples with a learned judge or human reviewers. We argue that this *generate-and-rank* paradigm misses the problem of sparse feedback. Within a declared research problem, an agent follows the control loop of a greybox fuzzer: it proposes a candidate, executes it, observes feedback, and chooses what to try next. A fuzzer rarely finds a bug, but coverage makes partial progress observable on every execution. Fuzzers then use that signal to mutate inputs and allocate effort, rather than only to rank completed runs. Auto-research needs the same two capabilities. First, each experiment should expose a cheap, dense signal of epistemic progress before final scientific validation is available. Second, that signal should determine the next intervention so that the agent searches rather than repeatedly samples. Because the optimized progress signal is guidance rather than a verdict, final validation must still decide what counts as a discovery using evidence protected from adaptive reuse. We propose controlled tests of whether candidate signals predict validated progress, whether feedback-directed search yields more validated discoveries per unit cost than repeated sampling, and whether protected validation reduces false discoveries. Feedback architecture, not only generation, is a central bottleneck in auto-research.
9. 【2608.09836】Mismatch Matters: On-Policy Distillation Beyond Token Agreement
链接:https://arxiv.org/abs/2608.09836
作者:Zichao Yu,Chengzhi Yu,Shengze Xu,Yujin Han,Bingqing Jiang,Xu Wang,Difan Zou
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:LLM post-training pipelines, modern LLM post-training, exploit repetitive loops, On-policy distillation, modern LLM
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from agreement to teacher-student mismatch, and find that mismatch tokens can be mainly categorized into two types: student-excess tokens and student-deficit tokens. Student-excess tokens are generated by the student but assigned near-zero probability by the teacher; their log-ratio corrections grow unbounded and destabilize the update. Student-deficit tokens, in contrast, are preferred by the teacher but rarely sampled by the student; their absence blocks the transfer of the teacher's reasoning patterns. To tackle these mismatch directions, we propose TIDE (Token-level Independent Deficit-Excess correction), which applies bounded Hellinger shaping to suppress the most severe sampled excesses and an analytic teacher top-$K$ injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher-student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. Moreover, the gains of TIDE are more pronounced under strong teacher-student mismatch, where it improves Avg@8 from 6.9% to 20.3%, reduces average response length by a factor of 3.6, and substantially reduces formatting failures. Code is available at this https URL
10. 【2608.09834】RA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classification
链接:https://arxiv.org/abs/2608.09834
作者:Fan Zhang,Jiaming Li
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:analysis converts unstructured, converts unstructured financial, sentiment analysis converts, Financial sentiment analysis, support market analysis
备注: 12 pages, 6 figures, 2 tables. Fan Zhang and Jiaming Li are co-first authors. Corresponding author: Jiaming Li
点击查看摘要
Abstract:Financial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making. Existing work on resource-efficient financial NLP has largely focused on compressing or adapting pretrained language models, with less attention to combining contextual representations with lightweight rule-derived features. This study develops Rule-Aware FinBERT (RA-FinBERT), a parameter-efficient framework that integrates low-rank adaptation (LoRA) with three continuous VADER-derived sentiment proportions (positive, negative, and neutral) and a source-level metadata feature. The standardized four-dimensional feature vector is directly concatenated with the 768-dimensional final-layer FinBERT [CLS] representation and passed through a lightweight classification head. This design introduces only 1,024 additional trainable weights relative to a structurally matched text-only FinBERT model. RA-FinBERT was evaluated against text-only FinBERT and a lightweight DistilBERT baseline for three-class sentiment classification of financial-news titles and descriptions. On the held-out test set, RA-FinBERT achieved 69.89% accuracy and a macro F1 score of 0.634, compared with 63.44% and 0.526 for text-only FinBERT. Neutral-class recall increased from 18.18% to 45.45%. The framework supports both CPU and GPU execution, offering a lightweight and practical approach to financial sentiment classification under constrained computational resources. These findings indicate that rule-derived sentiment information and source metadata can provide complementary signals to contextual FinBERT representations and improve performance with minimal additional model complexity.
11. 【2608.09819】Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
链接:https://arxiv.org/abs/2608.09819
作者:Mind Lab:Vin Bo,Asher Cai,Jingwei Cao,Song Cao,Vic Cao,Amelia Chen,Andrew Chen,Kaijie Chen,Cleon Cheng,Steven Chiang,Kaixuan Fan,Hera Feng,Huan Feng,Arthur Fu,Jun Gao,Pyke Han,Nolan Ho,Ori Hong,Hailee Hou,Piers Hua,Charles Huang,Miles Jiang,Nora Jiang,Yuyi Jiang,Qiuyu Jin,Fancy Kong,Kuss Koo,Jaron Lee,Andrew Lei,Alexy Li,Dawn Li,Lucian Li,Ray Li,Ricardo Li,Smith Li,Theo Li,Allen Lin,Elliot Lin,Fan Lin,Chen Ling,Kairus Liu,Kieran Liu,Logan Liu,Neo Liu,Xiang Liu,Yuxin Lu,Maeve Luo,Pony Ma,Verity Niu,Cole Qiao,Guian Qiu,Vince Qu,Sentry,Niko Song,Vincent Wang,Bo Wu,Rio Yang,Evelyn Ye,Fiona Ye,Ina Ye,Regis Ye,Josh Ying,Atlas Zeng,Danney Zeng,Salmon Zhan,Anya Zhang,Di Zhang,Mia Zhang,Sueky Zhang,Wei Zhao,Ada Zhou,Adrian Zhou,Yuhua Zhou,Juno Zhu,Murphy Zhuang
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:open agent-model family, agent-model family, family for experiential, real environments, environments and continuing
备注: 49 pages, technical report
点击查看摘要
Abstract:Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its successor. Collaboration is pursued via the Mixture-of-LoRA (MoL) architecture that freezes a base model, composes specialist LoRA adapters, and selects one LoRA per user turn. The flagship Macaron-V1-Venti combines a 744B GLM-5.2 base with four LoRAs for chat, agent, coding, and GenUI; the Qwen3.6-based Macaron-V1-Tall (50B) uses the same design for local deployment. This report presents Macaron-V1 as a co-designed system spanning architecture, algorithms, and infrastructure. The MoL architecture supports continual learning through extensible LoRA specialists. The algorithm combines Model-Harness Co-design and recursive self-improvement loop, including the UI4A component-native GenUI harness, a stateful action substrate, versioned HCP contract, and the agentic RL framework MindForge. The supporting infrastructure includes the post-training platform MinT, the long-context RL method LongStraw, and stability techniques for sparse MoE and DSA base models. We evaluate Macaron-V1 on Personal Intelligence, GenUI, and general capability benchmarks against frontier baselines. Our results validate the current system, while compounding gains from continual learning and collective intelligence remain open questions.
12. 【2608.09805】Parameter Exploration for RLVR via Variational Learning
链接:https://arxiv.org/abs/2608.09805
作者:Vatsal Venkatkrishna,Nico Daheim,Iryna Gurevych
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:reinforcement learning research, long time, reinforcement learning, learning research, Exploration
备注:
点击查看摘要
Abstract:Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performance. Many existing methods control exploration in the action-space, for example, using temperature scaling. However, these methods cannot reorder tokens but only influence the variance in the output distribution. This limits exploration and can lead to divergence or stalled training. Here, we investigate parameter-space exploration, where rollouts are generated by sampling different policies from a posterior that may each explore different rollouts. Sampling less or more diverse policies is then a complementary control lever over exploration. We introduce a family of methods called Perturbed Parameter Policy Optimization (3PO) which use different sampling strategies and different rollout grouping for reward estimation. Experiments on OLMo-3-1025-7B and Qwen2.5-Math-7B across mathematical reasoning and code generation tasks show that these approaches consistently improve average downstream performance over standard GRPO at a near-identical FLOPs cost. Moreover, using multiple parameter samples consistently produces fewer zero-advantage groups and malformed or incorrect rollouts during training than GRPO and action-space baselines. Overall, our work presents evidence that parameter-space exploration can improve reinforcement learning for LLMs.
13. 【2608.09802】SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
链接:https://arxiv.org/abs/2608.09802
作者:Yuling Shi,Jinghan Xu,Kelin Fu,Wenhao Zeng,Shilin He,Lei Zhang,Yue Liu,Zelin Zhao,Terry Yue Zhuo,Jialun Cao,Siyu Ye,Tianyu Liu,Kai Cai,Shing-Chi Cheung,Xiaodong Gu
类目:Computation and Language (cs.CL); Software Engineering (cs.SE)
关键词:unsolved SWE-bench Verified, long-horizon software engineering, check unstated requirements, recent audit found, reject correct solutions
备注: Published as a conference paper at COLM 2026
点击查看摘要
Abstract:As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and that frontier models can verbatim reproduce gold patches from training data. Code refactoring, which requires coordinated, behavior-preserving changes across many files, offers a substantially harder and more realistic test of agent capability, yet remains underserved by current benchmarks. We introduce SWE-Bench ProMax, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust). Every instance undergoes rigorous, multi-stage curation that directly addresses the quality problems identified in prior benchmarks: issue descriptions are rewritten from scratch to provide precise, unambiguous specifications, and test suites are manually reviewed to remove overly narrow and overly broad tests. Tasks with insufficient complexity or limited cross-file scope are filtered out, yielding a benchmark of challenging, large-scale refactoring tasks that average 11.4 modified files and 261.6 lines of code per instance, substantially exceeding the scale of existing benchmarks. Experiments with frontier models under two agent scaffolds show that the best model achieves only 41.2% resolve rate, confirming that SWE-Bench ProMax presents a meaningful and unsaturated challenge for current AI coding agents. Our benchmark is available at this https URL.
14. 【2608.09792】Comparing British and American Audio Description of Movies
链接:https://arxiv.org/abs/2608.09792
作者:Igor Sterner,Alex Lascarides,Frank Keller
类目:Computation and Language (cs.CL)
关键词:Narrating the visual, audio description, audio description created, visual component, British audio description
备注: CMN 2026 Workshop
点击查看摘要
Abstract:Narrating the visual component of movies is known as audio description. It is a narrative technique designed to enable blind and visually impaired individuals to follow the story. However, it is far more constrained than most narratives: the descriptions not only need to convey the story in the movie, but they must also fit into gaps between dialogue and they need to conform to guidelines that exist in each region. In this work, we compare audio description created in the United Kingdom against audio description created in the United States. We use guidelines written for these two regions, alongside the impressions from a practitioner in the field, to motivate specific hypotheses about the differences. We test these hypotheses against our pre-existing corpus, which provides both human-authored American and British audio description for each of 206 movies. Results provide quantitative evidence to uphold all tested hypotheses, including differences in lexicon, the use of the progressive aspect, the use of passive constructions, the use of subjective adjectives and modifiers, when characters are named, how scenes are cued, and degree of overlap with movie dialogue and music. Our work offers a quantitative lens into the narrative technique of audio description.
15. 【2608.09779】KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs
链接:https://arxiv.org/abs/2608.09779
作者:Ghanshyam Verma,Simanta Sarkar,Devishree Pillai,Hotaka Shiokawa,Yourong Xu,Fiona Veazey,Peter Hubbert,Hui Su,Paul Buitelaar
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Retrieval-Augmented Generation, Answering complex conditional, Language Models, Large Language
备注:
点击查看摘要
Abstract:Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly in domain-specific contexts where general-purpose LLMs and RAG tend to underperform. We hypothesize that augmenting RAG with unstructured and structured knowledge, extracted from both documents and knowledge graphs (KGs), can improve reasoning and answer accuracy for such tasks. To test this, we propose KGCaRe, a hybrid approach that combines neural retrieval with symbolic reasoning over LLM-generated KGs. KGCaRe constructs a KG from documents using a multi-prompt extraction strategy and stores it in a graph database. Simultaneously, the documents are embedded into a vector store to enable neural retrieval. KGCaRe performs innovative iterative graph traversal guided by the LLM to extract relevant triples, prune irrelevant information, and uses additional clue entities to traverse the graph again if the initial traversal does not provide satisfactory context to generate the answer. The relevant triples extracted from the KG in path form, along with semantically retrieved text passages, are then fed into custom KGCaRe prompts to generate answers to the complex conditional questions with explanations. We evaluate KGCaRe on two complex conditional QA datasets. Our results on these datasets show that KGCaRe consistently outperforms existing baselines, including Vanilla LLM, Code Prompt, Text Prompt, Think-on-Graph, Vanilla RAG, and HybridContextQA, across multiple LLMs such as Mistral, Mixtral, GPT-3.5, and GPT-4o. We publicly release the software pipeline that we developed to implement the proposed KGCaRe approach.
Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:
arXiv:2608.09779 [cs.CL]
(or
arXiv:2608.09779v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.09779
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
16. 【2608.09772】PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models
链接:https://arxiv.org/abs/2608.09772
作者:Zhanna Mukhametsharip(1),Vera Demberg(1 and 2),Varsha Suresh(2) ((1) Saarland University, Germany, (2) Max Planck Institute for Informatics, Germany)
类目:Computation and Language (cs.CL)
关键词:Large Vision-Language Models, demonstrated strong performance, Large Vision-Language, superficial correlations, demonstrated strong
备注: Under Review
点击查看摘要
Abstract:Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason about relationships between images and text or rely on superficial correlations, known as shortcut learning. This question is particularly important for multimodal sarcasm detection, where successful prediction depends on recognizing pragmatic incongruity rather than treating sarcasm as simple image-text mismatch. We introduce PragMatch, a controlled benchmark of 3,000 image-text pairs derived from MMSD2.0, including original sarcastic examples and constructed literal and hard-negative pairs. We identify influential shortcut cues through systematic masking and evaluate their impact through targeted injection experiments. Our results show that LVLM predictions are sensitive to lexical, OCR-derived and stylistic cues, with injected surface signals causing substantial changes in model predictions despite unchanged underlying image-text relationships. Our findings reveal limitations in current LVLMs while PragMatch provides a systematic testbed for evaluating multimodal pragmatic reasoning beyond surface-level image-text alignment.
17. 【2608.09767】Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
链接:https://arxiv.org/abs/2608.09767
作者:Abner Hernandez,Tomás Arias Vergara,Daiqi Liu,Andreas Maier,Paula Andrea Pérez-Toro
类目:Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
关键词:Real-time MRI makes, Real-time MRI, categories remains challenging, observe vocal-tract articulation, MRI makes
备注: Submitted for review at SLT 2026
点击查看摘要
Abstract:Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling. Specifically, we extract representations from PhonoQ's Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.
18. 【2608.09766】Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
链接:https://arxiv.org/abs/2608.09766
作者:Pinzhen Chen,Koel Dutta Chowdhury,Xiaoya Xu,David Tan,Doreen Osmelak,Ona de Gibert,Ariun-Erdene Tumurchuluun,Ashok Urlana,Fedor Sizov,Hale Sirin,Jesujoba Alabi,Karrar Talib Abed,Mateusz Klimaszewski,Nikolay Bogoychev,Niyati Bafna,Patricia Schmidtova,Preksha Manjunath Shanbhag,Sherrie Shen,Vilem Zouhar,Vivek Iyer,Yasser Hamidullah,Yusser Al Ghussin,Zheng Zhao
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Multilingual translation benchmarks, treating language pairs, sourced in English, English and translated, Multilingual translation
备注:
点击查看摘要
Abstract:Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.
19. 【2608.09765】REFRAMED: Towards Realistic Audio Description Generation for Movies
链接:https://arxiv.org/abs/2608.09765
作者:Igor Sterner,Mirella Lapata,Alex Lascarides,Frank Keller
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:visually impaired audiences, key visual content, Audio Description, enabling access, impaired audiences
备注: COLM 2026
点击查看摘要
Abstract:Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being told. However, existing approaches formulate AD generation in an artificial setting where both the content and timing of descriptions are pre-specified, reducing the task to clip-level captioning. They further rely on noisy transcription and alignment pipelines, and lack the rich parallel data required for modeling narrative context. We introduce a new formulation of AD generation in which models must jointly decide what to describe and when to do it. To support this, we present REFRAMED, a high-quality dataset of 2,023 videos that span 3,302 scenes from 206 movies, with professional AD transcripts (both American and British versions), professional subtitles and aligned screenplays. We also provide a manually curated challenge set that pairs full movies with multiple AD references, together with evaluation protocols that leverage dialogue gaps and multi-reference comparisons. Experiments with state-of-the-art AD systems and multimodal LLMs show that they outperform trivial baselines but fall far short of expert human performance. Our dataset and benchmark establish a new foundation for research on video understanding.
20. 【2608.09717】How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
链接:https://arxiv.org/abs/2608.09717
作者:Hasan Mahmud,Khawaja Abaid Ullah,Mohammad Javad Khojasteh,Jamison Heard,Prabu David
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
关键词:Large language models, judges remains unclear, perform subjective evaluations, subjective evaluations traditionally, evaluations traditionally made
备注: 9 pages, 2 figures, 2 tables. Includes technical supplement
点击查看摘要
Abstract:Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.
21. 【2608.09703】Matryoshka Language Model Suites
链接:https://arxiv.org/abs/2608.09703
作者:Nathan Godey,Yoav Artzi
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:classically requires training, suite classically requires, classically requires, separately and serving, Training
备注:
点击查看摘要
Abstract:Training a language model suite classically requires training each model separately and serving them independently. We improve both training and inference efficiency by stacking sub-models of increasing size into a single nested architecture trained end-to-end. This Matryoshka training framework reduces the total parameter count of the suite, enables low-cost distillation from the largest to all smaller sub-models at every training step, and is well-suited for speculative decoding as the draft model is contained within the verifier. We validate our approach by training a Matryoshka suite comprising 500M, 1.5B, and 3B sub-models. Our suite is on par with independently trained baselines on benchmark performance and validation and out-of-domain perplexities, while using 36% less training compute and improving the throughput of speculative decoding by 14-26%. We also ablate key architectural choices, offering guidance for building strong Matryoshka LM suites.
22. 【2608.09698】VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding
链接:https://arxiv.org/abs/2608.09698
作者:Ruqi Sun,Jiaping Li,Wenhui Tao,Ximing Zheng,Yuefeng Tan,Jiahao Wei,Yuxin Ma
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)
关键词:Great fiction earns, weaving domain expertise, pierce armor gaps, Great fiction, precise details
备注: 14 pages, 5 figures, Accepted by UIST'26
点击查看摘要
Abstract:Great fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot yet smell of decay, weaving domain expertise into the fabric of invented worlds. Current AI writing tools offer limited support for discovering and integrating unfamiliar domain knowledge into narrative. They require explicit queries that authors cannot formulate, generate finished prose that risks homogenizing voice, or assist only within the boundaries of what authors already know. We argue that AI should reveal latent knowledge gaps to writers while preserving their agency to transform discovered knowledge into authentic prose. Grounded in formative interviews with 9 fiction writers, we present VeriForge, a mixed-initiative writing system that divides cognitive labor so that the system assumes initiative over domain discovery while the author retains full initiative over narrative synthesis. VeriForge realizes this through three complementary mechanisms. Proactive inline highlighting flags potential knowledge gaps as authors draft. Dual-stream querying pairs conversational responses with source-anchored Knowledge Cards for direct fact extraction. A spatial Knowledge Canvas allows authors to organize and connect discovered knowledge across their writing. These mechanisms are powered by a graph-based retrieval-augmented generation pipeline grounded in domain-specific source materials. A within-subjects user study (N=12) provides preliminary evidence that this paradigm helps authors recognize previously overlooked knowledge gaps, supports creative exploration, and is perceived by expert raters to produce passages with stronger domain grounding in a controlled cold-start writing task.
23. 【2608.09650】Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking
链接:https://arxiv.org/abs/2608.09650
作者:Matan Fainzilber,Shlomit Plavner
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:substantial lexical gap, Reranking medical procedures, insurance information retrieval, health insurance information, information retrieval
备注: 10 pages, 6 figures, 4 tables. Code available at [this https URL](https://github.com/matanf-healthee/listwise-crossencoder-reranking)
点击查看摘要
Abstract:Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial lexical gap between patient language and clinical nomenclature. We present a systematic comparison of two reranking paradigms for this production task: (1) small cross-encoders (MedCPT, MiniLM-L12) fine-tuned with listwise learning-to-rank objectives across layer freezing configurations, and (2) Qwen3-Reranker-4B, a 4B-parameter instruction reranker whose prompt is iteratively refined via an agentic optimization loop driven by GPT-4.1. On a purpose-built dataset of 2,647 queries across 708 insurance services, we find that a 109M-parameter cross-encoder fine-tuned with ListNet outperforms the 4B-parameter model by 2.6 percentage points on NDCG@3 and 13.3 points on Spearman correlation - at 37x fewer parameters. We report practical findings, a scalable LLM based dataset construction pipeline, and deployment trade-offs relevant to production reranking systems. We release our code and a sample dataset to support reproducibility and adaptation to other domains.
24. 【2608.09638】Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
链接:https://arxiv.org/abs/2608.09638
作者:Yen-Shan Chen,Yu Chian Duan,Chih-En Kuo,Jian-Bin Wu,Yun-Nung Chen
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Computer Science and Game Theory (cs.GT)
关键词:Theory of Mind, limited diagnostic insight, provide limited diagnostic, agent interactions, essential for agent
备注:
点击查看摘要
Abstract:Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\times$2 taxonomy -- epistemic versus motivational reasoning crossed with inference versus action -- using human-crafted, perspective-constrained queries. Benchmarking 28 LLMs reveals three insights: 1) Reasoning, not knowledge. Models show strong game-rule comprehension but markedly weaker ToM abilities, isolating failures to social reasoning rather than missing domain knowledge. 2) Expression, not representation. Mechanistic analyses via linear probing and activation steering show that models frequently represent correct mental-state inferences in their hidden states but fail to express them during generation -- linear probes recover 77-82% accuracy versus 62-70% from the models' own chain-of-thought. 3) Policy, not deliberation. Dedicated reasoning training yields substantial improvements whereas test-time chain-of-thought provides only marginal gains (+11.0 versus +1.1 points on average), suggesting that robust ToM depends on a learned reasoning policy rather than increased inference-time deliberation.
25. 【2608.09624】Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
链接:https://arxiv.org/abs/2608.09624
作者:Mingyu Luo,Ming Deng,Zilang Qiu,Yiming Cheng,Ci Tao,Xue Tan,Sijin Sun,Yangfu Li,Ping Chen,Jun Dai,Xiaoyan Sun
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
关键词:Internal safety scores, Internal safety, text is generated, safety scores judge, separate harmful prompts
备注:
点击查看摘要
Abstract:Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks that would have failed anyway. In this paper we audit that inference. Attention based measurements are usually read from prompt dependent locations, so a wrapper changes both the content being judged and the place the signal is taken from. We therefore introduce Active Attention Probing, which supplies a fixed content independent measurement coordinate. We pair every base goal with a plain and a wrapped version and generate real completions from the target models. On Llama, wrapping raises harmful generation from 0.05 to 0.27 while harmful intent AUROC falls from 0.936 to 0.803, so the attacks grow more dangerous while the prompts look safer to the score. Among wrapped harmful prompts the outcome AUROC is 0.220, which places the attacks that succeeded below the attacks that failed. Rare token, passive, and detector derived channels reproduce the reversal on the same matched design, and the reversal itself persists across three target models, seven attack families, and two independent judges. Distribution shift then degrades calibration and threshold transfer before it degrades ranking.
26. 【2608.09588】MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL
链接:https://arxiv.org/abs/2608.09588
作者:Beiyu Xu,Zhenyu Wu,Jiaoyan Chen,Riza theresa Batista-navarro
类目:Computation and Language (cs.CL)
关键词:heterogeneous database collection, research and benchmarks, benchmarks assume, overlooking settings, Traditional
备注:
点击查看摘要
Abstract:Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collection. We therefore study schema linking in a multi-database setting, where the system must first locate the target database and then construct a compact, SQL-relevant schema for generation. We propose MDB-Link, a hierarchical schema-linking framework that retrieves question-relevant columns from a global index, aggregates retrieval evidence to shortlist databases, and uses a budget-aware large language model (LLM) for database reranking, table selection, and column grounding. With Qwen2.5-14B, MDB-Link outperforms LinkAlign on MMQA, Spider2-Snow, and BIRD-dev in database localization and column selection while producing schema subsets close in size to the gold schemas. Exact match improves from 16.88 to 51.41 on MMQA, 2.50 to 9.17 on Spider2-Snow, and 12.52 to 38.01 on BIRD-dev. MDB-Link also runs faster than LinkAlign and AutoLink, demonstrating the effectiveness of hierarchical schema reduction for downstream SQL generation.
27. 【2608.09568】Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization
链接:https://arxiv.org/abs/2608.09568
作者:Wenxiao Zhao,Shu Wang,Ying Nian Wu
类目:Computation and Language (cs.CL)
关键词:Direct Preference Optimization, aggregates token-level log-probability, token-level log-probability ratios, Direct Preference, Preference Optimization
备注: 16 pages, 2 figures, COLM2026
点击查看摘要
Abstract:Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing equally to the preference signal. However, the contribution of individual tokens to the preference signal varies. We introduce token credit, which modulates each token's KL regularization based on its contribution to the preference outcome. We derive that effective token credit is proportional to the magnitude of each token's implicit reward, and observe that this quantity evolves substantially during training. This implies that static token credit becomes increasingly misaligned as training progresses. In this work, we propose Se-DPO (Self-Evolving Token Credit for DPO), a live mechanism that derives token credit from the model's own evolving internal signals during DPO training. Since the reward signal varies in reliability across positions, Se-DPO calibrates token credit based on both the strength and the confidence of each token's contribution. Se-DPO requires no external models, adding only a lightweight calibration network with minimal computational overhead. Experiments show that Se-DPO improves over DPO by up to 9.8 points on AlpacaEval~2 and 12.2 points on Arena-Hard.
28. 【2608.09551】Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models
链接:https://arxiv.org/abs/2608.09551
作者:Bocheng Chen,Han Zi,Roucheng Ou,Yawei Liu,Minyue Chen,Zimo Qi,Rongrong Wang,Guangliang Liu
类目:Computation and Language (cs.CL)
关键词:exploit explicit linguistic, explicit linguistic cues, manipulate natural language, directly exploit explicit, manipulate natural
备注:
点击查看摘要
Abstract:In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit explicit linguistic cues in user prompts to bypass the safety mechanism of LLMs. However, such attacks can often be mitigated by existing safety alignment algorithms. On the other hand, human language is inherently grounded in pragmatics, necessitating typical context to interpret language, e.g., world knowledge, social norms. However, such contexts are often implicit because they are not directly expressed in human language and are not sufficiently leveraged in safety alignment, creating a fundamental mismatch between human language interpretation and safety alignment approaches. In this paper, we demonstrate that this mismatch exposes vulnerabilities in LLMs. We refer to this vulnerability as the pragmatic attack surface, which can be exploited to achieve high attack success rates. The experimental results demonstrate that our proposed approach outperforms baseline attack methods across various open-source and closed-source models by a substantial margin.
29. 【2608.09548】ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
链接:https://arxiv.org/abs/2608.09548
作者:Yilin Jiang,Xiaorong Zhu,Fei Tan,Zicheng Zhang,Kaiyi Huang,Yang Yu,Zexuan Fei,Yiming Luo,Keqian Li,Hao Hao,Aimin Zhou,Guangtao Zhai
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
关键词:Large language models, Large language, increasingly deployed, models, Large
备注: 13 pages, 6 figures, 8 tables. Benchmark data: [this https URL](https://huggingface.co/datasets/ZeroLoss-Lab/ELBench)
点击查看摘要
Abstract:Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.
30. 【2608.09539】Mawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance Detection
链接:https://arxiv.org/abs/2608.09539
作者:Rasha Albalawi,Nuha Albadi,Hamzah Luqman,Maram Kurdi,Saad Ezzini,Asma Yamani,Ahmed Ashraf
类目:Computation and Language (cs.CL)
关键词:detection remain limited, original Mawqif, original Mawqif dataset, remain limited, Women Driving
备注:
点击查看摘要
Abstract:Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This paper presents the Mawqif-v2 Extension, consisting of 996 manually annotated Arabic tweets collected from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is annotated with stance, sentiment, and sarcasm labels following the original Mawqif annotation scheme. The released extension is intended as a held-out evaluation set for assessing model generalization to both semantically related and previously unseen targets, while the original Mawqif dataset is used for training and development. In addition, we establish baseline results using several Arabic and multilingual transformer models, as well as zero-shot large language models (LLMs), to facilitate reproducible evaluation. Together with the original Mawqif dataset, the Mawqif-v2 Extension provides a benchmark for evaluating cross-target generalization in Arabic stance detection.
31. 【2608.09538】CS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
链接:https://arxiv.org/abs/2608.09538
作者:Vincent Cohen-Addad,Dimitris Paparas,Ernest van Wijland,Max Springer,Julien Canitrot-Paradis,Honghao Lin,David Woodruff,Adarsh Kumarappan,Rajesh Jayaram,Rudrajit Das,Lalit Jain,Ola Svensson,Silvio Lattanzi,Mislav Balunovic,Theophane Weber,Vahab Mirrokni
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:evaluating Large Language, Large Language Models, Theoretical Computer Science, Large Language, research-level Theoretical Computer
备注:
点击查看摘要
Abstract:We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA). Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark. We verify the correctness of generated proofs via a verification agent, and further benchmark the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs. Our reference verifier achieves over 90% accuracy on the expert labeled set.
32. 【2608.09510】Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts
链接:https://arxiv.org/abs/2608.09510
作者:Kevin Thomas,Milosz Kasprzyk,Reuel C Igbokwe Onuigbo,Elliott Pert,Cameron Tovey,João A. Leite,Olesya Razuvayevskaya,Carolina Scarton
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Social and Information Networks (cs.SI)
关键词:Detecting machine-generated disinformation, rewrite misleading content, Detecting machine-generated, large language models, make it easier
备注: Under review
点击查看摘要
Abstract:Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detector performance on fixed held-out datasets, do not capture how detectors behave when posts are deliberately transformed to evade classification. This paper adapts the Build it, Break it, Fix it framework into Build it, Break it, Repeat (BiBiR): iterative sessions designed to stress-test detectors' robustness under iterative adversarial conditions, evaluating whether models remain reliable when disinformation posts are systematically transformed to evade classification. Across five iterations, the findings show that the best adversarial breakers' transformations came from a combination of back-translation and LLM persona-based rewriting, with the best performing technique achieving a 95% label flip rate (LFR), whilst still preserving the meaning of the original posts. The best builders' model was a triplet contrastive model with a dynamic anchor switching (DASS) architecture, which achieved an average accuracy of 72.68%, outperforming the strong baseline (a fine-tuned e5-small-LoRA) by 15 percentage points on the most robust set of breakers' adversarial attacks. The results demonstrate that an iterative framework best exposes detector weaknesses and pushes robustness improvements; however, it may still require semantic preservation analysis to distinguish valid adversarial evasion from transformations that changed the original disinformation claims' meaning.
33. 【2608.09507】Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
链接:https://arxiv.org/abs/2608.09507
作者:Yuting Liu,Wei Wu,Jianzhe Zhao,Guibing Guo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Natural language user, Natural language, interface for LLM, provide an interpretable, interpretable interface
备注:
点击查看摘要
Abstract:Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), \textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8\% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.
34. 【2608.09444】Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
链接:https://arxiv.org/abs/2608.09444
作者:Kristian Schwethelm,Daniel Rueckert,Georgios Kaissis
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)
关键词:looped language models, main promise, language models, depth-adaptive inference, forward pass
备注:
点击查看摘要
Abstract:A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, this adaptivity breaks standard batching: tokens in the same batch now require a different number of loops, so there is no unified forward pass, making efficient inference difficult. Standard inference frameworks like vLLM schedule on the token level and cannot handle this because tokens need to be removed from the batch within the forward pass. Loop-level scheduling has been proposed as a solution, but never implemented end to end. The key challenge is that looped architectures also contain non-looped boundary stages (e.g., token embedding and LM head) that must be scheduled at different frequencies than the loop. We introduce continuous depth batching (CDB), which schedules at the granularity of individual loop iterations. CDB handles boundary stages and loop steps in separate priority queues, makes exit decisions one step ahead, and overlaps all scheduling work with GPU computation. On Ouro 1.4B and Huginn 3.5B, CDB can realize up to $99\%$ of the theoretical maximum speed-up from adaptive-depth, translating to $1.5$-$1.9\times$ higher offline throughput and $45$-$90\%$ lower normalized latency under dynamic serving load.
35. 【2608.09432】ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models
链接:https://arxiv.org/abs/2608.09432
作者:Róisín Luo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Transformer-based language models, representing token order, Transformer-based language, lacks an intrinsic, intrinsic mechanism
备注:
点击查看摘要
Abstract:Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order. Existing architectures address this limitation by explicitly incorporating positional information through learned positional embeddings or hand-crafted positional encodings, such as rotary positional encoding (RoPE), treating positional information as an architecturally acquired capability rather than an inherent property of the model. Motivated by the pursuit of positional-encoding-free architectures, this work explores a language model architecture that integrates causal state-space equations to implicitly encode positional information before attention computation. Specifically, each model block applies a causal state-space equation before self-attention, allowing recurrent state dynamics to encode sequential information into token representations. Consequently, subsequent attention layers operate on position-aware representations without requiring explicit positional encodings while retaining the expressive modeling capacity of self-attention. We present \textsc{ZetaGPT}, a compact hybrid language model designed for research, rapid prototyping, algorithm verification, and educational applications. In addition to the proposed architecture, \textsc{ZetaGPT} provides a fully open-source, end-to-end training pipeline encompassing dataset construction, tokenizer training, pretraining, supervised fine-tuning, reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning via pure reinforcement learning. To the best of our knowledge, \textsc{ZetaGPT} is the first open-source small language model without explicit positional encoding and establishes a compact, reproducible reference implementation for the development and empirical study of positional-encoding-free language models.
36. 【2608.09424】Reducing Pretraining-Generation Mismatch in Diffusion Language Models
链接:https://arxiv.org/abs/2608.09424
作者:Xiaocheng Lu,Huabin Liu,Song Guo,Jianguo Li
类目:Computation and Language (cs.CL)
关键词:predicts future tokens, language models align, clean left context, training predicts future, Autoregressive language models
备注: 12 pages, 9 figures, 1 table
点击查看摘要
Abstract:Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language models offer parallel denoising, but native dLLM pretraining can randomly corrupt prompt and continuation tokens together, weakening the clean-prefix interface needed for prompt-conditioned generation. We identify this mismatch for prompt continuation and propose PCD (Prefix-Conditioned Diffusion), a pretraining objective that combines AR prefix supervision with no-shift suffix denoising. At the training-objective level, PCD changes the attention mask, corruption mask, and label construction in continued pretraining; it does not require an autoregressive decoder, verifier, or new inference mode. By supervising the clean-prefix side autoregressively and applying diffusion only to the unknown continuation, PCD makes the local training interface resemble how block-diffusion models are queried at evaluation time. We further separate intra-sample prefix conditioning from inter-sample objective mixing, allowing us to identify the local alignment signal separately from the optional batch-level mixing knob. Across LLaDA2-Mini and Qwen-1.7B backbones, PCD consistently improves over same-family native dLLM stable baselines, reaching a 4.2% relative gain on the main LLaDA2-Mini six-benchmark average (+2.56 points) and a 14.2% relative gain in the primary Qwen mechanism comparison (+4.86 points). These results suggest that aligning the pretraining context distribution with prompt-conditioned generation can recover a measurable part of the dLLM continuation gap without changing inference.
37. 【2608.09420】Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation
链接:https://arxiv.org/abs/2608.09420
作者:Bo Wang,Ruixing Zhang,Yunqi Liu,Yang Zhang,Liangzhe Han,Tongyu Zhu,Leilei Sun
类目:Computation and Language (cs.CL)
关键词:evaluating interactive assistants, interactive assistants, simulators are widely, scalable environments, environments for training
备注: 26 pages, 7 figures, 16 tables. Code: [this https URL](https://github.com/ptwang773/UserIDA)
点击查看摘要
Abstract:User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6\% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7\% of evaluated dialogue states, compared with 22.9\% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.
38. 【2608.09393】mporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
链接:https://arxiv.org/abs/2608.09393
作者:Rose Cymbler,Daniel Guez,Laurent Fabre
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:quantify temporal misgrounding, Standard legal RAG, identify and quantify, earlier or future, legal RAG treats
备注: 13 pages, 1 figure, 4 tables. Accepted at the ICML 2026 Workshop on AI for Law (AI4Law), Seoul. Code and data: [this https URL](https://github.com/rosecymbler/fiscal-fr-bench)
点击查看摘要
Abstract:We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem. We introduce FiscalQA Pro, pairing a versioned corpus of 32,436 article-versions of the French tax code (93 years, 1938-2031) with an all-model-hard temporal-reasoning track: 209 scored, expert-reviewed questions across 33 CGI articles (221 released; twelve flagged out of the answerable scope). At selection time, no evaluated model recovered its date-applicable answer closed-book in any of four sampling draws, and the currently in-force text lacks the gold value for all but one of the scored questions. Answers are scored deterministically via atomic ground-truth "nuggets" (regex and numeric-with-tolerance), never LLM-as-judge: an LLM judge would inherit the temporal bias it is meant to score. Across eleven models (five frontier closed-API systems plus Gemini 2.5 Pro as a substitute entry, and five open-weight), parametric knowledge yields 3.0% mean strict accuracy and RAG over a static current-version corpus 2.7%. Static RAG retrieves the date-applicable version 0% of the time, confidently citing a real but inapplicable version. Our end-to-end retriever over a multi-version index, with no oracle, reaches 98.3% mean strict; an oracle-article ablation reaches 99.1%, locating the residual gap in first-stage recall, not version selection. We additionally release a version-aware jurisprudence dataset of 69,208 citation links, together with the corpus, benchmark, model responses, and pipeline code.
39. 【2608.09356】Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages
链接:https://arxiv.org/abs/2608.09356
作者:Zijie Zhang
类目:Computation and Language (cs.CL)
关键词:Closely related languages, limiting cross-lingual transfer, Closely related, Common Turkic Script, related languages written
备注:
点击查看摘要
Abstract:Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare two approaches to script unification: the general-purpose uroman romanizer and the family-specific Common Turkic Script (CTS). We train matched fastText models on transliterated Wikipedia corpora from 11 Turkic languages and evaluate them on WikiANN named entity recognition and Universal Dependencies part-of-speech tagging. CTS and uroman show no significant difference on NER, while both substantially outperform the official monolingual fastText baselines. POS results reveal no universal winner: language-specific differences are associated with the cross-lingual character n-gram coverage induced by each representation, while within-language coverage becomes more important when target-language supervision is available. Although CANINE-c achieves higher overall POS averages, the substantially simpler fastText-based systems remain competitive on several treebanks. Overall, the effectiveness of script unification depends on the language, the induced subword overlap, and the available supervision.
40. 【2608.09292】Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents
链接:https://arxiv.org/abs/2608.09292
作者:Bingzhen Liu,Xiaomeng Fan,Yuwei Wu,Zhi Gao,Mingyang Gao,Chuanhao Li,Yunde Jia
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Self-evolving methods improve, Self-evolving methods, improve the capabilities, capability boundary, underlying LLMs
备注:
点击查看摘要
Abstract:Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample correct trajectories on difficult examples for further improvements. In this paper, we propose a zeroth-order self-evolution framework that enables agents to learn beyond their capability boundary by perturbing LLM parameters to adapt to difficult examples without any trajectory annotations. Specifically, we perturb LoRA parameters of LLMs, run the agent, compute the losses under the perturbed and original parameters, and use the loss difference to estimate gradients and further update the LoRA parameters. We sample trajectories using the updated LLMs for supervised fine-tuning to break through the capability boundary of the agents, forming a closed self-evolution loop. We introduce a parallel perturbation inference mechanism and an adaptive lookup mechanism to reduce time consumption in zeroth-order optimization, with an answer perplexity loss that provides smooth and stable zeroth-order loss values. Experiments on multiple deep research benchmarks show that our method obtains substantially more successful trajectories and consistently outperforms strong baselines, especially on difficult examples. The code and released artifacts are available at this https URL.
41. 【2608.09289】Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing
链接:https://arxiv.org/abs/2608.09289
作者:Steve Woollaston,Brendan Flanagan,Hiroaki Ogata
类目:Computation and Language (cs.CL)
关键词:research distinguishes grammatical, language writing research, writing research distinguishes, automated writing evaluation, distinguishes grammatical accuracy
备注: APCLC submission
点击查看摘要
Abstract:Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that isolates structural errors from unnaturalness by generating literal error corrections and idiomatic revisions for 3,830 English writing samples from 120 Japanese junior high school students. Applying the regex-based CEFR-J grammar extractor, we quantify two diagnostic measures: accuracy gaps (structures attempted but incorrectly produced) and idiomatic gaps (grammatically correct structures underused or overused relative to native norms). Results reveal distinct patterns: definite articles, third-person singular -s, and modals (would, could) exhibit significant accuracy difficulties, while -ing forms and hypothetical modals (would) show the largest idiomatic underuse, with simple present verbs, subject-verb-object patterns, and modal can conversely exhibiting the most pronounced overuse. A two-dimensional instructional typology maps error rates against idiomatic gaps, distinguishing accurate but overused grammar items from error-prone or avoided complex forms requiring targeted production practice. This framework advances pedagogical feedback by enabling teachers to diagnose whether learner difficulties arise from inaccurate execution, structural avoidance, or L1-mapped overreliance, supporting evidence-based interventions tailored to the specific needs of each learner.
42. 【2608.09282】ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
链接:https://arxiv.org/abs/2608.09282
作者:Adrian Li,Kelong Mao,Yudong Guo,Heming Xia,Xinwei Yang,Lirui Luo,Jace Wong,Pu Yao,Sulong Xu,Simiu Gu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Real-world shopping, requires constructing, retrieving a single, complementary items, Real-world
备注:
点击查看摘要
Abstract:Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.
43. 【2608.09280】Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025
链接:https://arxiv.org/abs/2608.09280
作者:Nusrath Jinnath,Wei Zhao
类目:Computation and Language (cs.CL)
关键词:Responsible NLP Checklist, Responsible NLP practice, NLP practice includes, Responsible NLP, NLP Checklist aims
备注:
点击查看摘要
Abstract:Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to aid transparency on the current research practice, which we focus on. We curate and release the first two datasets of: a) all the checklist responses and justifications from the EMNLP 2025 Main and Finding tracks; b) checklist reference linking to paper sections. We also provide the first analysis of recent EMNLP Checklists, by examining $73,922$ responses and justifications to them. For the Main track, we find that authors isolate ethics questions of the Checklist from the paper's bulk, mimicking the trend of ethics being an afterthought. We then examine \texttt{NO} responses. We find $44.9\%$ of justifications are poor or bad-faith, being brief or empty. Then, we find significant issues with the checklist design and effort of authors, namely that $6\%$ of all checklists contained logical contradictions between parent and child responses. We also find evidence of surface compliance for responsible ethics, with $53\%$ authors dismissing potential risks or social impacts of their work, for which there should be none. We compare this to the Findings track, noticing a similar trend in both tracks. Lastly, we discuss the implications of the checklist design and provide recommendations for future checklist iterations. Including: a) enforcing a minimum word count, b) enforcing more scrutiny on the risks of appliances.
44. 【2608.09276】Verifiably grounded machine interpretation of lunar geology
链接:https://arxiv.org/abs/2608.09276
作者:Tom Sander,Kay Wohlfarth,Christian Wöhler
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Planetary geology relies, reconstruct past events, Planetary geology, interpretive reasoning, diverse observations
备注:
点击查看摘要
Abstract:Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an automated "machine intelligence geologist" by embedding this distinct methodology of geologic knowledge discovery and inference into a multimodal vision-language architecture. Focusing on the stratigraphy of lunar basaltic mare volcanism, we train a model to generate verifiably grounded geologic interpretations directly from co-registered topographic, spectral, and geologic maps. We demonstrate that while the system successfully balances established geological priors with local visual evidence to accurately describe stratigraphy and terrain, numeric age dating derived solely from vision defaults to memorized priors. Integrating an open-book retrieval mechanism resolves this, enabling the model to faithfully cite published chronologies. Our findings delineate the necessary architecture for automated geologic inference: site evidence must be visually interpreted from local data, while quantitative historical context must be retrieved from the scientific record.
45. 【2608.09254】Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
链接:https://arxiv.org/abs/2608.09254
作者:Morris Lee
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:LLM analytics agents, SQL syntax accuracy, valid business definitions, LLM analytics, SQL syntax
备注: 20 pages, 7 figures, 11 tables. Benchmark, code, run records, pre-registration and one-command reproduction: [this https URL](https://github.com/k-w-lee/query_proof)
点击查看摘要
Abstract:LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number. No execution-match metric can score them. This paper introduces WarehouseReliabilityBench, 400 frozen tasks over two synthetic warehouses in which roughly half the correct responses are a clarification, an abstention or a refusal, with pinned denominators and a pre-registered paired bootstrap fixing each claim verb before the numbers existed. QueryProof, a 7B agent, uses rules derived from a semantic layer and physical catalog to determine its behaviour, and gates every answer on deterministic post-execution checks. On an 80-task synthetic test split evaluated once, QueryProof outperforms a direct-prompted 32B baseline by +0.237 [+0.112, +0.375] Business Truth Rate at 71.0% lower cost per correct answer; against a cost-matched few-shot baseline the accuracy gain holds but the cost difference does not resolve. This compares systems rather than model sizes: the 32B baseline receives none of the scaffolding. False success falls from 0.754 to 0.351 of returned answers, and no wrong number was returned on an answerable task (0 of 24), though 13 answers went to questions requiring clarification or abstention. Removing the routing layer changes little (0.562 against 0.537), so the result does not depend on escalation. Routing tuned on validation over-abstains on test, and the fitted confidence model loses to the heuristic it replaced. Resampling template families rather than tasks widens both accuracy intervals to include zero, so the effect's direction is better supported than its magnitude. The gain tracks the deterministic layer, though no component ablation was run.
46. 【2608.09251】MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
链接:https://arxiv.org/abs/2608.09251
作者:Peiwen Li,Shiyang Zhang,Yangtian Zhang,Sizhuang He,David van Dijk,Rex Ying
类目:Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large language model-based, recently shown strong, shown strong potential, Large language, language model-based multi-agent
备注: 25 pages, 8 figures, 9 tables
点击查看摘要
Abstract:Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing methods mainly rely on coarse prompt-level differentiation without parameter adaptation for diverse subtasks, resulting in insufficient inter-agent heterogeneity and limited specialized capability that bottleneck performance on tasks with complex requirements. To address this, we introduce a Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts (MoRSE) that distinguishes agents with (role, subtask)-conditional specialization at both the task structure and parameter levels. To make agents' responsibility explicit at the task structure level, we formulate a task-oriented multi-agent system that decomposes each task into a dependency-aware Directed Acyclic Graph of subtasks and assigns each agent a specific (role, subtask), introducing task-level specialization across collaborating agents. Additionally, to address the diverse role and subtask parameter adaptation demands, we propose a dynamic Mixture of (role, subtask) LoRA Experts module with a prototype-based semantic router for subtasks, augmenting agents with parameter-level specialization on a shared LLM substrate cost-effectively. Then, to co-optimize experts and router stably under sparse task rewards, we further propose a hierarchical group-relative policy optimization with two-layer credit assignment that isolates expert updates from the cross-route variance introduced by routing decisions, disentangling expert quality from routing quality. Experiments on code-generation benchmarks across three backbones demonstrate the effectiveness of our approach, with improvements in both whole-task and step-wise performance, and the gains from trained specialization generalize across held-out task categories and domains.
47. 【2608.09222】Reading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision Model
链接:https://arxiv.org/abs/2608.09222
作者:Jiawen Kang,Dongrui Han,Xixin Wu,Helen Meng
类目:Computation and Language (cs.CL); Neurons and Cognition (q-bio.NC)
关键词:modeling infers latent, infers latent properties, existing formulations rely, formulations rely primarily, observed behavior
备注:
点击查看摘要
Abstract:Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action trajectories. In verbalized cognitive tasks, task execution also produces response dynamics that action-only formulations leave unmodeled, such as verbal production, interaction, and hesitation. We propose a factorized inverse decision model (FIDM) that decomposes each individual's task-execution likelihood into an action factor and an effort factor, governed by separate individual-specific parameters. From raw verbal transcripts, a language model produces structured task-execution traces for factorized inference. On data from 400 older adults performing a grocery-shopping dialog task for cognitive screening, controlled recovery shows selective estimation of the intended factors, while matched semi-synthetic conditions show that FIDM preserves action-execution distinctions even when aggregate behavioral summaries are matched. Action evidence further localizes task-defined deviations across participants. In cognitive-status classification, FIDM provides information complementary to clinical scores, trajectory summaries, and frozen language representations, with consistent gains across all evaluated baselines in the binary setting.
48. 【2608.09209】UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
链接:https://arxiv.org/abs/2608.09209
作者:Chidaksh Ravuru,Shashank Srivastava
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:boosting benchmark performance, Neural language models, large crowdsourced corpora, crowdsourced corpora frequently, corpora frequently exploit
备注: Accepted at COLM 2026
点击查看摘要
Abstract:Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed. We present U N M ASK, a fully automated pipeline that discovers, causally verifies, and mitigates spurious correlations in text classifiers without additional human annotation. Given unlabeled training examples, U N M ASK generates candidate surface patterns as executable boolean expressions, filters them through a statistical validation protocol with independent replication, and establishes causal model dependence via verified counterfactual interventions. Causally confirmed features then serve as annotation-free group definitions for Deep Feature Reweighting, eliminating the group labels that standard DFR requires. Applied to BERT and RoBERTa trained on MNLI, our pipeline independently rediscovers established lexical-overlap and negation biases, verifying 9 of 10 features on BERT and 6 on RoBERTa, and improving HANS accuracy by up to 12.58 pp. On CivilComments-WILDS, programmatic groups match the 70.1% worst- group accuracy of hand-labeled DFR (Kirichenko et al., 2023) without demographic annotation. We further demonstrate that the discovery and validation stages generalize to reward model preference data, surfacing interpretable spurious correlations in RewardBench2.
49. 【2608.09189】EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
链接:https://arxiv.org/abs/2608.09189
作者:Junyu Wang,Siyuan Zhang,Peiyuan Jiang,Jian Zong,Jingyu Zhang,Tianrui Wang,Yuqin Lin,Zhenghui Chen,Shuqing Xie,Ziyang Ma,Meng Ge,Xiaobao Wang,Longbiao Wang,Jianwu Dang
类目:Computation and Language (cs.CL)
关键词:rudimentary paralinguistic perception, Emotional Intelligence, Spoken Language Models, theory-driven cognitive framework, lacking a systematic
备注: Accepted at ACM Multimedia 2026 (MM '26)
点击查看摘要
Abstract:Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
50. 【2608.09187】Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System
链接:https://arxiv.org/abs/2608.09187
作者:Yanlin Yu
类目:Computation and Language (cs.CL)
关键词:long-form translation request, request can succeed, produce an unusable, API layer, long-form translation
备注: 9 pages, 2 figures. A sanitized reference implementation is included as ancillary material
点击查看摘要
Abstract:A long-form translation request can succeed at the API layer and still produce an unusable result. The output may be empty, truncated, filtered, dominated by source or prompt material, or interrupted after producing text worth keeping. This report describes a recovery protocol developed for a deployed translation system with heterogeneous inputs and provider APIs. It delays the first visible release behind a 64-character window, validates the assembled output, and uses typed stream events to distinguish replacement from continuation. Interrupted work is retained only when a paragraph or sentence prefix can be re-derived from the source. Further attempts follow a stable model order and a shared deadline before entering a provenance-marked fallback path. A sanitized companion artifact implements the protocol and passes 38 public tests. Its fixed cases reproduce all 14 configured completion labels, contain four early-invalid prefixes before any of their 235 characters become visible, retain 31 boundary-safe characters across four interrupted streams, and satisfy the attempt, event, and provenance rules in two end-to-end scenarios. These results are executable checks of the published control flow. Translation quality and detector performance on naturally occurring outputs require a different evaluation.
51. 【2608.09154】UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following
链接:https://arxiv.org/abs/2608.09154
作者:Jeet Sharma,Balpreet Kaur,Jeremiah Hong,Hamed Zamani,Haw-Shiuan Chang
类目:Computation and Language (cs.CL)
关键词:follow long lists, follow complex instructions, enhance LLMs' ability, follow complex, Large language models
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a reference document (i.e., back-translation) is a widely used method to measure/enhance LLMs' ability to follow complex instructions. However, this method introduces a critical loophole: the constraint synthesis model copies text from the reference as a very specific constraint and the evaluated LLM trivially satisfies the constraint by copying its text in the response. To address these issues, we propose UNSPECIFIC, a novel framework that synthesizes constraints common to two similar reference articles to reduce copy-pasting, selectively hardens only trivially satisfied constraints to balance difficulty and naturalness, and evaluates satisfaction on both the generated article and its summary to penalize superficial instruction following. Consequently, we built the UNSPECIFIC benchmark on news, story, and blog domains to analyze the copy-pasting behavior of LLMs. Our results show that our synthesized constraints are not only more challenging (e.g., the satisfaction rate of GPT-5 Mini drops from 90% to 78%) and natural (LLM win-rate gap improves by 30%) from a human perspective but also mitigate the copy-pasting. We also find that a large portion of constraints are satisfied superficially (i.e., not satisfied in the core narrative of the article). The code and datasets are released at this https URL.
52. 【2608.09142】An Agentic Generative Large Language Model for Treatment Planning of Colorectal Cancer
链接:https://arxiv.org/abs/2608.09142
作者:Mengxian Lyu,Cheng Peng,Tim Jang,Ang Li,Mengyuan Zhang,Ziyi Chen,Leighton Elliott,Tianshi Liu,Lidice Galindo,Chiranjeevi Sainatham,Oscar F. Borja-Montes,Kaleb E. Smith,Ying Zhang,Lichao Sun,Jiang Bian,Gloria Lipori,Duane A. Mitchell,Elizabeth A. Shenkman,Yi Guo,Thomas J. George,Yonghui Wu
类目:Computation and Language (cs.CL)
关键词:ensure guideline-concordant care, precision oncology requires, oncology requires synthesizing, requires synthesizing heterogeneous, synthesizing heterogeneous patient
备注:
点击查看摘要
Abstract:Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure guideline-concordant care. While large language models (LLMs) show promise in many diagnostic tasks, their adoption for high-stakes treatment planning is hindered by complex reasoning, adherence to timely clinical guidelines, and safety concerns. In this study, we present GatorOnco, an agentic LLM for colorectal cancer (CRC) treatment planning. GatorOnco is developed using a total of 282 billion tokens of biomedical text, including healthcare system-scale clinical text comprising 166 billion tokens from UF Health. We implemented a domain-adaptation method that integrates pre-training, model merging, a two-stage post-training approach, and agent-based reinforcement learning. An agentic retrieval-augmented generation (RAG) approach dynamically integrates time-sensitive clinical guidelines into the reasoning process. In a blind, randomized clinical evaluation conducted by five UF Health oncologists, GatorOnco significantly outperformed open-source LLMs (P 0.01) and achieved expert-level performance comparable to UF Health oncologists. Compared with expert oncologists, GatorOnco received significantly higher ratings for readability (4.46 vs. 4.19, P 0.01) and completeness (3.91 vs. 3.52, P 0.01), while showing statistically comparable performance in correctness (4.09 vs. 4.11, P = 0.921), currency (4.04 vs. 3.98, P = 0.478), and safety (4.22 vs. 4.22, P = 0.999). These findings demonstrate that integrating agentic reasoning with large-scale domain adaptation can help bridge the gap for generative AI in high-stakes cancer treatment planning.
53. 【2608.09140】Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation
链接:https://arxiv.org/abs/2608.09140
作者:Li Siyan,Zhou Yu,Julia Hirschberg
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:Recent work, focuses on direct, personally-identifiable information, captured by standard, standard detectors
备注: Accepted into HAIPS workshop at COLM 2026
点击查看摘要
Abstract:Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information (PII) captured by standard detectors. One such approach is Privacy-Conscious Delegation (PCD), where a local LLM acts as an intermediary. However, privacy risk does not stem solely from explicit identifiers but also PII-free self-disclosures, leaving users identifiable through combinations of quasi-identifying traits. We investigate a probabilistic variant of PCD, where we augment its objectives with an LLM-driven probabilistic estimation of k-anonymity. To facilitate this, we first create the PUPA-SD dataset, which contains naturalistic user queries with self-disclosure. Our preliminary results indicate that optimizing PAPILLON on PUPA-SD improves quality on unseen conversations across a variety of local models and produces the best privacy-utility balance for Llama-3.2-3B, while smaller models struggle to jointly optimize quality and privacy. We propose k-anonymity as a useful auxiliary metric for tackling PCD.
54. 【2608.09128】Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
链接:https://arxiv.org/abs/2608.09128
作者:Keyu He,Xuhui Zhou,Maarten Sap
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
关键词:multi-agent social settings, improving LLM social, LLM social reasoning, social, increasingly deployed
备注:
点击查看摘要
Abstract:LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B's performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.
55. 【2608.09126】Subjective Multi-Bias Detection with Large Language Models
链接:https://arxiv.org/abs/2608.09126
作者:Ruiyu Li,Zhiying Zhu
类目:Computation and Language (cs.CL)
关键词:bias, pervasive challenge, subjective bias, bias detection, subjective biases
备注:
点击查看摘要
Abstract:In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identification of subjective bias, a type of bias that introduces improper attitudes or portrays a statement at odds with the actual truth. The subjective bias can jeopardize the authenticity and reliability of texts, leading to misconceptions and potential social tensions, especially when expressed through offensive language. Following prior work [1], we tackled with three different types of subjective biases in text: (1) framing bias with the use of one-sided words or phrases containing a particular point of view; (2) epistemological bias which includes subtle linguistic features that can affect the believability of the texts; (3) demographic bias with word/phrase usage under presuppositions of a particular demographic factor (i.e., gender or religion). In terms of the data we utilize, the input consists of texts that may harbor subjective biases. The output is a classification or annotation that reveals the presence or absence of such biases within the provided content. More specifically, we detected three different types of multi-span biases in corpus WIKIBIAS [2] with more than 4,000 sentence pairs from Wikipedia edits. The data is labelled by bias type for span pairs with the following categories: (1) framing bias, (2) epistemological bias, (3) demographic bias, and (4) no bias. The project codes are released at this https URL.
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2608.09126 [cs.CL]
(or
arXiv:2608.09126v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.09126
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
56. 【2608.09106】LexKairos: Benchmarking Legal Temporal Capabilities in LLMs
链接:https://arxiv.org/abs/2608.09106
作者:Chenyang Li,Zejia Feng,Yuqin Huang,Yuxiao Ye,Huiyuan Xie
类目:Computation and Language (cs.CL)
关键词:Large language models, Large language, demonstrated strong performance, demonstrated strong, wide range
备注: 15 pages, 5 figures
点击查看摘要
Abstract:Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that governs the validity of statutes, the progression of legal cases, and the enforcement of procedural deadlines. However, legal temporal capabilities remain underexplored in existing legal AI benchmarks. To address this gap, we propose LexKairos, a comprehensive benchmark for evaluating the temporal capabilities of LLMs in the Chinese legal context across three dimensions: statutory temporal knowledge, case temporal modeling, and statute-case temporal reasoning. LexKairos comprises nine sub-tasks drawn from real-world Chinese judicial cases and statutes. We conduct systematic evaluations of eight LLMs under multiple inference settings, including vanilla, Chain-of-Thought (CoT), and thinking modes. Our results show that Gemini-3-Flash achieves the strongest overall performance, yet even the best-performing model exhibits notable limitations on tasks demanding precise time-sensitive statutory metadata recall or complex reasoning in time limits, indicating that legal temporal knowledge and reasoning remain open challenges for current LLMs. Data and code are available at this https URL.
57. 【2608.09096】Evo-Bench: Can Language Models Improve Agent Harness?
链接:https://arxiv.org/abs/2608.09096
作者:Lisheng Huang,Chen Yang,Hao Zhou,Huatong Song,Zongchao Chen,Ran Le,Yang Song,Wayne Xin Zhao,Tao Zhang
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Large Language, driven rapid progress, static task solving, Language Models
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
58. 【2608.09093】he Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora
链接:https://arxiv.org/abs/2608.09093
作者:E. M. Freeburg
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:document arrangement, arrangement is written, notation, dataset card records, choices change model
备注: 44 pages, 10 tables, 7 figures. Pre-registered protocols and their amendment history ship with the repository. Code, data, and instruments: [this https URL](https://github.com/emfreeburg/announcement-carries-cue)
点击查看摘要
Abstract:How a document's arrangement is written down, its notation, is a training variable that no dataset card records. The field has established that text-extraction choices change model behaviour, and has never once measured the notation of what those choices put into the corpus. We define clean-window survival, a deterministic count of how much of a stream still demands the boundary inference, and measure notation on three fronts. What corpora carry: a census of thirteen public corpora, where survival falls to 0.153 in a vision-converted PDF slice against 0.889 in C4; the scarce resource is not unmarked text but long unmarked text; a pre-registered supply test finds what remains institutional, not consumer. Our own pre-registered prediction failed: converters do not fabricate structure on prose, and that null forced the reliability mechanism that survives it. What readers use: across five base models spanning 0.6B to 8.2B and two pipelines, deleting a structural announcement makes the following prose measurably harder to predict, while swapping its notation moves nothing. That zero does not make notation unimportant; it relocates the variable: the operative cue is the announcement, not the sigil. What writers impose: a bounded null. Base models do not impose the marked register above the authored baseline, and handed prose with every announcement deleted they do not put one back, at a rate indistinguishable from zero against an authored reference of zero. We ship the format those measurements imply: the pure frame, paragraphs in authored order, every announcement deleted into a reversible sidecar, mixed against the marked copy over announcement presence rather than notation. Choose format operators by the capability they train, not by the fidelity they preserve, and record extractor identity and survival on data cards.
59. 【2608.09080】When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
链接:https://arxiv.org/abs/2608.09080
作者:Maryam Tahermazandarani,Adnan Mahmood,Fahmida Islam,Quan Z. Sheng
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
关键词:Large Language Models, Large Language, achieved strong performance, clinical reasoning tasks, reasoning tasks
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.
60. 【2608.09049】Security and Privacy Taxonomy Generation from Mobile App Reviews
链接:https://arxiv.org/abs/2608.09049
作者:Moghis Fereidouni,Vinaik Chhetri,Umar Farooq,A.B. Siddique
类目:Computation and Language (cs.CL)
关键词:Mobile app reviews, continuously renewing source, users experience privacy, Mobile app, continuously renewing
备注:
点击查看摘要
Abstract:Mobile app reviews are a rich, continuously renewing source of how users experience privacy and security, yet existing taxonomies of these concerns are hand-crafted and cannot keep pace with the evolving nature of the data. Automating taxonomy construction is the natural response, but scalability is the core challenge: current LLM- and clustering-based methods are developed for scientific corpora of a few thousand documents and do not extend to app review collections numbering in the hundreds of thousands. We address this gap in two ways. First, we filter app reviews for privacy- and security-related content, yielding a comprehensive corpus of over 600K reviews. Second, we introduce TaxoScale, a pipeline that handles taxonomy construction at this scale by extending an expert-defined taxonomy via Recursive Hierarchical Clustering and LLM-based node naming. TaxoScale outperforms strong automatic-taxonomy baselines on path, level, coverage, and novelty metrics, and discovers novel branches absent from prior taxonomies.
61. 【2608.09046】Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities
链接:https://arxiv.org/abs/2608.09046
作者:Avijit Roy,Proma Roy,Hrishitva Patel
类目:Computation and Language (cs.CL); Computers and Society (cs.CY)
关键词:Large language models, Large language, increasingly deployed, deployed as general-purpose, Tokenization Equity Audit
备注: Accepted at IJCAI 2026 Workshop ( [this https URL](https://lm4uc.github.io/) )
点击查看摘要
Abstract:Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked. We introduce the Tokenization Equity Audit (TEA), a reproducible benchmark for measuring tokenization premiums in technical tutoring content. TEA evaluates three widely used tokenizers, GPT-4o's o200k base, Qwen2.5-7B, and Mistral-7B, on a 120-item Python debugging corpus translated from English into Bengali, Hindi, Arabic, Tamil, and Yoruba. Bengali and Hindi serve as the primary validated cases, while the remaining languages provide exploratory cross-script and cross-family comparisons. Across this corpus, Bengali requires (1.56\times) as many GPT-4o tokens as English, reducing a nominal 128k-token context window to an effective 82k-token English-equivalent capacity for the same semantic content. With the Qwen2.5 and Mistral tokenizers, Bengali requires up to (4.5\times) the English token count. Yoruba, despite using the Latin script, exhibits the highest GPT-4o tokenization premium at (2.37\times), indicating that tokenization inequity cannot be explained by script family alone. These results demonstrate that tokenization can create measurable economic and functional barriers, highlighting the need to treat tokenization as an equity-relevant infrastructure layer for underserved language communities, particularly where educational systems depend on low-cost or offline-capable AI tools.
62. 【2608.09045】Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production
链接:https://arxiv.org/abs/2608.09045
作者:Xiao Liu,Shiwei Gan,Yafeng Yin,Jiaxin Yin,Bowen Guo,Yaqi Sun,Zhiwei Jiang,Lei Xie
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:sign language recognition, sign language, language recognition, Recent advances, sign language understanding
备注:
点击查看摘要
Abstract:Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress. Meanwhile, sign language production (SLP), which generates sign sequences from text, has also attracted growing attention. This naturally raises an important question: can sign language understanding and production be unified within a single framework? Compared with unifying SLU subtasks, this problem is substantially more challenging. Existing SLU tasks largely share the same direction of mapping, namely from sign inputs to linguistic outputs, whereas SLT and SLP lie in opposite directions of sign-text mapping. A unified framework must therefore address two key challenges: (1) bridging the modality gap between continuous sign motions and discrete text tokens through a shared sign tokenizer that supports both linguistic abstraction and motion reconstruction; and (2) learning a single conditional autoregressive model that can take either sign or text as input and generate the corresponding target sequence in the opposite modality. To this end, we propose Uni-SLTP, a unified framework for SLT and SLP with two key components: (1) a shared sign tokenizer that converts sign sequences into discrete tokens and latent representations, capturing both semantic and reconstructive information; and (2) a unified autoregressive generation model that formulates both tasks as conditional sequence generation. Experiments on widely used public datasets show that Uni-SLTP achieves superior motion accuracy for SLP while maintaining competitive SLT performance.
63. 【2608.09044】ree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents
链接:https://arxiv.org/abs/2608.09044
作者:Zihao Deng,Yining Zhu,Leiming Wang,Jingfei Lu,Junbo Wang,Chuncheng Ran,Yu Yang,Dixuan Yang,Jikun Shen
类目:Computation and Language (cs.CL)
关键词:Continual self-evolution requires, self-evolution requires LLM, Continual self-evolution, requires LLM agents, transform environmental interactions
备注:
点击查看摘要
Abstract:Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related trajectories, but their experience representations are often disconnected from the underlying reasoning process. This limits feedback attribution, cross-task transfer, and update and retrieval efficiency, particularly in complex reasoning tasks with outcome-level feedback. To overcome this limitation, we propose \textbf{T}ree-\textbf{o}f-\textbf{E}xperience (ToE), a structured experience-management framework that aligns experience organization with the hierarchical reasoning process of LLM agents. Specifically, ToE organizes the experience into a shared tree of analytical perspectives and reasoning paths, whose reliability is calibrated through environmental outcomes to support systematic updating, transfer, and efficient retrieval. The experimental results on \textsc{Game of 24} and \textsc{FinEvolveBench} show that ToE substantially improves both problem-solving performance and efficiency. On \textsc{Game of 24}, ToE achieves a 31.4\% relative improvement in accuracy over the experience-free ToT baseline. On \textsc{FinEvolveBench}, ToE improves tsIC by an average of 41.24\% over the experience-free pipeline across 12 evaluation settings, whereas conventional experience-management methods often underperform experience-free baselines.
64. 【2608.09043】Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
链接:https://arxiv.org/abs/2608.09043
作者:Hyangsuk Min,Hwanjun Song
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:modern platforms repeatedly, Users of modern, modern platforms, platforms repeatedly, repeatedly need summaries
备注: 36 pages, 17 figures, 10 tables
点击查看摘要
Abstract:Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on its own. We formalize this setting as streaming dialogue summarization, where a system must summarize a current window using selective memory from an unbounded history under a fixed budget. We show that the central challenge is not how much history is accessed, but whether memory recovers the evidence that the current window presupposes. We construct a benchmark and evaluation protocol that separately assesses whether memory contains gap-resolving evidence and whether the generated summary reflects it. We propose ReMEMBER, a missing-evidence memory framework that conditions retrieval on unresolved window dependencies and refines retrieved chunks into evidence-dense memory under a fixed budget. Experiments on dialogues with histories up to 160K tokens show that ReMEMBER improves memory recall and gap-resolution completeness over memory construction baselines under the same budget.
65. 【2608.09028】PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs
链接:https://arxiv.org/abs/2608.09028
作者:Ponkrit Kaewsawee,Chaklam Silpasuwanchai,Chutiporn Anutariya
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB); Logic in Computer Science (cs.LO)
关键词:check compliance demand, compliance demand machine-readable, demand machine-readable constraints, Toggle, Institutional policies stay
备注: 23 pages, 3 figures, 4 tables. Under review at IJCKG 2026, Bangkok, Thailand
点击查看摘要
Abstract:Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that gap is still done by hand. PolicyKG closes the loop. It is an LLM pipeline that reads a policy PDF, classifies each sentence as an obligation, permission, or prohibition, lifts the label into first-order deontic logic, and emits SHACL constraints. Four stages run on a LangGraph state machine with per-stage validators. The piece that matters most is the Corpus Adapter: a YAML vocabulary registry that grounds LLM predicates in a target ontology. Retargeting to a new domain means swapping the registry, not retraining a model. On the Asian Institute of Technology Policies and Procedures corpus (1,663 sentences, 443 rules), PolicyKG reaches 86.9% deontic classification accuracy (Cohen's kappa = .709). Three annotators independently re-label a 50-item sample and agree at Fleiss' kappa = .844. SHACL shape correctness on a 69-shape subset is F1 = .866. The FOL path handles 79.2% of rules; the rest go through a direct NL-to-SHACL fallback. We audited every one of the 443 rules for second- or higher-order constructs. An automated regex checklist flagged none, and a first-author pass on the 92 FOL-fallback cases confirmed the same. The exact upper 95% Clopper-Pearson bound on the true HOL rate is 0.67%. This is an audit finding for one corpus, not a proof of FOL sufficiency for institutional policy. Swapping the AIT registry for a GDPR registry raises exact property alignment from 1/15 to 11/15 (Fisher's exact p .001; Cohen's h = 1.53). On the LexDeMod lease-contract benchmark (N = 200), Macro F1 drops to .370 because lease English uses "shall be entitled" for permission -- exactly the vocabulary mismatch registry swap is meant to fix. Repeated runs produce hash-identical SHACL outputs.
Comments:
23 pages, 3 figures, 4 tables. Under review at IJCKG 2026, Bangkok, Thailand
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB); Logic in Computer Science (cs.LO)
ACMclasses:
I.2.4; I.2.7; H.2.8
Cite as:
arXiv:2608.09028 [cs.AI]
(or
arXiv:2608.09028v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2608.09028
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Ponkrit Kaewsawee [view email] [v1]
Mon, 10 Aug 2026 02:28:57 UTC (146 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs, by Ponkrit Kaewsawee and 2 other authorsView PDFHTML (experimental)TeX Source
view license
Current browse context:
cs.AI
prev
|
next
new
|
recent
| 2026-08
Change to browse by:
cs
cs.CL
cs.DB
cs.LO
References Citations
NASA ADSGoogle Scholar
Semantic Scholar
export BibTeX citation
Loading…
BibTeX formatted citation
loading…
Data provided by:
Bookmark
checked="checked"class=“labs-tab-input”>
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.
Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)
mathjaxToggle();
We gratefully acknowledge support from
our major funders,
member institutions, ,
and all contributors.
About
Help
Contact
Subscribe
Copyright
Privacy
Accessibility
Operational Status (opens in new tab)
Major funding support from
66. 【2608.09024】ELICITED: EHR-grounded Longitudinal Interactive Conversations for Information-seeking Triage Evaluation and Decision-making
链接:https://arxiv.org/abs/2608.09024
作者:Haohao Zhu,Xiaolin Shi,Jiayu Zhou
类目:Computation and Language (cs.CL)
关键词:limited clinical resources, prioritize limited clinical, safely wait, prioritize limited, rapidly identify patients
备注:
点击查看摘要
Abstract:Emergency-department (ED) triage requires clinicians to rapidly identify patients who need immediate attention, determine who can safely wait, and prioritize limited clinical resources. At presentation, however, information may be limited to a chief complaint and initial vital signs. Clinically important details, including symptom onset and progression, associated symptoms, medical history, and medication use, are often obtained through focused conversation. Effective triage therefore requires clinicians to identify information gaps, ask appropriate follow-up questions, and update their assessment as new evidence becomes available. Most existing ED benchmarks evaluate acuity prediction from a fixed clinical snapshot. Although this formulation measures predictive performance after patient information has been assembled, it does not capture the interactive process through which triage-relevant evidence is elicited and interpreted. Existing medical dialogue datasets support the study of clinical communication, but dialogue statements are not always linked to temporally ordered events in the electronic health record (EHR). We introduce EHR2Dial-Triage, an agentic conversation-generation framework and benchmark grounded in MIMIC-IV-ED. The framework constructs triage conversations under explicit role-based and temporal information boundaries. Each accepted patient disclosure is linked to its supporting EHR event and the first dialogue turn at which it becomes available. EHR2Dial-Triage enables controlled evaluation of information elicitation, evidence use, five-level Emergency Severity Index prediction, and patient-facing communication across models and patient personas. It provides a structured setting for studying conversational triage as a dynamic process of clinical information acquisition, reasoning, and communication.
67. 【2608.09019】How People Evaluate AI-, Expert-, and Peer-Style Financial Advice
链接:https://arxiv.org/abs/2608.09019
作者:Aryan Ramchandra Kapadia,Eshwar Chandrasekharan,Koustuv Saha
类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Social and Information Networks (cs.SI)
关键词:Certified Financial Planner, evaluate AI-generated financial, including financial choices, people evaluate AI-generated, Online Community Forum
备注:
点击查看摘要
Abstract:As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
68. 【2608.08994】Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web Intelligence
链接:https://arxiv.org/abs/2608.08994
作者:Joshua Castillo,Santosh Nukavarapu,Ravi Mukkamala
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Retrieving relevant evidence, noisy web data, Retrieving relevant, heterogeneous language, data is challenging
备注: 8 pages, 2 figures. Accepted as a Short Paper at KDIR 2026 (International Conference on Knowledge Discovery and Information Retrieval)
点击查看摘要
Abstract:Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogeneous language, and irrelevant content. We present Guardian Crawler, a reproducible retrieval-first testbed for controlled experiments on knowledge discovery and evidence-grounded summarization over synthetic web-like corpora. The architecture combines BM25 retrieval with risk-aware, embedding-augmented, and hybrid reranking, followed by constrained retrieval-augmented generation with explicit document citations. Experiments on a synthetic 900-document corpus and 10 queries produced the highest descriptive retrieval scores under risk-based reranking, with P@10 = 1.00 and NDCG@10 = 0.94, compared with 0.94 and 0.81 for BM25. The best hybrid and BM25+Semantic configurations reached NDCG@10 values of 0.94 and 0.88, respectively. All 41 evaluable generated bullets passed the lexical coverage threshold; an automated LLM judge classified 36 as supported, one as partially supported, and four as unsupported. These results demonstrate the feasibility of Guardian Crawler as a controlled testbed but do not establish statistical superiority, human-validated faithfulness, or transfer to live-web investigative environments.
69. 【2608.08989】How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology
链接:https://arxiv.org/abs/2608.08989
作者:Wu Hangyu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Public infant cry, Public infant, infant cry corpora, cry corpora, Public
备注: 18 pages, 7 figures. Under review at AAAI 2027
点击查看摘要
Abstract:Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides and what fixes it. Across four cry corpora screened by a multi-level leakage audit (byte-level and embedding-level deduplication plus a within-corpus train-test near-duplicate audit), we probe four frozen encoders and a handcrafted baseline under a unified five-class need ontology and shared task formulations. The audit exposes what single-corpus evaluation conceals: within-domain macro-F1 swings by 0.57-0.80 for the same encoder, cross-corpus transfer is negative on average (negative-transfer ratio 0.19-0.35, significant in 18 of 30 directed cells, BH-FDR), and 349 content-identical clip groups carry conflicting metadata labels across corpus distributions. The same audit, however, reveals a consistent way forward. Transfer into the noisiest corpus is consistently positive in effect size at matched training size and after near-duplicate removal, offering a practical recipe for small, noisy corpora. Frozen probes saturate at modest label budgets, while stabilized fine-tuning wins with full labels; domain-adaptive pretraining significantly beats stabilized fine-tuning at 5-10-shot (the 1-shot advantage is not robust to optimization-seed variance) but shows no significant advantage at 50-shot or beyond. In the tested binary, shared-label settings, ontology-mapped joint training wins in all four encoder-by-target combinations, whereas naively merging unmapped labels costs up to 37 F1 points. We release the ontology, mapping code, and audit pipeline, turning incompatible cry corpora into a usable joint-training resource.
70. 【2608.08975】How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
链接:https://arxiv.org/abs/2608.08975
作者:Ming Li,Chenguang Wang,Xirui Li,Xinyue Zeng,Dianqi Li,Peng Shi,Dawei Zhou,Tianyi Zhou
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large language models, language models increasingly, models increasingly participate, choices shape AI-review, shape AI-review judgments
备注:
点击查看摘要
Abstract:As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.
71. 【2608.08942】Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access
链接:https://arxiv.org/abs/2608.08942
作者:Lier Jin,Lan Hu,Binqi Shen,Hanyu Cai,Yuting Xin
类目:Computation and Language (cs.CL)
关键词:everyday decision making, Large language models, Large language, public services, prompt
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide comparable assistance regardless of a user's literacy, communication style, or prompt-engineering expertise. However, existing research on prompt robustness primarily focuses on adversarial attacks, prompt injection, and prompt optimization, while overlooking whether semantically equivalent requests receive different responses simply because they are phrased differently. We refer to this accessibility challenge as "Prompt Privilege": users with greater prompting expertise systematically obtain better model performance despite expressing the same underlying intent. To address this problem, we present a unified framework for measuring and mitigating accessibility disparities in LLM interactions. We introduce Prompt Equity Score (PES), a quantitative metric for evaluating performance consistency across user populations, and Prompt Equity Transformer (PET), an LLM-based agent that automatically transforms user requests into semantically equivalent, accessibility-oriented prompts while preserving their intent. PET shifts prompt optimization from the user to the AI system, functioning as an intelligent accessibility layer between users and foundation models. Experiments on the MedQA benchmark demonstrate measurable prompt privilege, with statistically significant performance disparities between low-literacy and expert-prompting cohorts. Applying PET eliminates these disparities while preserving semantic fidelity, demonstrating that accessibility-oriented prompt normalization can improve equitable AI access. By introducing prompt privilege as a new dimension of AI accessibility and PET as a practical solution, this work advances system-centered accessibility and provides a foundation for more fair, trustworthy, and inclusive AI systems.
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2608.08942 [cs.CL]
(or
arXiv:2608.08942v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.08942
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
72. 【2608.08915】Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue
链接:https://arxiv.org/abs/2608.08915
作者:Esam Ghaleb,Hugh Mee Wong,Kristina Kobrock
类目:Computation and Language (cs.CL)
关键词:Situated language, Situated, multimodal, speech, gesture
备注:
点击查看摘要
Abstract:Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet dialogue models typically rely on transcripts alone. We study how much referential information gestures and their combination with speech carry in multimodal dialogue under different partner visibility conditions. % We build models that identify the intended referent in a video-mediated referential communication game based on either the speech transcript, the skeletal representation of gesture, or both modalities. Our results show that gesture alone is predictive of the intended referent and that multimodal fusion is most beneficial when the transcript-based model is uncertain. Training-only alignment of learned representations with the referent image further improves the fusion model performance. % In a comparison with human interaction data, we further see pragmatic effects of interlocutor visibility on gesture production and informativeness as well as an entrainment effect in speech and multimodal, but not gesture, performance across rounds of repeated interaction. We thus make contributions to the technical modelling of multimodal information in human dialogue and the analysis of human interaction data via trained model representations.
73. 【2608.08910】d Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving
链接:https://arxiv.org/abs/2608.08910
作者:Matteo Grella
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:PTQTP decomposes LLM, decomposes LLM weight, LLM weight matrices, free per-group scales, decomposes LLM
备注:
点击查看摘要
Abstract:PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses the decomposition into a single uniform nine-level quantizer, a known balanced-ternary identity. To our knowledge, at the time of writing, this work is the first to impose that identity as a constraint inside PTQTP's solver. The two trit planes then fold losslessly into one 4-bit code plane that we make the persistent serving representation: disk bytes, expert-cache bytes, and kernel input are the same 4.0625-bits/weight blocks, consumed in one integer dot pass. For this conjunction (ratio-3 nine-level code, CPU-SIMD kernels, SSD expert streaming, identical persistent bytes) we likewise found no precedent. We apply this to the routed experts of DeepSeek-V4-Flash-0731, a 284B-A13B mixture-of-experts model, quantizing in one shot from the released MXFP4 expert weights and streaming experts from SSD on a 64 GB laptop. Against a 4.5-bit Q4_K baseline, measured one process per fixture with an expert-lossless anchor arm as reference control, the tied model matches the official serving API on 5/5 fixtures at step 0 (Q4_K: 4/5) and 12/14 captured continuation steps (11/14), scores 86 vs. 84 on a 100-item MMLU subset, decodes 6.7% faster in decode phase, and ships 9% smaller files: no detected fidelity difference at these small evaluation sizes, and every fixture-level difference between the arms traces to a single measured near-tie cell. The tied fit nevertheless shows higher weight-reconstruction error and worse perplexity, a measured dissociation between proxy metrics and reference fidelity. A cumulative trunk-ternarization ladder and bitwise-pinned aarch64/x86-64 kernels complete the report. All code, formats, and evaluation artifacts are open source in the fucina inference stack.
74. 【2608.08881】heory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration
链接:https://arxiv.org/abs/2608.08881
作者:David M. Markowitz,Timothy R. Levine
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Retrieval-Augmented Generation, leading deception theories, current work developed, work developed, developed seven Retrieval-Augmented
备注:
点击查看摘要
Abstract:The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models. Across 700 statements drawn from five published deception datasets, four large language models (gpt-4o, claude-sonnet-4-6, ollama/llama3, deepseek-v4-flash), and two run-types (RAG vs. baseline), a total of 39,200 deception judgments were rendered. Detection accuracies were consistent with typical human accuracies and not statistically different across RAG (54.5%) and baseline models (54.6%). RAG-based models (57.0%) were less truth-biased than baseline models (59.7%), but the effect size was quite small. Theoretical perspective mattered little for accuracy yet mattered substantially for response bias, which ranged from highly lie-biased (the verifiability approach, 32.2%) to highly truth-biased (truth-default theory, 88.1%). Content effects and model effects further moderated the results. Theory-guided AI judgments are unreliable with current parameters, yet they might show promise with additional datasets, model testing, and theory-to-data matching.
75. 【2608.08869】Position Bias in Ordinal Classification: A Systematic Evaluation
链接:https://arxiv.org/abs/2608.08869
作者:Yu Wang,Jeffrey Zhou,Menglin Liu,Ge Shi
类目:Computation and Language (cs.CL)
关键词:Large language models, Large language, ordinal classification, alter their predictions, semantically equivalent
备注:
点击查看摘要
Abstract:Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from label order, demonstration order, and demonstration placement. First, we apply the three probes to ten frontier LLMs on a common ordinal-classification task; every model is sensitive to all three positional sources, showing that the problem is pervasive. Second, we vary eight prompt-, task-, and model-level factors across five datasets; accuracy and stability are often misaligned, and only lower scale cardinality consistently improves both. Third, we compare pointwise, pairwise, and listwise inference, alternative aggregation and debiasing methods, and joint configurations; the tested corrections do not provide a reliable remedy, while a comparison-based listwise formulation offers the best balance but transfers unevenly across models and bias sources. These findings show that positional robustness depends on the full system configuration rather than the model alone. Ordinal-classification systems should therefore be selected jointly for predictive performance and stability.
76. 【2608.08868】Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State
链接:https://arxiv.org/abs/2608.08868
作者:Lily Chen,Ted Mau,Michael Gensheimer,Brian Anthony Nuyen,Nancy Jiang,James Zou
类目:Computation and Language (cs.CL)
关键词:systems analyze conversational, analyze conversational traces, implicitly assuming, recoverable from conversation, modern AI systems
备注: COLM 2026
点击查看摘要
Abstract:Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recoverable from conversation. We study observability: whether a target is recoverable from conversational transcripts alone. Observability is difficult to assess because transcripts may provide only a partial view of many targets, and large-scale analysis requires model-based annotation, making true limits of the conversational signal hard to distinguish from annotator error. We therefore study clinical encounters, where patient-reported outcome measures (PROMs) provide an external anchor for patient state, and visits follow broadly structured patterns. We study observability of patient state and conversational phase structure using 439 real-world clinical encounter transcripts spanning 134 hours, including 245 ENT transcripts paired with 273 PROM surveys. We operationalize patient state using PROM scores for voice, cough, and swallowing; phase structure using conversational phase segmentation. To make these analyses credible at scale, we use a PHI-compliant GPT-5 deployment for transcript annotation and conduct 40 hours of manual validation, reducing the risk that apparent limits of observability simply reflect annotator error. Our core finding is an observability asymmetry: phase structure is observable and useful for characterizing clinical encounter organization, while patient state is only partially observable, even in a setting designed to elicit patient symptoms and experiences, cautioning against transcript-only inference of human state.
77. 【2608.08847】Explicit Boundary Markers for Subword Vocabularies
链接:https://arxiv.org/abs/2608.08847
作者:Sander Land,Clara Meister
类目:Computation and Language (cs.CL)
关键词:Subword tokenizers represent, space-using writing systems, Subword tokenizers, writing systems, tokenizers represent
备注:
点击查看摘要
Abstract:Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have separate embeddings in models, so occurrences of one word are divided across rows that are trained independently, and the two forms need not even segment the string the same way: " together" may be a single entry while the same word without a preceding space is tokenized as "to|gether". Capitalization divides a word further, into as many as six forms. We introduce an alternative to standard whitespace conventions using an explicit word boundary marker, which prevents such duplication. Words are delimited by the boundary markers, and spaces between words are represented as pairs of such markers. Two shift codes do the same for title case and upper case, allowing one internal representation of a word to be re-used across different settings. Switching to this convention mitigates the duplicate-entry issue, but does not improve tokenization compression: for both vocabulary-learning algorithms, the best marker scheme stays within one percent of the baseline in characters per token, averaged across six languages. It does result in better language modeling performance. Every marker scheme tested downstream reaches lower bits per byte than the baseline, suggesting that duplication carries a cost that compression does not capture.
78. 【2608.08829】Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
链接:https://arxiv.org/abs/2608.08829
作者:Muhammad Faishal Adly Nelwan,Alfan Farizki Wicaksono
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Activation steering edits, current practice fixes, injection layers globally, frozen language model, Activation steering
备注: 43 pages, 24 figures, 30 tables. Under review at ACL Rolling Review (August 2026)
点击查看摘要
Abstract:Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task. We argue that the best layers are an instance-level decision, and we make per-instance, multi-layer selection both well understood and deployable. On two open-weight 8B models and six binary persona traits, a per-instance oracle over layer subsets shows that the best layers vary from one input to the next: on most trait-model pairs, no fixed global layer set recovers the per-instance benefit. A greedy rule that ranks layers by single-layer marginal effect recovers nearly all of the oracle's benefit, but both must score candidate layers against the gold answer, so neither can run at deployment; the rule instead becomes the target a prompt-only predictor is trained to reproduce. Our deployable recipe needs no label at inference: a per-instance layer ranker read off the prompt embedding, a classifier that infers the steering direction, and an adaptive gate that scores short steered passes against that inferred direction and steers no more layers than necessary. The recipe recovers most of the oracle's lift (the bulk on the stronger model, a clear majority on the harder one), never drives any trait-model pair below its unsteered alignment baseline on average, and largely avoids the fluency collapse that strong global selection incurs at higher layer counts. A mechanistic account, "direction over magnitude", explains the behavioural flip under a mis-directed global set, the output collapse from steering too many layers, and the ceiling of unsteerable inputs.
79. 【2608.08822】Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models
链接:https://arxiv.org/abs/2608.08822
作者:Abdalla Doleh,Toni Somers,Ratna Babu Chinnam
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:decision-making research depends, carefully controlled complexity, Cognitive decision-making research, production is slow, decision-making research
备注: 38 pages, 7 figures
点击查看摘要
Abstract:Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured decision scenarios and validates their complexity through a composite framework rooted in established task-complexity theory. We evaluated 4,238 scenarios across multiple domains and complexity tiers. Measurement validation met rigorous psychometric standards. Agreement among five independent model families was nearly perfect, with an intraclass correlation coefficient of 0.997 and a kappa of 0.971. Known-groups validity demonstrated large separation between tiers, with an eta-squared of 0.587 and all pairwise comparisons significant at p less than .001. Factor analysis revealed a dominant complexity construct, with loadings between 0.87 and 0.96 across three frameworks, while interactivity formed a weaker secondary dimension at 0.34. Discriminant validity was limited by a strong relationship between complexity and text length that persisted after controlling for tier, yielding a partial correlation of 0.86. This constrains construct purity but does not undermine the instrument's tier-grading function. Model analyses showed a negative association between throughput and schema pass rate (r = -0.967, p = .007, n = 5), suggesting a speed-quality trade-off, though largely driven by one high-throughput model. Llama 4 Maverick generated scenarios fastest at 134 per minute versus 25 for DeepSeek Chat V3.2, but underproduced complex-tier scenarios, whereas DeepSeek Chat V3.2 balanced domain coverage with high schema compliance. The system demonstrated strong psychometric properties, enabling reliable classification into Simple, Moderate, and Complex tiers and providing the measurement infrastructure needed for downstream cognitive assessment of AI systems
80. 【2608.08809】vatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers
链接:https://arxiv.org/abs/2608.08809
作者:Yu Wang,Shengyao Zhuang,Xueguang Ma,Zongyu Wu,Jimmy Lin,Vivek Srikumar,Zhichao Xu
类目:Computation and Language (cs.CL)
关键词:model scale challenges, production retrieval system, single model scale, scale challenges, challenges the flexibility
备注:
点击查看摘要
Abstract:A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with the workload. In the context of information retrieval (IR), a transformer-based model can be made smaller in three ways---using fewer layers, passing fewer tokens through the upper layers, or producing a shorter embedding---and each way saves a different compute resource. These options have been studied one at a time, each as its own method with its own code and training setup, which makes them hard to combine or adapt to a new model. We present~\ours to bring all three under one simple abstraction: a single object names any size the model can run at, and a short schedule lists the sizes to train. Training then produces one checkpoint that serves all of those sizes, and at deployment the user picks any of them. The same abstraction covers both retrievers and rerankers and both encoder and decoder models, as it works through interfaces that Hugging Face transformers already expose; a new backbone is a configuration change, not new modeling code. Prior methods---Matryoshka embeddings, early exit, 2D~Matryoshka (e.g., Starbucks), and layerwise token compression---become special cases of our unified abstraction. The same interface also enables Matryoshka~LTC (MLTC), which jointly trains several token-compression ratios in one retriever checkpoint. To validate our framework, we train 20 checkpoints across three backbones and two tasks: the quality curves are smooth, one checkpoint costs little over a model trained for a single size, and a controlled study confirms the wallclock speedups. We release the framework and all checkpoints as a resource for building elastic retrieval systems.
81. 【2608.08801】IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements
链接:https://arxiv.org/abs/2608.08801
作者:Shiva Ahir
类目:Computation and Language (cs.CL); Hardware Architecture (cs.AR); Emerging Technologies (cs.ET)
关键词:altering numerical constraints, Translating technical requirements, introduce semantic drift, Semantic Requirement Representation, Translating technical
备注:
点击查看摘要
Abstract:Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other specification-critical meaning. IDRAAK is presented as an interpretable framework for detecting such drift using a language-independent Semantic Requirement Representation (SRR), with six detection workflows evaluated, ranging from deterministic comparison to multi-agent verification and few-shot prompting. On 890 synthetic perturbations across 300 requirements from 10 engineering domains, a single LLM call with six few-shot examples achieves MCC=0.888 and F1=0.983, outperforming the evaluated structured and multi-stage alternatives. Further evaluation on PAWS-X (805 pairs, 5 languages) and XNLI (700 pairs, 7 languages) exposes complementary strengths and limitations of structured and LLM-based approaches. Deterministic SRR comparison performs strongly on technical requirements (F1=0.898) but poorly on general-domain text (F1=0.012), while structured evidence improves performance on adversarial paraphrases. Post-hoc Platt scaling further improves confidence calibration. The results demonstrate that increased agentic complexity does not necessarily improve semantic-drift detection and that simple few-shot prompting can provide a strong and efficient alternative.
82. 【2608.08800】Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages
链接:https://arxiv.org/abs/2608.08800
作者:Sofiia Riazhskykh,Nam Luu,Ondřej Bojar
类目:Computation and Language (cs.CL)
关键词:reportedly increase token, increase token efficiency, training tokens needed, LLMs on artificial, reportedly increase
备注:
点击查看摘要
Abstract:Pretraining LLMs on artificial languages ("pre-pretraining") is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33% of training tokens needed to reach a certain performance. We validate this prior result for English on a larger set of natural languages across four language families, using two different tokenizers and varying model sizes. We also relate the observed gains (or losses) in token efficiency to quantified linguistic properties of the languages, such as sentence length, morphological richness, and features of dependency syntactic trees (tree depth, number of children, number of crossing dependencies). Our empirical results indicate that the reported gains depend heavily on the experiment setup and the choice of random seed, although we can confirm the trend of stable gains with 128-Dyck pretraining of small models with the Llama tokenizer for most of the examined languages. On a general note, we argue that multiple training runs should be carried out at least for a subset of experiments to avoid the community adopting unstable approaches.
83. 【2608.08795】oward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection
链接:https://arxiv.org/abs/2608.08795
作者:Sihan Hou,Xinmeng Hou,Zhijun Zhang,Zehao Wang,Xuhong Ren,Sibo Qin,Kuntharrgyal Khysru,Qing Guo
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:Tool-using large language, malicious instructions embedded, external observations manipulate, observations manipulate subsequent, manipulate subsequent agent
备注:
点击查看摘要
Abstract:Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external observations manipulate subsequent agent decisions and actions. Most existing adaptive attacks rely on repeatedly querying and refining against the target agent, whereas realistic attackers may have only a single opportunity to interact with an unknown target agent. We propose SAVOR (Strategy Abstraction Via Outcome-Conditioned Reflection), which shifts attack adaptation from test-time iteration to offline strategy distillation. SAVOR performs outcome-conditioned reflection over successful and failed trajectories collected from disjoint training environments, validates context-conditioned candidate strategies, and iteratively consolidates them into a reusable strategy memory. At test time, the frozen memory guides the generation of a single payload for each unseen target, requiring only one target-agent query and no target-agent feedback. Across two benchmarks and three victim models, SAVOR attains the highest average attack success rate in all six settings, leading the strongest prior attack by 2.5 to 11.8 points and the same injection channel without strategy learning by 23.1 points on Agent Security Bench, which holds out attacker tools, and 28.6 points on OpenClaw-IPI, an executable benchmark we introduce that holds out attack goals and verifies attacks through tool interactions and execution receipts. A memory learned under one defense also transfers to another.
84. 【2608.08793】Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents
链接:https://arxiv.org/abs/2608.08793
作者:Xueping Gao
类目:Computation and Language (cs.CL)
关键词:Skills package reusable, package reusable instructions, tool-using language-model agents, Agent Skills package, Skill Runtime Intelligence
备注: 17 pages, 1 figure, 6 tables. Submitted to PROFES 2026. Code and artifacts: [this https URL](https://github.com/hellogxp/skill-runtime-intelligence)
点击查看摘要
Abstract:Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly represented by session-, model-, or tool-centric traces: a Skill can be discovered but not activated, activated without instructions, or appear successful without an independently verified outcome. We present Skill Runtime Intelligence, a passive runtime-intelligence system that reconstructs supported Skill-lifecycle stages across heterogeneous harnesses while preserving unsupported stages as unknown. Its Run Panorama separates immutable events, deterministic relations, inferred diagnoses, and controlled outcomes with four evidence grades; optional trace import and OTLP/HTTP export support existing observability deployments. Across six frozen repository profiles, three coding agents, and seven clean or fault-injected conditions, all 126 executions preserve source worktrees and each correlates to exactly one source session. Yet adapters expose three distinct semantics: no Skill runs; complete runs but no failure-like events; or failure-like events in every operational-failure and clean session. In a seven-template diagnostic study, semantic aliases and Panorama localize the same six non-clean boundaries but differ in exact/status behavior; both Raw views emit a failure status on all 18 clean cases, while Panorama emits none. A known-rule graph conforms to 126/126 frozen contracts, whereas a second model completes only 228/378 calls. These observations motivate executable adapter qualification and show that event presence is not boundary fidelity, composite exact scores mask distinct errors, and model explanations must not overwrite deterministic facts.
Comments:
17 pages, 1 figure, 6 tables. Submitted to PROFES 2026. Code and artifacts: this https URL
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2608.08793 [cs.CL]
(or
arXiv:2608.08793v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.08793
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
85. 【2608.08791】Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models
链接:https://arxiv.org/abs/2608.08791
作者:Saurabh Yadav,Badri Narayana Patro,Vijay Srinivas Agneeswaran
类目:Computation and Language (cs.CL)
关键词:Diffusion language models, broad context, context to create, Diffusion language, handle input noise
备注:
点击查看摘要
Abstract:Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this is only partially true. Internally, diffusion models detect text errors highly accurately. Externally, their reported certainty ignores this signal. As accuracy drops due to noise, confidence stays near its maximum and the ability to correctly rank answers degrades toward random chance. We call this mismatch the representation confidence gap. The visible concentration of high certainty scores is a misleading surface symptom. Standard math adjustments remove this concentration but fail to fix the underlying loss of ranking order. This ranking deficit favors standard models under noisy conditions and resists common remedies. Matching training recovers accuracy but not ranking, while score recalibration and input level error signals cannot reorder the final answers. However, the information needed to properly evaluate an answer survives in the hidden states. A lightweight extraction tool uses this signal to improve ranking. This approach is highly efficient because it leaves the base model completely frozen and requires zero additional text generation steps. We present this tool to prove the signal exists, while clearly noting its limits. Ultimately, certainty reliability is a more pressing limit than overall accuracy under noisy conditions.
86. 【2608.08775】OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents
链接:https://arxiv.org/abs/2608.08775
作者:Andrea Caciolai,Pere-Lluís Huguet Cabot,Chierh Cheng,Albert Ventayol-Boada,Gabriel Mejia Gonzalez,Christophe Ropers,Lucas Bandarkar,Sebastian Ruder,Darlene Sakakihara,Elliot Yun,Pierre Andrews,Grégoire Mialon,Romain Froger,Marta R. Costa-jussà
类目:Computation and Language (cs.CL)
关键词:realistic multi-tool environments, Agentic benchmarks aim, exclusively in English, multi-tool environments, aim to measure
备注:
点击查看摘要
Abstract:Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.
87. 【2608.08772】Multilingual Emotion Neurons in Large Audio-Language Models
链接:https://arxiv.org/abs/2608.08772
作者:Xiutian Zhao,Philipp Koehn,Björn Schuller,Berrak Sisman
类目:Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
关键词:human communication, central to human, expression varies, Emotion, languages
备注:
点击查看摘要
Abstract:Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.
88. 【2608.08744】Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs
链接:https://arxiv.org/abs/2608.08744
作者:Sourav Das,Tanmay Joshi,Kripabandhu Ghosh
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Emerging Technologies (cs.ET); Machine Learning (cs.LG)
关键词:deployed Large Language, Large Language Model, Large Language, deployed Large, model substantially exceeds
备注: 13 Pages, 6 Figures, Submitted to ARR Cycle
点击查看摘要
Abstract:The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds the one-time cost of fine-tuning. Yet most efficiency interventions target either pre-training scale or post-hoc compression. We ask whether folding a calibrated, differentiable energy surrogate into the fine-tuning objective can produce inference behavior that gains task accuracy at zero or near-zero carbon cost, a break-even configuration. We propose a joint loss mechanism with a per-model carbon-emission parameter, a linear surrogate over parameter norm, FLOP proxy, and a memory proxy, fit from on-hardware energy profiling. We fine-tune three architecturally distinct families: Gemma-2 2B, Llama-3.1 8B, and Qwen-2.5 14B, and evaluate inference F1 and CO$_2$ emissions on three MMLU subjects: abstract algebra, philosophy, and formal logic. We discover from several outcomes that the carbon term behaves as either harmful interference or beneficial regularization depending on the task structure. We position calibrated carbon-aware fine-tuning as a lightweight, drop-in regularizer with a non-empty but model and task-dependent break-even region. This is an ongoing work, and we will release our codebase soon.
89. 【2608.08732】AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval
链接:https://arxiv.org/abs/2608.08732
作者:Haoyu Zuo,Yibo Yan,Xin Zou,Shuliang Liu,Yi Cao,Mingdong Ou,Xuming Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Multi-vector vision-language retrievers, fine-grained Visual Document, incurs substantial overhead, vision-language retrievers enable, retrievers enable fine-grained
备注: 24 pages, 7 figures
点击查看摘要
Abstract:Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive compression, whereas merging does not explicitly prioritize important regions when forming representatives. We introduce AnchorFold, a training-free focus-then-fold framework for document-side index compression. AnchorFold applies Recursive Attention Propagation over visual self-attention graphs, performing multi-step propagation within each attention head and integrating scores across heads and layers. The focus stage selects the highest-centrality tokens as anchors. The fold stage assigns remaining tokens to their most similar anchors in the normalized retrieval space and summarizes each anchor-centered group through centrality-weighted aggregation. This preserves non-anchor contributions while concentrating capacity on structurally important tokens. Across ViDoRe v1/v2 and REAL-MM-RAG with three diverse retrieval backbones, AnchorFold consistently outperforms all evaluated training-free baselines at $\gamma \leq 0.20$. On ViDoRe v1/v2, it retains 98.3% of full-index NDCG@5 on average at $5\times$ compression, achieving near-lossless compression, and 92.4% at $20\times$ compression.
90. 【2608.08721】LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization
链接:https://arxiv.org/abs/2608.08721
作者:Zexun Lin,Yuan Feng,Junlin Lv,Kevin S. Zhou,Xike Xie
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:accelerates large language, efficiency critically determined, decoding accelerates large, large language model, language model inference
备注:
点击查看摘要
Abstract:Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency critically determined by the speculative length selected at each decoding round. Existing dynamic speculation methods select the speculation length by estimating how many tokens will be accepted, which is reasonable for autoregressive drafters that generates tokens sequentially. The recent wave of diffusion-based drafters, however, generates candidate blocks in parallel at substantially lower drafting cost, shifting the key question from how many tokens to generate to how many generated tokens are worth verifying. We therefore reformulate dynamic speculative-length selection as expected-speedup optimization and derive a marginal criterion that extends the speculative sequence only when its acceptance gain outweighs the additional verification cost. Building on this criterion, we develop \textit{LibraSpec}, a training-free and plug-and-play algorithm that iteratively determines the speculative length using drafter confidence scores. Theoretically, we prove that LibraSpec monotonically converges toward the optimal speculative length. Experiments across six target models, three diffusion-based speculative decoding methods, and math, coding, and chat benchmarks show consistent improvements under both greedy and sampling settings, achieving a further $0.5\sim1.5\times$ improvement over baselines and up to $8.49\times$ speedup over autoregressive decoding.
91. 【2608.08650】he Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism
链接:https://arxiv.org/abs/2608.08650
作者:Jiguo Li
类目:Computation and Language (cs.CL)
关键词:models increase parameter, increase parameter capacity, models increase, model releases, capacity while keeping
备注:
点击查看摘要
Abstract:Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution cannot be explained by a chronological list of model releases alone. This technical survey synthesizes primary papers, official technical reports, and prior surveys to organize modern Mixture-of-Experts systems along five coupled dimensions: expert granularity, expert topology, routing freedom, the scope of load balancing, and execution structure. We describe eight architectural milestones as a dependency graph with six mainline developments and two orthogonal branches, rather than as eight successive generations. We then analyze individual systems through four control planes: Expert Topology, Routing, Balance, and Expert Parallelism. These planes specify which experts exist, which experts process each token, how aggregate load is controlled, and how selected computation is mapped onto physical devices. The framework connects algorithmic choices such as Top-k routing, shared experts, fine-grained experts, and dynamic expert composition with systems concerns including token dispatch, device placement, all-to-all communication, and communication-computation overlap. We conclude with equal-budget pretraining experiments, quality and systems metrics, and open research questions. The main trend is a shift from merely activating more sparse parameters toward decoupling semantic routing, computational budgets, and physical execution.
92. 【2608.08638】CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
链接:https://arxiv.org/abs/2608.08638
作者:Yuqian Zhang,Yao Shi,Kexin Huang,Botian Jiang,Zhe Xu,Yiwei Zhao,Min Liang,Shuang Chen,Xipeng Qiu
类目:ound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:supports interactive assistants, personalized media, interactive assistants, accessibility tools, supports interactive
备注: 20 pages, 6 figures, 10 tables. Technical report
点击查看摘要
Abstract:Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require faithful linguistic rendering, consistent speaker identity, and low-latency response. Yet compact streaming systems must preserve sufficient acoustic detail in a predictable low-rate latent sequence, while iterative diffusion sampling and classifier-free guidance multiply inference cost at every autoregressive step. To strike a balance between high-fidelity synthesis and low-latency inference, we present CuteTTS, a compact continuous-autoregressive TTS system. It combines semantically aligned causal VAE latents with patch-level autoregression, explicit speaker conditioning, and a bidirectional flow-matching head. We further introduce guidance-step distillation, which absorbs classifier-free guidance and multiple solver steps into a single interval-conditioned student. Evaluations on LibriSpeech and Seed-TTS-Eval demonstrate competitive intelligibility and speaker similarity in zero-shot voice cloning, while distillation lowers first-audio latency by 23.3% and real-time factor by 40.8% relative to the base model with comparable objective and subjective quality. These results provide a practical path toward continuous-autoregressive TTS that reconciles high-fidelity generation with the latency demands of real-time interaction.
93. 【2608.08636】Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach
链接:https://arxiv.org/abs/2608.08636
作者:Tong Bao,Yi Zhao,Heng Zhang,Chengzhi Zhang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
关键词:entity, entity type, entity type information, plays a crucial, named entity recognition
备注:
点击查看摘要
Abstract:Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts. Recently, large language models (LLMs) have demonstrated the capacity to achieve competitive SciNER performance with minimal human effort. Existing research highlights the importance of incorporating candidate entity type information for accurate entity recognition and classification by LLMs. However, when too many candidate entity types are provided in the prompt, LLMs struggle to accurately recognize and label entities in scientific texts, where entity types are more complex than in general domains. To address this challenge, we propose TdSciNER, a type-driven approach that effectively leverages entity type information to enhance SciNER performance. In TdSciNER, we first design an entity type filter model to identify the most likely entity types present in a given sentence. Subsequently, we introduce an auxiliary multi-class entity typing task within a multi-task learning framework alongside SciNER to obtain richer contextual representations. Then, we develop a novel demonstration selection strategy based on sentence similarity and entity type diversity to activate the in-context learning capabilities of LLMs, thereby improving entity recognition accuracy across diverse scientific domains. Experiments on three datasets demonstrate that our method achieves performance comparable to fully supervised models. Further analysis validates that each entity type-driven component in TdSciNER contributes to the improvement of SciNER performance. This work provides valuable insights for future advancements in SciNER and broader information extraction tasks in scientific text mining.
94. 【2608.08634】Can Open-Weight Models Compete on Financial Text Comprehension?
链接:https://arxiv.org/abs/2608.08634
作者:Jan Spörer
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); General Finance (q-fin.GN)
关键词:labs caught, Open-weight language models, proprietary frontier models, Financial Touchstone benchmark, recent months
备注: To be presented at the workshop International Symposium on Large Language Models for Financial Services (FinLLM@IJCAI2026)
点击查看摘要
Abstract:Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba's proprietary flagship Qwen3-Max. Anthropic's Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google's Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.
95. 【2608.08618】RAG-Based Auto-Configuration for Industrial Fieldbus Devices
链接:https://arxiv.org/abs/2608.08618
作者:Aadil Gani Ganie,Saad Ezzini,Naveed Farooz Marazi
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Industrial device commissioning, heterogeneous PDF manuals, commissioning requires engineers, manually extract hundreds, device commissioning requires
备注:
点击查看摘要
Abstract:Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and transcribe them into supervisory control systems, a time-intensive, error-prone workflow. This paper presents SysName, a production-oriented pipeline that automates device configuration end-to-end for Modbus RTU, OPC-UA, Profibus DP, and CANopen. It builds a hybrid dense-sparse retrieval index augmented by an ontology graph derived from ECLASS, AAS, and SOSA/SSN, using a BGE-M3 encoder with a cross-encoder reranker to surface relevant manual passages. A local LLM (T=0.1) generates ontology-aligned JSON-LD configurations via protocol-specific prompts and a four-step repair pipeline. A two-stage abstention gate, combining a reranker-score threshold and an IRI resolution ratio, blocks unsafe LLM invocations and filters low-coverage configurations before SHACL validation. On a gold set of 28 field-level queries, the hybrid retriever reaches 0.96 HitRate@10, and the reranker raises MRR@10 from 0.56 to 0.63 with perfect score separation for abstention. The generator attains field-level F1=0.87 with exact match on 9 of 12 runs. End-to-end runs on an H100 GPU complete in 2.6-6.6s per device with zero unsafe writes and zero silent failures on a five-device benchmark; every unsuccessful run is flagged by abstention or deployment verification. Component-wise evaluation localises the single systematic failure to OPC-UA generation, invisible to end-to-end metrics alone. A case study commissions a physics-simulated Universal Robots UR5e robot from unmodified vendor documentation (254-page manual, 8-page register list, 496 chunks), reaching field-level F1=1.0 over three runs with read-back and joint-consistency verification. An ablation study and comparison with five industrial-LLM systems complete the analysis.
96. 【2608.08607】North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings
链接:https://arxiv.org/abs/2608.08607
作者:Meriem Laifa,Abdallah Bengueddoudj
类目:Computation and Language (cs.CL)
关键词:Natural Language Processing, Language Processing, Algeria mental healthcare, mental healthcare, Natural Language
备注:
点击查看摘要
Abstract:Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained concentrated in high-income, English-language settings. North Africa, and Algeria in particular, is largely absent from this literature despite its unique linguistic, historical, and healthcare context. We present the first conceptual framework examining the potential role of NLP within Algeria's mental healthcare system. Drawing on narrative synthesis of global NLP mental health research, Algerian healthcare literature, and low-resource NLP methodologies, we identify four structural barriers to mental healthcare: the language-of-care gap, geographic inequities in access, stigma-related barriers to help-seeking, and the absence of research and digital infrastructure. We then map existing NLP capabilities to each barrier, outlining their potential applications, implementation constraints, and the technical, institutional, and governance requirements necessary for deployment. Based on this analysis, we propose a research and policy roadmap that prioritizes data resources, multilingual language technologies, evaluation frameworks, and regulatory capacity. Although grounded in the Algerian context, the framework addresses challenges common to many multilingual, low-resource, and post-colonial settings. This work provides a foundation for future research on culturally and linguistically appropriate NLP for mental healthcare and offers a practical roadmap for developing responsible AI-enabled mental health systems in underrepresented regions.
97. 【2608.08606】Mitigating Gender Bias in English to Romanian Machine Translation
链接:https://arxiv.org/abs/2608.08606
作者:Ioana Grigore,Sergiu Nisioi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:correctly translate gender, gendered target language, fail to correctly, correctly translate, Machine translation
备注:
点击查看摘要
Abstract:Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations that default to masculine forms or reinforce gender stereotypes. We propose a hybrid pipeline to mitigate this issue by combining large language model (LLM)-based gender classification with neural machine translation (NMT). Our system uses a fine-tuned LLM to detect the intended gender of target words in English sentences and insert inline gender hint tags. These tagged sentences are then passed to a Transformer model fine-tuned to generate morphologically correct Romanian translations. To support this, we introduce three novel datasets for gender disambiguation and translation. Our approach improves gender accuracy on the WinoMT and WinoGender benchmarks by over 40 percentage points compared to a baseline MT system. This is the first method to explicitly address and evaluate gender bias in English-Romanian MT using both LLM inference and tag-aware translation.
98. 【2608.08557】OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories
链接:https://arxiv.org/abs/2608.08557
作者:Changhao Xiang,Shilin Zhang,Zheng Ma,Kanzhi Cheng,Ruize Ma,Yi Feng,Jianbing Zhang,Zhi Wang,Zhen Wu,Xinyu Dai,Lewei Lu
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:fixed image encoding, actively acquire evidence, image encoding, multimodal agents, agents to actively
备注:
点击查看摘要
Abstract:Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.
99. 【2608.08510】From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
链接:https://arxiv.org/abs/2608.08510
作者:Thai-Binh Nguyen,Zhaolin Li,Jan Niehues,Alexander Waibel
类目:Computation and Language (cs.CL)
关键词:spontaneous informal conversations, remarkable ability, ability to engage, engage in spontaneous, spontaneous informal
备注: Accepted at ICMI 2026
点击查看摘要
Abstract:Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out competing speech from nearby conversations. This "cocktail party" scenario still presents severe challenges to speech recognition systems. The CHiME-9 MCoRec task provides a testbed where systems must recognize groups of speakers and transcribe each of their conversations from audio-visual input. In this work, we analyze a diverse set of systems, representing different design directions for addressing the cocktail-party scenario, where the best system achieves up to 57% relative error reduction. We identify three main strategies: (1) explicit or implicit audio-visual target speech separation, (2) improved audio-visual speech recognition for each target speaker, and (3) the use of large language models to group speakers into conversations and enhance conversational consistency. Our analysis shows that these directions address complementary failure modes of the cocktail-party problem, and that high speech overlap alone does not explain performance differences, challenging the common assumption that overlap is the primary source of difficulty in cocktail-party recognition.
100. 【2608.08506】Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs
链接:https://arxiv.org/abs/2608.08506
作者:Mohanad Odema,Gabrielle De Micheli,Dayin Gou,Nilesh Malpeddi,Prathamesh Vaste,Jacob Song
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Performance (cs.PF)
关键词:prominence for LLM, reducing model parameter, model parameter count, maintaining task-level accuracy, Training-free low-rank compression
备注: COLM 2026
点击查看摘要
Abstract:Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. However, existing SOTA frameworks share two key limitations: (1) residual errors in calibration data activations accumulate across layers during compression, causing misalignment between representations simulated at compression time and those experienced at inference; (2) the assumption that layer importance distribution is preserved post-compression does not hold. Together, these two effects introduce misalignment in the compression process in relation to the deployed model. We study these effects and propose a simple, training-free methodology compatible with existing frameworks to mitigate them, comprising: (1) Layer-by-Layer Compression with Calibration Correction; (2) Iterative Compression with Rank Allocation Correction. Implemented atop an existing SOTA decomposition framework, and evaluated on Llama and Qwen3 models across various benchmarks and compression rates, our approach demonstrates up to ~1-2.5 accuracy point improvements over per-weight and joint decomposition baselines on zero-shot tasks.
101. 【2608.08503】MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
链接:https://arxiv.org/abs/2608.08503
作者:Rahma Simin Ali,Jawad Hossain
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:reasoning remains challenging, Mathematical reasoning remains, Bangla mathematical reasoning, remains challenging, challenging in low-resource
备注:
点击查看摘要
Abstract:Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical reasoning dataset with GPT-5.4-generated rationales, and fine-tune four 4B--7B student models under a matched protocol in which answer-only and CoT conditions share data splits, response-only loss masking, decoding, and scoring, differing only in the training target. In-domain, CoT provides no significant improvement over answer-only fine-tuning for three stronger backbones (paired bootstrap 95\% CIs include zero; exact McNemar $p \geq 0.17$), despite generating 15--52$\times$ more tokens, but significantly improves the weaker 4B model by 18.56 points ($p 0.0001$). On the larger, contamination-audited BanglaMATH benchmark, this pattern reverses: CoT significantly outperforms answer-only supervision for all four models by 20.1--28.1 points (all $p 0.0001$). Answer-only fine-tuning also reduces out-of-domain accuracy below the base model for three models, whereas CoT preserves or improves it for all four. A human study with two co-author annotators, external-expert adjudication, and Cohen's $\kappa = 0.76$--$1.00$ finds no significant CoT improvement over the base model on reasoning-content criteria; instead, its measurable effect is target-language adherence and producing inspectable reasoning. Overall, rationale supervision's value depends on backbone capability and distribution shift: in this setting, its main benefits are Bangla adherence, auditable reasoning, and out-of-domain robustness rather than improved in-domain reasoning validity.
102. 【2608.08485】HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails
链接:https://arxiv.org/abs/2608.08485
作者:Tak Ho Alex Li,Kaijie Liu,Lik-Hang Lee,Kin Chung Ho,Ping Shum,Michael K. Ng
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Current LLM safety, Current LLM, fine-tuning distorts pre-trained, LLM safety guardrails, generative judges incur
备注: Preprint, August 2026. 10 tables, 2 figures
点击查看摘要
Abstract:Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur prohibitive inference costs. We challenge the prevailing paradigm by asking: can safety be achieved through pure geometric reasoning over frozen semantic representations? We present HoloAegis, a minimally parametric topological inference framework that decouples representation from reasoning. We term our approach minimally parametric because the only free parameters are the anchor count K and the temperature tau, both fixed after construction and requiring no gradient-based training. An un-fine-tuned encoder maps text to a unit sphere, after which all decisions are purely geometric. We formalize safety evaluation as a Gibbs-Boltzmann Free Energy computation over a pre-computed System Topology Anchor Bank, and we introduce Dual Time-Scale Exponential Moving Averages to detect progressive multi-turn semantic drift. Our key theoretical insight is a Topological Boundary Stability Conjecture: we provide theoretical motivation and strong empirical evidence that sparse anchor centroids stabilize the decision boundary against high-frequency lexical perturbations far better than full vector space methods. Evaluated across 8 benchmarks, HoloAegis achieves state-of-the-art accuracy (1.0000 AUC on AuthenHallu, 0.9802 on HarmBench) with sub-millisecond latency, zero cold-start data, and cross-lingual transfer (0.9758 AUC on Chinese CHIFRAUD).
103. 【2608.08477】VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use
链接:https://arxiv.org/abs/2608.08477
作者:Juan S. Santillana
类目:Computation and Language (cs.CL)
关键词:LATAM cybersecurity imagery, LATAM security decoder, LATAM cybersecurity, LATAM security, Model Context Protocol
备注: 11 pages, 1 figure
点击查看摘要
Abstract:We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to a 1.04B Spanish/LATAM security decoder via an MLP. To our knowledge, it is the first sub-2B VLM specialized for cyber UI (IDA, Ghidra, Wireshark, Nmap, Metasploit, Volatility) that answers in Spanish, emits structured reasoning via native |think| tokens, invokes tools via Model Context Protocol (|tool_call|), and exports to this http URL's LLaVA mmproj format for air-gapped deployment. We report a negative preliminary visual-grounding result: despite fully functional pipelines, the current vision SFT (400-1900 steps, ~16M tokens) yields near-zero B6 scores (0.08 tool-identification), ignoring image content. We specify remediation (longer SFT, =60% replay, lower LR) and expose a checkpoint-loader bug (unstripped llm. prefix) masquerading as training collapse. Crucially, we introduce a 3-variant ablation matrix (V0: NoPE-every-4, V1: all-RoPE, V2: NoPE+learned 2D) to study if periodic no-positional-encoding (NoPE) layers help or hurt attention over the 729-token visual block. Code, configs, and weights are released to establish priority on this architectural question. We provide B1-B5 for the text backbone, text controls, preliminary B6/B7 scores, wall times, GGUF efficiency on CPU, and a corpus of 14,596 QA pairs across 10 domains. We open-source all models and trajectories: jsantillana/vectrayx-1b, jsantillana/vectrayx-vision-1b, and jsantillana/vectrayx-vision-1b-checks.
104. 【2608.08467】LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
链接:https://arxiv.org/abs/2608.08467
作者:Minhan Cho,Soyoung Park,Kihyeon Jeong,Byeongkyu Jeon,Daejin Choi,Jinyoung Han
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Model Context Protocol, Large Language Models, Context Protocol, Large Language, Model Context
备注: 4 pages, 1 table. Accepted at the AgentSearch Workshop at SIGIR 2026, Melbourne, Australia (non-archival). Code and data: [this https URL](https://github.com/rabqatab/llm-in-mcp-matters)
点击查看摘要
Abstract:The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server instructions: the system-prompt text a server hands to the host application. When a query concerns an entry of the embedded table, the model can act on it immediately instead of re-discovering the same information through a search tool. We test whether client LLMs actually consume such instruction-embedded data, reporting a 54,000-trial study across 24 LLMs (9 Claude, 6 Gemini, 9 GPT) on a production legal-information MCP server. A diagnostic condition that removes the competing search tool shows that failures are dominated by behavioral preference rather than missing capability. With search unavailable, 23 of 24 models read the embedded data reliably (hit ratio at least 98%); with a search tool merely present, 9 models drop below 15%. A 2^3 factorial analysis of three instruction-level interventions reveals strong interaction effects: combining all three restores at least 86% for 20 of 24 models, but individual interventions can backfire for specific model families. Per-server prompt engineering is therefore a workaround rather than a fix; we argue that MCP host applications should provide an explicit mechanism that places server instructions ahead of tool selection in the client LLM's deliberation.
105. 【2608.08459】Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
链接:https://arxiv.org/abs/2608.08459
作者:Zhuowen Liang,Zhengxuan Zhang,Jiayang Wang,Jiazhuo Chen,Nan Tang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Databases (cs.DB)
关键词:queryable relational databases, turn long, isolated spreadsheets, queryable relational, Practical AI systems
备注: 24 pages, 13 figures, 7 tables
点击查看摘要
Abstract:Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domains such as finance, healthcare, education, transportation, and enterprise operations, downstream workflows rely on normalized schemas, entity identities, keys, cross-table relationships, and integrity constraints for analytics, compliance, auditing, and SQL-backed decision making. Existing Document-to-Table benchmarks are insufficient for this setting: flattening evidence into single tables can duplicate entities, obscure many-to-many relationships, create sparse records, and avoid testing whether extracted facts form a valid database instance. This creates an urgent need to evaluate document understanding as database construction rather than field extraction. We introduce Doc2DB-Bench, a benchmark for Document-to-Database construction, containing 203 long-document instances across 42 schemas and seven domain groups, with 117 entity tables, 132 relationship tables, 7,341 rows, and 41,935 cells. Built through a controllable DB-to-Doc synthesis pipeline and organized by a taxonomy of intra-table extraction and inter-table reasoning, the generated documents undergo authenticity verification, proving indistinguishable from real-world references. Doc2DB-Bench thus provides a testbed for reliable, auditable, and relationally faithful LLM-based data systems. The benchmark is publicly available at this https URL.
106. 【2608.08451】Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization
链接:https://arxiv.org/abs/2608.08451
作者:Haojie Yu,Ziyou Jiang,Junjie Wang,Mingyang Li,Yuekai Huang,Jie Huang,Qing Wang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Ordered Reasoning Chain, changing lexical expressions, Reasoning Chain, share invariant principles, frequently changing lexical
备注: 9 pages, 4 figures, conference
点击查看摘要
Abstract:Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an Ordered Reasoning Chain (ORC) of recurring topics, harm language indicators, severity hierarchies, and type characteristics, which can help us capture the key information in the frequently changing lexical expressions. We propose BRACE, which encodes the ORC as four differentiable stages (Topic - Indicator - Severity - Type) with intermediate supervision, serving as a structured regularizer blended with direct heads, and supported by prototype-based feature augmentation and feature path disentanglement. The evaluation results show that, across 4 domains and 5 harm categories, BRACE achieves harm-type macro F1 of 0.934 (RoBERTa-wwm-ext, 3-seed mean), with decoder backbones (Qwen3-1.7B LoRA) reaching 0.949. Ablation studies show that all components contribute to BRACE, and the structural decomposition of ORC enables BRACE to distinguish harmful types with semantic ambiguity. Disclaimer: This paper may contain content that is disturbing to some readers.
107. 【2608.08447】Hidden Language Consistency Phenomena in Reasoning LLMs
链接:https://arxiv.org/abs/2608.08447
作者:Muhammad Ali Shafique,Kelly Marchisio
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:correct answer, commonly evaluated, preserve the intended, consistency, Multilingual
备注:
点击查看摘要
Abstract:Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding. This omission conceals important multilingual behaviors that emerge as tasks become harder. In this paper, we study task difficulty, task accuracy, thinking-language consistency (TC), and answer-language consistency (AC) across reasoning models using PolyMath benchmark in eight languages and four difficulty levels. We uncover four findings: (1) language consistency exhibits four difficulty-dependent behaviors: output-language consistency remains aligned with input, remains misaligned, degrades gradually, or collapses abruptly. (2) We identify the language consistency breakdown effect, where increasing difficulty can cause a sudden drop in output-language consistency, especially in less strongly represented and non-Latin-script languages. (3) Due to this breakdown effect, accuracy can be preserved or even improved at a harder difficulty level as the model shifts to its internal dominant language. (4) Quantization can improve or degrade output-language consistency independently of its effect on accuracy, with GPTQ and AWQ often outperforming AutoRound under tolerance-based voting with {\epsilon} = 1.0. These results show that multilingual capability cannot be characterized by accuracy alone; reliable evaluation should jointly consider task accuracy, language consistency, and task difficulty for multilingual benchmarks.
108. 【2608.08392】CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
链接:https://arxiv.org/abs/2608.08392
作者:Zejun Xu,Taiyi Chen,Jin Li,Yongtong Gu,Qi Cheng,Aixuan Lv,Shuai Zhu,Pengfei Zhu,Kaichen Yang,Boyu Sun,Yixian Yang,Mulong Xie,Xin Liu,Dagang Li,Xiaoteng Ma,Hongru Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large language models, Large language, language models, models are increasingly, increasingly deployed
备注: Accepted to COLM 2026. Project page: [this https URL](https://warriorxu0302.github.io/CAP-Bench/)
点击查看摘要
Abstract:Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven by benchmarks that evaluate end-to-end task success, these evaluations largely overlook two fundamental sources of difficulty in real web browsing: complex actions over rich user interfaces and visual perception of dynamically rendered content, especially in workflows that span multiple websites. We introduce CAP, a scalable benchmark for evaluating browser agents on cross-site, human-like web tasks that require non-trivial UI interactions and visual understanding. Specifically, we adopt a decomposition-and-recomposition pipeline that first abstracts each website into a structured site card capturing user-facing functions, complex execution operations, and perceptual requirements, and then recomposes these components into realistic cross-site workflows. Each task is therefore grounded in multiple specific operations on each website, enabling fine-grained diagnosis. Built on this framework, we construct 420 tasks across 108 real-world websites and 24 domains under careful quality control. Experiments on state-of-the-art browser agents using our verifiable agent-as-a-judge evaluation framework show low success rates and reveal that perception-heavy interactions remain a major bottleneck, exposing substantial gaps between current agents and real-world web browsing demands.
109. 【2608.08383】Safety Cost of Steering Vectors Is Separable and Reducible
链接:https://arxiv.org/abs/2608.08383
作者:Yuxiao Li,Gjergji Kasneci
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:controlling LLM behavior, controlling LLM, LLM behavior, Steering, lightweight tool
备注: COLM 2026
点击查看摘要
Abstract:Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally compromise a model's safety mechanisms and increase compliance with harmful requests, while no effective mitigation yet exists. In this work, we show that this safety degradation arises from a separable component in the vector that disrupts the model's safety mechanisms but contributes little to the steering objective. We identify and remove this safety-degrading component, formulating the task as a constrained optimization problem solved through primal-dual updates, subject to preserving the intended steering effect and bounding false refusal. The resulting solution is both interpretable and surgical: the optimization recovers a single direction whose ablation from the steering vector restores model safety with minimal utility cost. Across models, steering behaviors, and attack suites, including unseen attacks types, our method substantially reduces steering-induced safety degradation while preserving the original steering effect with minimal impact on false refusal. Our method offers a post-hoc correction to steering vectors that mitigates their safety cost, and more broadly, it provides a general recipe for applying activation-level model interventions without paying a safety tax.
110. 【2608.08300】Mitigating Over-Personalization in LLMs via Structured Memory
链接:https://arxiv.org/abs/2608.08300
作者:Hakeem Hannoon,Andrew Zhao,Mihir Narayan,Sharvin Goyal,Ivaxi Sheth
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Conversational assistants increasingly, assistants increasingly rely, Conversational assistants, persistent long-term memory, assistants increasingly
备注:
点击查看摘要
Abstract:Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored user information is reintroduced into the model context, it can also influence responses in inappropriate or unrelated settings. We study two such failure modes in memory-augmented LLMs: cross-domain leakage, where memories from one life domain affect responses in another, and memory-induced sycophancy, where stored user beliefs make models more likely to agree with the user rather than respond truthfully. We apply a simple inference-time modification to how memories are presented to the model, without changing the model or the memory contents. Across seven models on PersistBench, we compare the commonly used all-in context format, where memories are injected as an unstructured list, with structured formats that partition memories by domain. This simple modification consistently reduces cross-domain leakage while preserving utility, with our strongest method reducing leakage by $8.8\%$ on average relative to the baseline.
111. 【2608.08283】Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
链接:https://arxiv.org/abs/2608.08283
作者:Osvaldo Quinjica,Eric Bennett,Xinchen Yang,Andrew Schonebaum,Marine Carpuat
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:digital humanities workflows, large language models, historical languages surprisingly, models can translate, translate some historical
备注:
点击查看摘要
Abstract:Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation. We investigate whether existing automatic evaluation metrics developed for modern languages are reliable in this setting, using translation from Classical Chinese to English as a test case. We introduce a diagnostic framework based on minimal pairs capturing error types salient in scholarly use, probing both reference-based and reference-free metrics for error sensitivity and tolerance to valid variation. We find that all metrics exhibit blind spots, however MetricX24 performs best overall. Our findings highlight the need for more robust and interpretable metrics for historically and culturally distinct translation settings.
112. 【2608.08256】AraSSM: A bidirectional state-space encoder for Arabic masked language modeling
链接:https://arxiv.org/abs/2608.08256
作者:Ahmed Amine Aliane,Hassina Aliane,Nasredine Semmar
类目:Computation and Language (cs.CL)
关键词:Mamba encoder pretrained, self-attention mechanism scales, mechanism scales quadratically, bidirectional Mamba encoder, Pretrained Transformer encoders
备注:
点击查看摘要
Abstract:Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but their self-attention mechanism scales quadratically with sequence length, which limits efficiency on long documents. Mamba, a selective state-space model (SSM), offers linear-time sequence modeling as a competitive alternative to attention, yet no dedicated bidirectional Mamba encoder pretrained specifically for Arabic currently exists. We introduce AraSSM, a bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days. We evaluate AraSSM by fine-tuning on four established Arabic NLU benchmarks covering sentiment classification (HARD), named entity recognition (ANERcorp), extractive question answering (ARCD), and natural language inference (XNLI-ar), following the per-task evaluation protocol introduced by AraBERT, and report results as mean +/- standard deviation across three fine-tuning seeds. AraSSM matches or exceeds published base-sized Transformer baselines on sentiment classification (96.37 +/- 0.03% accuracy on HARD), is competitive on extractive QA (32.19 +/- 1.07 EM, 63.79 +/- 0.25 F1 on ARCD) and named entity recognition (81.54 +/- 0.30 entity-level F1 on ANERcorp), and trails the base-sized Transformer range on natural language inference (72.83 +/- 0.07% accuracy on XNLI-ar), despite being trained entirely from scratch on consumer hardware rather than large-scale accelerator clusters.
113. 【2608.08255】Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
链接:https://arxiv.org/abs/2608.08255
作者:Yifu Huo,Shunjie Xing,Chenglong Wang,Peinan Feng,Qiaozhi He,Yan Ding,Anxiang Ma,Yuxin Gao,Tongran Liu,Tong Xiao,Jingbo Zhu
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Agentic reinforcement learning, credit assignment, reinforcement learning, suffers from delayed, delayed and sparse
备注:
点击查看摘要
Abstract:Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge is credit assignment, which aims to decompose trajectory-level rewards and provide more fine-grained supervision for intermediate decisions. However, existing credit assignment approaches ignore the rich process information naturally generated during environment interaction, e.g., interaction history. We argue that such information provides valuable supervision for identifying the contribution of individual actions. To this end, we propose Environmental Feedback-based Credit Assignment (EFCA), a multi-timescale credit assignment approach for long-horizon agentic RL. EFCA complements the long-term outcome signal with two environment-grounded process signals: a short-term feedback signal that captures the immediate effect of the current action and a medium-term state-history signal that identifies ineffective patterns from recent interactions. Both signals are directly extracted from environment feedback and integrated through a return reweighting mechanism. Experiments on ALFWorld and WebShop demonstrate that EFCA consistently improves both task success and task quality over strong baselines, highlighting the effectiveness of environment-grounded multi-timescale credit assignment for long-horizon agentic RL.
114. 【2608.08239】he Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World
链接:https://arxiv.org/abs/2608.08239
作者:Ashritha Gonuguntla
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:LLM routers promise, routers promise efficiency, inside multi-step agents, cheapest adequate model, LLM routers
备注: 8 pages, 3 figures. Accepted at the Conference on Language Modeling 2026. Code: [this https URL](https://github.com/AshrithaG/replay-gap) Data: [this https URL](https://huggingface.co/datasets/ashritha0907/replay-gap-trajectories)
点击查看摘要
Abstract:LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents. Yet agentic routers are evaluated like single-turn routers: by replaying logged trajectories and substituting another model's recorded outputs, assuming the rest of the trajectory is unaffected. We test this assumption with branching rollouts: we fork live SWE-bench agent trajectories at controlled points, rebuild the environment, continue each fork with a different model, and compare against same-model control forks that isolate sampling and replay noise. Across six paired runs (~900 rollouts), swaps exceed their matched control floors by +0.25 to +0.66 normalized edit distance (multiplicity-corrected CIs exclude zero), rewriting 61-94% of post-fork actions; 74-77% of early swaps diverge at the first post-fork action, versus 6-35% of controls, leaving only 3% of replayed states valid. Divergence decreases with fork depth in both directions. All five outcome flips we observe occur in swap arms, upgrades rescuing unsolved instances and a downgrade losing the sole solve, and zero occur across 359 control forks. Scoring these same swaps with a log-stitching replay evaluator, replay mispredicts every success-relevant outcome call and predicts patches with 0.00-0.11 similarity to reality. Auditing the noise floor, temperature-0 "determinism" is configuration-dependent: FP8-served controls diverge on over 90% of forks while AWQ-served ones remain near-identical; and under tight budgets the stronger model more often exhausts its steps without submitting. Replay-based benchmarks score the wrong world for agentic routing; we release our harness and all trajectories.
115. 【2608.08237】SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems
链接:https://arxiv.org/abs/2608.08237
作者:Muhammad Faizan Raza,Shuo(Luna)Yang,Satish Mahadevan Srinivasan
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC); Information Retrieval (cs.IR)
关键词:service level objectives, strict service level, Retrieval-Augmented Generation, systems in production, level objectives
备注: 7 pages, 5 figures, 2 tables. Authors' accepted version of a paper published in Proc. IEEE CoDIT 2026. The version of record is available at the DOI below
点击查看摘要
Abstract:Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval pipelines rely on fixed retrieval budgets that ignore query difficulty, over-retrieving for easy queries and under-serving hard ones, forcing operators to trade answer quality against SLO compliance. This paper proposes SAGE, a learned SLO-aware adaptive retrieval policy that dynamically selects the number of passages k per query. SAGE uses lightweight features derived from initial retrieval (e.g., score distributions, rank gaps, lexical signals) and is trained offline via imitation learning from an oracle that approximates optimal latency-quality trade-offs. At inference, it adds no LLM calls and minimal overhead. On Natural Questions, under a 5s P95 latency SLO, SAGE achieves 95% SLO compliance versus 30% for the best static baseline (k=20), reduces P95 latency by 36% and retrieval cost by 51% with only 2 percentage points Exact Match (EM) loss. A single policy trained on Natural Questions generalizes across HotpotQA, UnSeenTimeQA, and four LLM families (Llama, Qwen, Mistral, Gemma), consistently yielding +45-52 point SLO improvements without quality degradation.
116. 【2608.08236】LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems
链接:https://arxiv.org/abs/2608.08236
作者:Heng Zhou,Lian Zhang,Yutao Fan,Tiancheng He,Siki Chen,Hejia Geng,Philip Torr,Zhenfei Yin
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Multi-agent LLM systems, Multi-agent LLM, candidate answers, systems often fail, lack of candidate
备注:
点击查看摘要
Abstract:Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which incompatible claim should currently be trusted. Majority vote, debate, and judge-based selection choose an output without recording which claim wins, which is contested, or why a later update supersedes it. We present \term{LatticeMind}, a conflict-aware structured memory that handles contradiction at write time. It maintains explicit item status, applies cheap symbolic conflict checks, and invokes LLM reconciliation only for unresolved semantic cases. On a label-blind ConflictBank evaluation that removes source-name hints, LatticeMind reaches 0.97 accuracy versus 0.61 for the strongest aggregation baseline, with the gap significant at $p10^{-6}$ by paired McNemar test. Ablations show that removing the checker or the reconciler costs 12 to 14 points. On four secondary planning benchmarks the picture is mixed: LatticeMind beats naive merge on three of four, but does not replace deliberation methods on tasks rewarding iterative search.
117. 【2608.08227】Focus particles and scalar inferences across humans and language models
链接:https://arxiv.org/abs/2608.08227
作者:Catherine M. Brousse,Nelu D. Radpour
类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:formal semantic theories, posit structured representations, Focus particles, central to formal, formal semantic
备注: 3 pages, 1 figure, presented at 9th annual Conference on Cognitive Computational Neuroscience
点击查看摘要
Abstract:Focus particles such as "even" and "only" are central to formal semantic theories that posit structured representations over sets of alternatives. "Even" highlights unexpected or extreme alternatives, while "only" enforces exclusivity. If such scalar representations are robust and generalizable, they should give rise to consistent judgments across contexts and systems. In this work, we test whether humans and large language models (LLMs) construct stable scalar representations from sentences containing these particles. Using a dataset of approximately 100 items, participants and models were asked to make scalar judgments. Preliminary results suggest that similar outputs across humans and LLMs may arise from different underlying mechanisms.
118. 【2608.08212】Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment
链接:https://arxiv.org/abs/2608.08212
作者:Peiyang Liu,Xi Wang,Ziqiang Cui,Di Liang,Wei Ye
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:induce emergent misalignment, In-context learning, emergent misalignment, induce emergent, narrow misaligned
备注:
点击查看摘要
Abstract:In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts, however, conflate harmful-text exposure with an invitation to continue assistant behavior. We hold harmful answers fixed while varying their delivery as demonstrations, evidence, assistant history, or tool output. Across ten independently sampled contexts, demonstration framing raises broad EM by $30$--$32$ percentage points on a susceptible Gemini model; the gap survives domain exclusion, semantic clustering, unseen questions, and four prompt templates. Format and length-matched controls show that harmful content is necessary but insufficient. A role times continuation factorial further reveals model-dependent provenance effects: Gemini follows both assistant and tool histories, whereas Grok largely resists tool-framed continuation. Several other frontier and open-weight models show no gap. Blinded human audits confirm every main contrast and show that the model judge underestimates active-condition failures. Thus continuation framing is a strong, model-dependent moderator of ICL-EM, not a universal consequence of harmful context.
119. 【2608.08188】Quantization Degradation in Large Language Models: A Signal-Noise Perspective
链接:https://arxiv.org/abs/2608.08188
作者:Chenxi Zhou,Pengfei Cao,Jinyu Ye,Bohan Yu,Haida Yu,Jiang Li,Jun Zhao,Kang Liu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Post-training quantization reduces, large language models, quantized model degrades, Post-training quantization, reduces the deployment
备注:
点击查看摘要
Abstract:Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone. We systematically study weight-only post-training quantization across bit-widths, quantization methods, model scales and downstream tasks on multiple model families. We observe that such degradation varies substantially across these factors: 4-bit quantization usually preserves performance, 2-bit often causes broad degradation, and at 3-bit, degradation becomes apparent but varies markedly with task type, quantization method and model scale. To explain this variability, we use the signal-to-noise ratio (SNR) to measure how strongly quantization perturbs full-precision representations. We trace degradation back to two linked processes: how quantization errors arise within individual modules, and how they accumulate across layers. First, a source SNR decomposition shows that newly introduced errors depend on three factors: the magnitude of the weight error, the strength of the task-specific signal, and how strongly the quantization error aligns with task-specific activations. Different factors affect these components in distinct ways. Second, a cross-layer propagation analysis shows that these errors can be attenuated, preserved, or amplified as they pass across layers, and that larger models benefit from weaker error amplification. Together, these results establish that quantization degradation is governed by how errors are introduced at the source and how they accumulate across the network.
120. 【2608.08180】A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization
链接:https://arxiv.org/abs/2608.08180
作者:Praveen Kumar Katwe,Rakesh Chandra Balabantaray,Kali Prasad Vittala,Naman Kabadi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:systems frequently generate, frequently generate fluent, Toggle, Abstractive text summarization, text summarization systems
备注: 6 pages, 4 figures, 6 tables
点击查看摘要
Abstract:Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships between entities and events. Such relation-level hallucinations undermine the reliability of generated summaries, particularly in high-stakes domains. In this work, we present a refined and grounded framework for evaluating relation hallucination in abstractive summarization. We present the empirical Relation Hallucination Index (RHI) by introducing a dependency-aware relation extraction algorithm that incorporates lemmatization-based normalization, named entity grounded subject resolution, passive agent recovery, negation-aware verb modeling, reporting verb filtering, nominal relation fallback, clausal propagation, and systematic deduplication. These enhancements improve the structural fidelity of extracted relation triples and reduce spurious matches during evaluation. In addition, we introduce a normalized formulation of RHI to ensure scale-invariant comparison between datasets and models. The revised metric decomposes hallucination into interpretable components, aggregates relation hallucination metric into a normalized relation faithfulness score. Extensive evaluation across multiple state-of-the-art summarization models demonstrates that the grounded extraction process yields more stable and discriminative hallucination measurements. The proposed framework advances automated relation-level faithfulness evaluation and supports coherence-aware, hallucination-sensitive model analysis.
Comments:
6 pages, 4 figures, 6 tables
Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
ACMclasses:
I.2.7; H.3.1
Cite as:
arXiv:2608.08180 [cs.CL]
(or
arXiv:2608.08180v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.08180
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Praveenkumar Katwe [view email] [v1]
Sat, 8 Aug 2026 15:16:51 UTC (353 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization, by Praveen Kumar Katwe and 3 other authorsView PDFHTML (experimental)TeX Source
view license
Current browse context:
cs.CL
prev
|
next
new
|
recent
| 2026-08
Change to browse by:
cs
cs.AI
References Citations
NASA ADSGoogle Scholar
Semantic Scholar
export BibTeX citation
Loading…
BibTeX formatted citation
loading…
Data provided by:
Bookmark
checked="checked"class=“labs-tab-input”>
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.
Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)
mathjaxToggle();
We gratefully acknowledge support from
our major funders,
member institutions, ,
and all contributors.
About
Help
Contact
Subscribe
Copyright
Privacy
Accessibility
Operational Status (opens in new tab)
Major funding support from
121. 【2608.08168】hinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders
链接:https://arxiv.org/abs/2608.08168
作者:Bo Cheng,Qiaolin Lu,Yi Chang,Yuan Wu
类目:Computation and Language (cs.CL)
关键词:Large Language Models, remain poorly understood, direct answer generation, Large Language, neural mechanisms distinguishing
备注:
点击查看摘要
Abstract:While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain poorly understood. To deconstruct this cognitive process, we apply Top-K Sparse Autoencoders (SAEs) to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examine the model's divergent behaviors across math-solving tasks of three distinct difficulty levels. Observationally, we identify a clear distinction in how the model functions under two reasoning modes: Thinking mode relies on sparse and high-intensity feature activations driving verbal deduction independent of problem complexity, whereas NoThinking mode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation. Causally, suppressing the three most active sparse features by Total Activation Volume reveals three principles: (i) reasoning and syntactic structure are tightly coupled, as interventions consistently degrade \LaTeX{} and boxed-solution formatting; (ii) Thinking responds to disruption with compensatory over-generation marked by increased metacognitive cues and repetitive, low-information continuations; and (iii) coherent CoT behavior depends on a fragile coordination among specialized features, yielding distinct failure modes under perturbation but a consistently impaired output structure.
122. 【2608.08164】STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs
链接:https://arxiv.org/abs/2608.08164
作者:Nuthakki Siva Gopala Krishna,Kanishka Jain
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:reducing computational costs, widely adopted technique, significantly reducing computational, large language models, large teacher model
备注: 15 pages
点击查看摘要
Abstract:Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.
123. 【2608.08160】Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
链接:https://arxiv.org/abs/2608.08160
作者:Yingpeng Ma,Jianhao Yan,Bei Shi,Ka Hou Kam,Runnan Wang,Xuebo Liu,Yulong Chen,Yue Zhang,Derek F. Wong
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Large Language, fluid interactive storytelling, advancement of Large, Games by enabling
备注: Accepted by ICML 2026
点击查看摘要
Abstract:The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this, we formulate this challenge as Narrative Commitment Preservation (NCP), and take interactive narrative as our testbed. We introduce NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses. Each environment includes a structured narrative specification (trajectory, commitments, and initial facts) that we can automatically check throughout the interaction between the player agent and the narrator agent. Experiments across state-of-the-art LLMs reveal a substantial long-horizon consistency gap: high linguistic quality does not guarantee commitment preservation; even strong models frequently generate logically conflicting content under adversarial interventions, with the best-performing model (GPT-5.2) achieving only 42% survival rate after 20 turns and fact conflict rates ranging from 40% to 68% across models, and only isolated runs satisfying all achievement commitments within the 100-turn limit.
124. 【2608.08143】DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate
链接:https://arxiv.org/abs/2608.08143
作者:Anthony Miyaguchi,Conor Johnston
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:ARC working-note submission, Retrieval-Augmented Debate task, ARC working-note, ARC submission consisted, evaluating debate responses
备注: 12 pages, 4 figures. Accepted for publication in the CLEF 2026 Best of Labs proceedings
点击查看摘要
Abstract:We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval-augmented prompting pipeline. We summarize the results from the working paper and explore whether multi-LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at this https URL and this https URL.
125. 【2608.08126】Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk
链接:https://arxiv.org/abs/2608.08126
作者:Gregorius Reynaldi Pratama,Kuo-Kun Tseng
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:scoring increasingly relies, Credit scoring increasingly, decision logic, adverse decisions, scoring increasingly
备注: 13 pages, 9 figures, 5 tables
点击查看摘要
Abstract:Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable. A common proposal closes that gap with a language model: compute feature attributions, hand them to an LLM, and let it write the rationale. We build such a system end to end and test whether the second half of the promise holds. The predictive component is a multi-scale stacking ensemble fusing four differently regularised gradient-boosting learners with a residual network through a neural meta-learner trained on out-of-fold predictions. On a public 32,581-application credit dataset it reaches test ROC-AUC 0.9539 (95% CI [0.9462, 0.9616]) and PR-AUC 0.9137, beating the best single model by Delta-AUC = 0.0143 (p = 0.016 under a conservative independence assumption). Our central finding is asymmetric. The ranking gain is real but operationally small: at the F1-optimal threshold the ensemble avoids only six additional missed defaults out of 1,422 against a tuned random forest, cutting cost-weighted loss by under 2%. The narrative layer fails in a way prompt engineering alone does not fix. In an audited case the model named three factors as risk-increasing that the supplied attributions scored as risk-reducing, omitted the dominant driver, and introduced a feature never given to it. We trace this to properties we measure rather than assume: SHAP and LIME agree on which features matter (overlap@10 = 0.80) but not on their order (tau = 0.43, p = 0.18), and the attribution sign for the model's most sensitive input is near a coin flip across applicants (modal-sign share 0.53). Calibration (ECS = 0.117) and perturbation stability (DPD = 0.078) both fall short of our own thresholds. Constrained prompting is necessary but not sufficient: grounding must be verified after generation, not assumed.
126. 【2608.08107】NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
链接:https://arxiv.org/abs/2608.08107
作者:Jiayue Jin,Jingwei Zhang,Chen Wang,Jing Liu,Longteng Guo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large language models, language intelligence acquired, Neuron-aware Plasticity Allocation, enables new perceptual, acquired during pretraining
备注:
点击查看摘要
Abstract:Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence acquired during pretraining. In this work, we investigate this phenomenon from the perspective of internal adaptation dynamics and discover that neurons in pretrained LLMs exhibit heterogeneous plasticity during multimodal learning: some neurons are critical for preserving language capabilities, while others are more adaptive to multimodal knowledge. Based on this insight, we propose NeuPAT (Neuron-aware Plasticity Allocation Tuning), a lightweight and architecture-agnostic framework that allocates neuron-wise update constraints during multimodal instruction tuning. NeuPAT uses a small-scale probing stage to estimate neuron adaptation patterns and selectively protects language-sensitive neurons while promoting multimodal adaptation through more plastic neurons. Experiments across diverse LLM families demonstrate that NeuPAT recovers 94.5\% of the language capability degradation caused by vanilla tuning on 11 language benchmarks while maintaining comparable multimodal performance, providing an effective approach for capability-preserving multimodal expansion.
127. 【2608.08090】Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs
链接:https://arxiv.org/abs/2608.08090
作者:Rama Alomair,Remas Alsubaie,Walaa Saifalislam,Rima Alsonbul,Mona Alnajjar,Razan Aldossari,Haya Alibrahim,Abeer Aldayel
类目:Computation and Language (cs.CL)
关键词:Culture Specific, figurative language identification, requires a clearer, clearer understanding, figurative
备注: This paper is under review
点击查看摘要
Abstract:Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clearer understanding of the contribution of translated multilingual supervision. We examine this question using 742 proverb concepts across 6,787 translated instances in seven languages. We evaluate five models, including multilingual encoders and instruction tuned LLMs, under progressively increasing levels of multilingual supervision. Moreover, we introduce a multidimensional annotation framework for proverbs that characterizes them through four complementary figurative forms: Metaphorical, Moral/Advisory, Cause-Effect, and Culture Specific. Our findings show that approximately 50% of the translated multilingual training data is sufficient to achieve near-optimal figurative language identification performance. We further show that combining diverse figurative forms yields the strongest overall performance. A notable finding is that the least frequent figurative form, Culture Specific, exhibits the largest performance gains under multilingual supervision. Furthermore, the Moral/Advisory and Culture Specific forms contribute most to the performance of instruction-tuned LLMs on figurative language identification. These findings motivate multilingual figurative language identification to move beyond metaphor-centric taxonomies toward concept level multidimensional frameworks that explicitly model complementary forms of figurative meaning.
Comments:
This paper is under review
Subjects:
Computation and Language (cs.CL)
Cite as:
arXiv:2608.08090 [cs.CL]
(or
arXiv:2608.08090v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.08090
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
128. 【2608.08086】Archer: Adaptive Reuse of Cached Hidden States for Efficient Rollback in Diffusion Language Models
链接:https://arxiv.org/abs/2608.08086
作者:Xuning He,Zinan Sheng,Yongding Tao,Huanyu Liu,Ge Li,Xue Jiang,Yihong Dong
类目:Computation and Language (cs.CL)
关键词:Diffusion language models, allowing earlier predictions, Diffusion language, language models, iteratively refine
备注:
点击查看摘要
Abstract:Diffusion language models (DLMs) iteratively refine a sequence, allowing earlier predictions to be revised as context evolves. This rollback capability distinguishes them from irreversible autoregressive generation, but makes inference costly. Every denoising update alters the global context, forcing both prompt and response states to be recomputed even though only response tokens are revisable. Key-value (KV) caching could reduce this cost, yet conventional caching assumes immutable historical states and is therefore difficult to reconcile with rollback. In this paper, we introduce Adaptive Reuse of Cached Hidden States for Efficient Rollback (Archer), a training-free KV caching method for rollback-capable DLMs. Archer asymmetrically keeps the mutable response synchronized with the current hypothesis while reusing prompt K/V within a bounded state neighborhood. Although prompt representations also change under bidirectional attention, their token identities remain fixed; bounded reuse therefore amortizes repeated prompt computation without caching mutable response states. It also delays feedback from tentative tokens, reducing premature reinforcement of transient high-confidence errors and giving rollback more opportunity to correct them. Our analysis characterizes prompt reuse as a reversibility-aligned cache boundary, bounds its state-dependent approximation error, and gives a decoder-margin condition for preserving full-refresh decisions. Existing DLM acceleration often trades quality for speed. Archer shifts this frontier, attaining the best mean performance of 33.63% together with a 2.57x mean speedup on the main suite. Across evaluated settings, it improves Pass@1 by up to 3.05 points and reaches up to 2.95x speedup. Controlled analyses connect the quality gain to delayed prompt feedback and validate state-aware refresh. Our code is available at this https URL.
129. 【2608.08082】Commitment Before Realization: When Classifier-Free Guidance Becomes Unnecessary in Masked Diffusion Language Models
链接:https://arxiv.org/abs/2608.08082
作者:Fan Zhou,Weitian Wang,Tim Van de Cruys
类目:Computation and Language (cs.CL)
关键词:diffusion language model, Classifier-free guidance, masked diffusion language, diffusion language, CFG
备注:
点击查看摘要
Abstract:Classifier-free guidance (CFG) is usually kept on throughout masked diffusion language model decoding, although its benefit varies across prompts and over time. We study when CFG is actually needed by comparing, from any partial output, the probability of eventual constraint satisfaction under continued CFG and under base-only continuation. Their difference defines the remaining value of guidance. Guidance dependence is highly prompt-specific. Many prompts already succeed without CFG, while for others it provides no measurable benefit or can be harmful. For prompts that do benefit, the gain is often concentrated early. We define the commitment horizon $\astar$ as the earliest point from which switching all remaining decoding to the base model reduces final success by no more than a chosen tolerance. Under the base model, the corresponding success probability, or committor, is a martingale. To first order, CFG's per-step effect is governed by the covariance between the guidance logit direction and the successor committor. This gives a local account of when guidance can help, but it does not by itself locate the horizon. Among prompts with an observed preterminal horizon, $\astar$ is usually early and varies more within constraint families than between them. Freezing each prompt at its own cross-fitted horizon is noninferior to full CFG on all 13 subtasks at the prespecified margin, even while many tokens remain masked. This separates commitment from realization. The boundary also identifies a later region in which higher parallelism adds only a small cost in constraint success, although fluency still degrades with parallel width. For failed trajectories, reopening committed positions improves recovery in both failure modes.
130. 【2608.08067】DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
链接:https://arxiv.org/abs/2608.08067
作者:Yi Shu,Tianyu Peng,Yingzhuo Deng,Wen Yang,Jun Lin,Changming Xie,Xinyu Yu,Jiajun Zhang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:speech dialogue, speech dialogue models, speech, speech dialogue model, dialogue
备注:
点击查看摘要
Abstract:Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to the scarcity of dialect speech data. Moreover, during dialect adaptation, the semantic representation space of speech dialogue models continuously evolves, while conventional speech supervision remains unchanged, leading to semantic inconsistency between hidden representations and speech targets and degrading speech stability and naturalness. To address these issues, we propose DialectS2S, an end-to-end speech dialogue model for Chinese dialects. We first develop a scalable dialect speech dialogue synthesis pipeline for efficient data construction. We further introduce a two-stage post-training strategy with self-aligned speech supervision, which aligns the semantic content of speech supervision with the evolved semantic representations of the model to improve dialect speech generation quality. Experimental results show that DialectS2S consistently outperforms existing baselines across multiple Chinese dialects in speech dialogue, achieving substantial improvements in dialect consistency, response quality, and speech intelligibility. Our work provides an efficient and scalable solution for end-to-end speech dialogue modeling in low-resource dialect scenarios. To facilitate future research and practical applications, we fully open-source the DialectS2S framework, including model checkpoints, training datasets, and fine-tuning code.
131. 【2608.08059】APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain
链接:https://arxiv.org/abs/2608.08059
作者:Marie Escribe,Tharindu Ranasinghe,Amal Haddad Haddad,Hansi Hettiarachchi,Damith Premasiri
类目:Computation and Language (cs.CL)
关键词:highly repetitive documents, Machine Translation, output often involves, Automatic Post-Editing, involves repeating
备注:
点击查看摘要
Abstract:Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, especially in specialised and highly repetitive documents. Despite substantial work on Automatic Post-Editing (APE), most available corpora operate at the sentence level, others are synthetic, and overall not designed to study how corrections propagate in realistic Computer-Assisted Translation (CAT) workflows. This paper presents the APEX-VW (Automatic Post-Editing eXperiments on Virtual Wards) Corpus, a new open English-Spanish (EN-ES) dataset built from recent NHS virtual-ward documents and professional PE in Trados Studio, with controlled MT, terminology, and quality assurance settings. The corpus contains seven document-coherent source texts totalling 42k words, translated with four MT systems representing different paradigms and then post-edited by professional translators. Unlike prior resources such as WMT APE corpora, eSCAPE, MLQE-PE, or LangMark, the dataset preserves document order and CAT-tool context, making it suitable for research on terminology normalisation, correction propagation, and human-in-the-loop translation support. The paper describes the corpus design, data preparation, PE setup, and initial corpus statistics, and positions the resource as a benchmark for document-level APE and propagation-aware assistive tools.
132. 【2608.08032】Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE
链接:https://arxiv.org/abs/2608.08032
作者:Ramakrishna P. Kompella,Aadit Mahajan
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:request in English, harmful request, lower-resource language, reliably refuses, refuses a harmful
备注: Accepted to the actionable Interpretability workshop at COLM 2026
点击查看摘要
Abstract:Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reasoning model, and find it is not a failure to detect harm. Harm is encoded as an internal direction that is nearly language-invariant in mid-network (English-vs-Indic cosine ${\approx}0.9$ at $L11$), and steering that direction upstream causally controls refusal. But the detection direction is orthogonal to the change that actually writes the refusal, which is late and assembled over the course of generation rather than read off in a single forward pass. We attribute the write to a specific, localizable circuit, a mixture-of-experts writer held in check by an attention opposer and price every way of intervening on it: damping the opposer is cheap and effective, amplifying the writer is a cost wall, and surgical edits to the responsible heads do nothing. The circuit's organization, and the gradient method that exposes it, recur in a second, unrelated MoE model, while the lever's strength is architecture-specific. The result is a cost-measured map of where a multilingual safety repair can land, and what it costs
133. 【2608.08024】Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
链接:https://arxiv.org/abs/2608.08024
作者:Zakhar Mrykhin,Valentin Malykh
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large language models, Large language, Prompt Embedding Probes, PEP, generate fluent
备注: 10 pages, 7 figures. Code available at [this https URL](https://github.com/zazamrykh/internal_probing)
点击查看摘要
Abstract:Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Probes (PEP), a white-box method for answer-level hallucination detection from the hidden states of a frozen LLM. PEP extends standard linear probes by augmenting the input with a small number of learnable prompt embeddings. We evaluate PEP on TriviaQA, GSM8K, and MedQA using Qwen3 models at multiple scales. PEP improves hidden-state-based detection over standard linear probes in the main in-distribution setting. We further evaluate PEP for pre-generation prediction, cross-model transfer, and out-of-distribution generalization. PEP remains effective in the pre-generation and cross-model settings, whereas robust cross-dataset transfer remains difficult. These results show that prompt-based adaptation can strengthen hidden-state probing while keeping the backbone frozen and adding only a small number of trainable parameters.
134. 【2608.07968】hinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
链接:https://arxiv.org/abs/2608.07968
作者:Chenrui Fan,Yize Cheng,Ming Li,Yongyuan Liang,Tianyi Zhou,Soheil Feizi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:language models increasingly, existing evaluations typically, evaluations typically study, improve performance, increasingly use test-time
备注:
点击查看摘要
Abstract:Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute one question at a time. Yet when multiple problems share an end-to-end cost or latency constraint, models must decide how to divide limited inference compute among them. We introduce an exam-style evaluation framework for studying this setting, in which a model must distribute one shared token budget across questions with different difficulty and point values to maximize its total score. Across several open and frontier reasoning models, we find that models fail to allocate a shared budget strategically across questions of varying difficulties and values. Models behave largely as greedy sequential solvers: they prioritize questions by presentation order, front-load effort on early questions, and remain insensitive to value, with these tendencies becoming more pronounced as the number of questions grows. Explicit planning prompts spread compute more evenly but do not produce value- or difficulty-aware prioritization. The same behavioral pattern extends from mathematical to code reasoning. These findings establish global budget allocation as a distinct capability that is not captured by conventional per-question evaluation and remains a challenge for current reasoning models.
135. 【2608.07933】EvoTrustRAG: Evolution-Aware Conflict Attribution and Evidence Handling for Reliable Retrieval-Augmented Generation
链接:https://arxiv.org/abs/2608.07933
作者:Xi Nie,Hongwei Li,Shenghao Wu,Wenshu Fan,Qiyang Song,Wenbo Jiang
类目:Cryptography and Security (cs.CR); Computation and Language (cs.CL)
关键词:large language models, adversarial environments, factuality of large, large language, language models
备注:
点击查看摘要
Abstract:Retrieval-Augmented Generation (RAG) improves the factuality of large language models with external knowledge, yet conflicting evidence remains a fundamental challenge in dynamic and adversarial environments. Existing approaches often treat conflicts as static inconsistencies and select more reliable knowledge, overlooking that the same conflict may arise from legitimate knowledge evolution, malicious manipulation, or unresolved uncertainty. We formulate conflict origin attribution as a new problem in RAG: identifying which explanation of conflicting evidence is supported by observable context rather than simply which fact should be trusted. We propose EvoTrustRAG, a training-free framework for evolution-aware conflict attribution and evidence handling before answer generation. EvoTrustRAG represents span-grounded retrieved facts as a conflict evidence graph, evaluates grounded evolution and directional intervention hypotheses using temporal relations, support structure, and auxiliary consistency, and projects local decisions onto a globally consistent explanation of each conflict group. The attribution determines whether earlier and later states are preserved as temporal knowledge, an intervention candidate is separated from the primary context, or an unresolved conflict remains visible to the generator. Unlike provenance-based approaches focused on post-hoc analysis, EvoTrustRAG determines during inference whether conflicting evidence follows plausible knowledge evolution, exhibits intervention-like support, or cannot be reliably attributed. Experiments show that EvoTrustRAG achieves 81.4% average accuracy on benchmark-native conflict settings, improves attribution macro-F1 from 72.2% to 79.1% over the strongest baseline, and reduces the error rate under the strongest coordinated attack from 31.2% to 16.0%.
136. 【2608.07921】Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention
链接:https://arxiv.org/abs/2608.07921
作者:Kasun Dewage,Marianna Pensky,Suranadi De Silva,T. H. Bandara
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:random matrix theory, random matrix, apply Marchenko-Pastur, matrix theory, pre-trained attention weights
备注: Accepted at the International Conference on Machine Learning and Applications (ICMLA 2026); to appear in IEEE proceedings
点击查看摘要
Abstract:We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a random-like bulk and a set of spectral outliers. We validate this decomposition causally: zeroing the MP-identified outliers (signal) in Mistral-7B drives HellaSwag, MMLU, and PIQA close to random-chance performance, whereas zeroing a count-matched subset of bulk singular values causes smaller but non-negligible degradation. Across 11 pre-trained transformers we identify five recurring patterns: spectral outliers encode a dominant component of the learned structure; Q projections carry the most outliers; V projections under grouped-query attention lack a clean signal/noise separation; entry-level outliers form structured row-bands in Q and column-bands in O; and specific residual-stream dimensions persist as band outliers across layers in K and O. We close by outlining how these observations could inform parameter-efficient fine-tuning and structured pruning.
137. 【2608.07891】Detection of Self-Introductions in Legislative Testimony
链接:https://arxiv.org/abs/2608.07891
作者:Sofija Dimitrijevic,Pallavi Das,Kasey Liu,Foaad Khosmood
类目:Computation and Language (cs.CL)
关键词:legislative committee testimonies, committee testimonies, legislative committee, BERT, fine-tuned BERT
备注: Presented at AAIRC-AI4 conference, Las Vegas, NV, USA August 2026 [this https URL](https://ai4.io/aairc-ai4/)
点击查看摘要
Abstract:Self-introductions are common in legislative committee testimonies. Successfully detecting them and extracting the speaker's name is enormously helpful in the task of speaker identification in the context of government meetings. In this paper, we present a pipeline for detection of self-introductions in legislative committee testimony using machine learning. We construct a training dataset from 1.54 million utterances spanning five state legislative sessions, apply a name-matching heuristic to generate automatic labels, and train three classifiers: a decision tree, random forest, and XGBoost to find self-introductions and extract the speaker's name. We construct a feature set combining bag-of-words, positional context, structural signals, introductory phrase indicators, and discourse context features. Among the three classifiers, XGBoost achieves the best performance with an F1 score of 0.9747 and the fewest total errors; adding fine-tuned BERT probability features improves this further. As an extension, we score the full candidate dataset with a fine-tuned BERT classifier and add BERT probability outputs as features. This BERT-augmented XGBoost model improves F1 from 0.9747 to 0.9782 and reduces total test errors from 241 to 207. The primary gain over the decision tree baseline (F1 0.9323) is driven by discourse context features and the boosting ensemble strategy; BERT provides a modest complementary signal. Analysis of false positives reveals that a minority are genuine self-introductions mislabeled due to name inconsistencies in the source data, indicating that measured metrics modestly understate true performance.
138. 【2608.07886】Vision-Language Grounding as Bidirectional Concept Correspondence
链接:https://arxiv.org/abs/2608.07886
作者:Jieyu Zhang,Ziqi Gao,Luke Zettlemoyer,Ranjay Krishna
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:unidirectional localization problem, existing formulations reduce, grounding connects language, visual content, connects language
备注:
点击查看摘要
Abstract:Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as $\textit{bidirectional concept correspondence}$ over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce $\textbf{ConCor-1}$, a grounding model built on top of a pretrained vision-language model. It uses learnable $\textit{bridge tokens}$ to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that $\textbf{ConCor-1}$ consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.
139. 【2608.07881】GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
链接:https://arxiv.org/abs/2608.07881
作者:Zihua Yang,Zhencheng Xie,Junyang Chen,Liang Xie,Yiqun Zhang,Mengke Li,Yang Lu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Theory (cs.IT); Machine Learning (cs.LG)
关键词:continuous numerical measurements, discrete categorical symbols, mixed tabular data, bridge the inherent, inherent heterogeneity
备注: 13 pages
点击查看摘要
Abstract:Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical relationships, which confines the learned metric to empirical co-occurrences and ignores conceptually obvious yet statistically unobserved affinities. Although LLMs offer external world knowledge, applying their text-centric reasoning to highly abstract tabular concepts presents significant challenges. Bridging this modality gap to construct a semantically complete metric typically requires embedding LLMs into iterative metric learning loops to dynamically optimize cross-modality representations. This incurs intractable computational overhead, forcing a compromise between semantic enrichment and scalability. Therefore, we propose GRACE, an LLM-grounded framework for scalable mixed-data clustering. GRACE shifts semantic acquisition to the attribute-value level via a multi-perspective LLM querying strategy, mapping heterogeneous values into knowledge-informed descriptions. Crucially, this one-shot grounding extracts general-purpose semantic representations that embed heterogeneous attributes into a unified space, decoupling expensive LLM invocation from iterative optimization. Furthermore, GRACE cross-validates these external semantics against dataset-internal statistical evidence to ensure alignment with the dataset-specific cluster structure. Ultimately, GRACE matches the scalability of conventional statistics-driven baselines while achieving superior clustering accuracy and conceptual interpretability over 11 competing methods. The source code is available at this https URL
140. 【2608.07862】SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs
链接:https://arxiv.org/abs/2608.07862
作者:Debopriyo Banerjee,Kapil Rajesh Kavitha,Angana Borah,Xudong Han,Yuxia Wang,Parameswari Krishnamurthy,Utkarsh Agarwal,Atharva Kulkarni,Swaran Lata,Ayush Munot,Dhruv Sahnan,Aaryamonvikram Singh,Preslav Nakov,Monojit Choudhury
类目:Computation and Language (cs.CL)
关键词:large language models, Western contexts, Existing safety evaluation, grounded safety risks, safety risks present
备注:
点击查看摘要
Abstract:Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state-of-the-art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at this https URL. Warning: This paper contains text that may be offensive or unsafe.
141. 【2608.07852】"Many Are My Names": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders
链接:https://arxiv.org/abs/2608.07852
作者:Adelaide Danilov,Aria Nourbakhsh,Oleksandr Marchenko Breneur,Salima Lamsiyah
类目:Computation and Language (cs.CL)
关键词:language model internally, model internally represents, remains underexplored, assigned roleplay persona, internally represents
备注: 38 pages, 9 tables, 4 figures, 2 listings
点击查看摘要
Abstract:How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker representations using a dataset of user-expressed emotional text and corresponding model responses. We decompose three generation settings (Assistant, Roleplay, and Story) into sparse autoencoder features extracted at turn-boundary and pronoun-token positions and selected through a filtering pipeline for different depths. We characterize each surviving feature through its steering effects and activation distribution. Our main finding is that the Assistant and roleplay personas are not independent alternatives: personas retain the Assistant-associated feature core while progressively differentiating from it across layers, starting from operational machinery towards behavioral and stylistic features. Meanwhile, generated story characters lack the Assistant-associated core. Both Story and Roleplay can be distinguished from the Assistant with Immersive Simulation Mode. However, the Assistant can sometimes enter or slowly drift into it even in the default setting.
142. 【2608.07851】EMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing
链接:https://arxiv.org/abs/2608.07851
作者:Yuxuan Gu,Wuyang Zhou,Huijun Xing,Danilo Mandic
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:training deep neural, static residual pathway, deep neural networks, Residual connections rely, textbf
备注:
点击查看摘要
Abstract:Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expressivity of residual routing by incorporating multiple residual streams and learning dynamic information flow, while manifold-constrained (mHC) variants stabilize training through doubly stochastic residual mixing. However, a generator-level bottleneck remains in existing methods: they use dense, unstructured generators for pre-branch aggregation, residual mixing, and post-branch redistribution, which results in parameter count growing rapidly with the number of streams. To address this issue, we propose \underline{\textbf{T}}ensorized \underline{\textbf{E}}fficient \underline{\textbf{M}}anifold-constrained \underline{\textbf{P}}arameterization for \underline{\textbf{E}}xpressive Residual \underline{\textbf{R}}outing (\textbf{TEMPER}), which represents these generators as multi-way tensors over the input-stream, feature, and output-stream modes, and parameterizes them using tensor networks. Such a structured low-rank formulation is shown to preserve token-dependent manifold-constrained routing interface while substantially reducing parameter growth. It also promotes interpretability and intuition, as: i) tensor ranks control the dimensionality of the learned routing subspace, with full ranks recovering dense routing; while ii) the generator approximation errors bound differences in routing logits and, consequently, in the routed-block outputs. Comprehensive experiments show that TEMPER matches or outperforms existing methods across language modeling and commonsense reasoning tasks, while requiring substantially fewer additional parameters. At eight residual streams, TEMPER achieves the best CORE score while using about $84\%$ fewer additional parameters than mHC, thus showing a stronger performance-parameter efficiency trade-off.
143. 【2608.07827】From token probabilities to calibrated confidence: An empirical study of mathematical question answering
链接:https://arxiv.org/abs/2608.07827
作者:Avery Ma,Lorne Schell,Vin Bhaskara,Leila Pishdad
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:large language models, empirical accuracy, large language, generated answer, token probabilities
备注:
点击查看摘要
Abstract:Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates with empirical accuracy. Prior work has shown that token probabilities are often overconfident, we investigate whether these readily available signals can nevertheless provide well-calibrated confidence estimation for mathematical question answering. We compare single-pass estimators, which reuse token probabilities from the original generation, with multi-pass estimators, which obtain additional confidence signals through verification or stochastic forward passes. While individual token probabilities can be highly saturated, we find that aggregating token probabilities over the full sequence captures small but consistent differences between correct and incorrect generations, yielding more informative confidence estimates. Multi-pass methods can yield calibrated confidence estimates. We study two such approaches: self-verification through re-prompting, including a lower-cost in-situ variant, and Monte Carlo Dropout, which derives confidence from variation across stochastic forward passes. We further evaluate two post-hoc calibration methods, Platt scaling and isotonic regression, both of which substantially reduce in-domain calibration error. However, their data efficiency varies with dataset difficulty, and the calibration mappings often transfer asymmetrically across datasets and models.
144. 【2608.07812】On the use of foundation models in cognitive science
链接:https://arxiv.org/abs/2608.07812
作者:Raj Sanjay Shah,Alex Warstadt,Michael Frank,Sashank Varma
类目:Computation and Language (cs.CL)
关键词:Foundation Models, host of recent, recent studies, studies have evaluated, Foundation
备注:
点击查看摘要
Abstract:A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.
145. 【2608.07763】Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
链接:https://arxiv.org/abs/2608.07763
作者:Anna Kołos,Grzegorz Statkiewicz,Karolina Seweryn,Katarzyna Kowol,Karolina Piosek,Wojciech Kusa
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:achieved strong performance, visual question answering, question answering, achieved strong, strong performance
备注: 28 pages. Preprint under review
点击查看摘要
Abstract:Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to failures in interpreting region-specific meanings, symbolic content, and context-dependent visual cues. Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper linguistic and pragmatic understanding in culturally situated settings. We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs. Overall, our dataset provides a controlled and challenging resource for assessing culturally grounded vision-language understanding beyond surface-level recognition.
146. 【2608.07737】he No-Meaning Falsity: The Structural Impossibility of the Arbitrary Sign in Classical Arabic
链接:https://arxiv.org/abs/2608.07737
作者:Elnaserledinellah Mahmoud Abdelwahab
类目:Computation and Language (cs.CL)
关键词:unrestricted semantic indeterminacy, foundational Saussurean axiom, architecture of Classical, Classical Arabic, Semantic Localization Theorem
备注: 45 pages
点击查看摘要
Abstract:This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary sign, are compatible with the structural architecture of Classical Arabic. We develop a formal mathematical model of Arabic non concatenative morphology in which lexical meaning is determined by the interaction between an invariant root and a morphosyntactic pattern. Within this framework, we establish a Morphological Correspondence Theorem, demonstrating that every lexical item is uniquely generated by a root pattern pair, and a Semantic Localization Theorem, proving that lexical meaning is determined at the derivational level prior to surface realization. To address Saussurean weaker notion of relative arbitrariness, we formalize it via conditional Kolmogorov complexity, defining arbitrariness algorithmically as the no rule property. We prove that general relative arbitrariness is formally undecidable, while Arabic relative arbitrariness is decidable and provably less than 1 for its motivated signifiers (Levels W and M), establishing a strict system complexity asymmetry over Indo-European languages.
147. 【2608.07727】Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages
链接:https://arxiv.org/abs/2608.07727
作者:Venkata Naga Sai Vishnu Rohit Pulipaka
类目:Computation and Language (cs.CL)
关键词:Dravidian languages, Malayalam make, train multilingual language, small part, per-language ability
备注:
点击查看摘要
Abstract:Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so it's not clear how much per-language ability these models actually keep. I have trained five GPT-2 architecture models from scratch to compare four monolingual models (one each for Tamil, Telugu, Kannada, and Malayalam, each with its own 32K-vocabulary subword tokenizer) against one multilingual model sharing a 64K-vocabulary subword tokenizer across all four languages. All the 5 models are trained on cleaned CC-100, Wikipedia, and Samanantar data. I have tested the models on perplexity, bits-per-byte, tokenizer efficiency, and fine-tuning results which are compared against mGPT. The monolingual models outperform mGPT on sentiment classification and named entity recognition, and their tokenizers proved more efficient than the shared multilingual model across all the languages tested.
148. 【2608.07693】CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
链接:https://arxiv.org/abs/2608.07693
作者:Quang Minh Dinh,Tuan Kiet Doan
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Generative traffic video, temporally coherent future, coherent future videos, traffic video forecasting, short observation history
备注: Accepted at ECCVW 2026
点击查看摘要
Abstract:Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions. In this paper, we present CosmosAlign, a generative traffic video forecasting framework built upon the pretrained Cosmos3-Nano world foundation model. Our approach is motivated by the observation that successfully adapting large pretrained world models to downstream forecasting tasks depends primarily on distribution alignment rather than increased model capacity. To this end, we propose a two-stage LoRA adaptation strategy that first aligns the conditioning-mode distribution with the target forecasting task, and then aligns the training captions with the model's native structured prompting interface through an LLM-based re-captioning pipeline. During inference, we further improve prediction quality using a fully training-free procedure consisting of consensus-based medoid sample selection and motion-adaptive blending of static scene regions. CosmosAlign achieves a final score of 76.49 on the AI City Challenge 2026 Track 5 benchmark, ranking first on the final leaderboard. Our code is publicly available at this https URL.
149. 【2608.07688】IntelliAudit: Using Large Language Models to Evaluate Audit Controls
链接:https://arxiv.org/abs/2608.07688
作者:Allison Wilson,Sina Moradi Sabet,Diar Shakimov,Panteha Shahrivar,Mohammad Reza Bagheri,Dean Konenkamp,Mohammad A. Tayebi
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)
关键词:satisfies semantic security, heterogeneous organizational evidence, organizational evidence satisfies, evidence satisfies semantic, audits require auditors
备注:
点击查看摘要
Abstract:IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational artifacts, and because audit conclusions depend on evidentiary sufficiency rather than keyword matching. We present IntelliAudit, a retrieval-grounded multi-agent system for IT audit evidence evaluation. Given a control and an evidence corpus, IntelliAudit retrieves relevant artifacts, generates an evidence-grounded assessment, challenges adverse findings, adjudicates disagreements, and produces an auditor-facing recommendation with cited evidence, rationale, missing-evidence analysis, and remediation guidance. We instantiate IntelliAudit on ISO/IEC 27001 and evaluate it across multiple simulated organizations using expert auditor review and audit-readiness user feedback. The evaluation shows that IntelliAudit can support control interpretation, evidence-grounded reasoning, and audit-preparation workflows, while also revealing the importance of human oversight for calibrating sufficiency judgments and correcting overly permissive recommendations. These results suggest that retrieval-grounded multi-agent systems can assist audit evidence review, but should remain decision-support tools rather than autonomous certification systems.
150. 【2608.07663】Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
链接:https://arxiv.org/abs/2608.07663
作者:Yeeun Choi,Youngbeom Yoo,Joon-Young Lee,Hyolim Kang,Seon Joo Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large Language Models, Multi-modal Large Language, current Multi-modal Large, Language Models, Multi-modal Large
备注: Accepted to ECCV 2026 (Oral). Project Page: [this https URL](https://choi-yeeun.github.io/MERIT/)
点击查看摘要
Abstract:When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).
151. 【2608.07641】SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators
链接:https://arxiv.org/abs/2608.07641
作者:Yuheng Zhang,Yuanchun Wang,Fanjin Zhang,Ruyu Zhao,Juanzi Li,Jie Tang,Jing Zhang
类目:Computation and Language (cs.CL)
关键词:large language models, months-long manual effort, transformed survey writing, automated process, rapid advancement
备注:
点击查看摘要
Abstract:The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of systematic frameworks for quantifying alignment with human reviewers. To address this gap, we propose SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation. We collect and annotate 675 survey papers with 1,630 review reports. We structure authentic peer-review reports by converting free-form comments into four-dimensional scores (Readability, Criticalness, Comprehensiveness, Structure) paired with supporting rationales. We further release standardized train/test splits and an evaluation protocol to measure alignment between automatic evaluators and human reviewers. To validate the benchmark, we develop SurveyAlign, a strong baseline evaluator by fine-tuning Qwen3-32B with LoRA on our annotated data, augmented with external knowledge for knowledge-intensive dimensions. On the test set, SurveyAlign substantially improves reviewer alignment over prompt-based judging with GPT-5.2, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 across all four dimensions. Our contributions are twofold: (1) we establish the first multi-dimensional, reviewer-aligned dataset with a reproducible evaluation framework for survey reviewing; (2) we develop a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research. Our code and data are available at this https URL
152. 【2608.07629】Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation
链接:https://arxiv.org/abs/2608.07629
作者:Samiratu Ntohsi,Neza David Tuyishimire,Anesu Kafesu,Marvin Ogore,Samuel Oluwajunwonlo Babalola,Oche Ankeli
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:leave thousands unsupported, Grassfields Bantu languages, including most Grassfields, Multilingual neural machine, Grassfields Bantu
备注:
点击查看摘要
Abstract:Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection. We implemented an embedding initialization strategy where a language token is the average of embeddings from multiple typologically related languages already in the mod el. We evaluate this approach on Limbum-to-English translation using a parallel corpus of 8,837 sentence pairs from New Testament text and a bilingual dictionary. We compare models: NLLB-200 zero-shot (chrF2++ = 12.5), a Transformer trained from scratch (chrF2++ = 14.5), NLLB-200 fine-tuned with a Swahili proxy token (chrF2++ = 47.3), and NLLB-200 with our averaged embedding initialization (chrF2++ = 46.7). We find that the multi-language initialization achieves performance comparable to the best single-language proxy. Both NLLB-200 variants improve over the from-scratch baseline by over 32 chrF2++ points. These results show that multilingual transfer is the dominant factor in extremely low-resource Bantu translation while eliminating the need for heuristic proxy selection. However, all systems fail to preserve tonal diacritics, highlighting an open challenge. We make our dataset and code available to support further research.
153. 【2608.07627】From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems
链接:https://arxiv.org/abs/2608.07627
作者:Manideep Dhar,Ritwik Singh,Sharat Chandra Kumar Manikonda
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
关键词:mounting technical debt, Fortune Business Insights, triage management, edge of production, exposing patients
备注: Peer-reviewed published article
点击查看摘要
Abstract:Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage management, documentation, scheduling, and revenue-cycle workflows, yet most deployments remain as fragmented pilots that stall at the edge of production, exposing patients and institutions to operational fragility, ungoverned risk, and mounting technical debt. At the same time, the global AI-in-healthcare market is projected to exceed nearly USD 1 trillion by 2034, according to the report of Fortune Business Insights, amplifying the financial consequences of architectural missteps and failed scaling strategies. This research proposes a compliance-first Agentic AI pattern catalogue and orchestration framework, purposely built for HIMS, moving beyond the single LLM chatbots and towards a governed ecosystem of autonomous and semi-autonomous agents. The framework extends by adding (i) a taxonomy of Agentic roles, (ii) a formal risk-stratification model that maps each pattern to risk tiers, human-in-the-loop checkpoints, and governance hooks, and (iii) a unified orchestration runtime capable of coordinating multi-agent workflows across EHR/HIMS landscapes such as Epic, Cerner, and MEDITECH. Technically the framework combines vLLM-based inference, optimized paging memory, confidential computing, and MCP based on-premise deployment, enforcing end-to-end encryption and policy-as-code controls aligned with HIPAA, GDPR, the EU AI Act, India's DPDP and DISHA Acts, ISO 27001, ISO 27002, ISO 14971 and IEC 62304. We exhibit how the proposed architecture is capable and efficient to reduce the documentation time, integration effort, and AI pilot attrition while constricting the governance and auditability, offering hospital leaders and governing authorities an urgently needed blueprint to convert AI investment into sustainable clinical, operational, and financial ROI
154. 【2608.07614】DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?
链接:https://arxiv.org/abs/2608.07614
作者:Susana Haing,Natan Vidra,Spurthi Setty
类目:oftware Engineering (cs.SE); Computation and Language (cs.CL)
关键词:standard benchmarks measure, Intent Violation Rate, standard benchmarks, developer implicit intentions, tests
备注: 8 pages, 2 figures, 8 tables
点击查看摘要
Abstract:Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code passes its stated test. We introduce the Intent Violation Rate (IVR) and a 49-problem pilot benchmark derived from HumanEval+. Each problem strips implicit constraints from a clarified prompt and encodes them as hidden constraint tests. IVR measures the fraction of LLM-generated solutions that pass the stated (visible) tests yet fail hidden constraint tests that capture unstated intent. Evaluating Claude Sonnet 4.6 and OpenAI GPT 4.1, we find both pass over 92\% of stated tests yet violate intent in over half of problems (54.5\% and 63.5\%), following a systematic, bimodal pattern consistent across both models. Out findings indicate that pass rates overstate how well generated code reflects developer intent.
155. 【2608.07594】Scaling Inherently Interpretable Language Models
链接:https://arxiv.org/abs/2608.07594
作者:Guide Labs Team,Andreas Madsen,Aya Abdelsalam Ismail,Giang Nguyen,Isaac Plant,Muawiz Chaudhary,Nathaniel Monson,Saqib Azim,Zhichen Guo,Julius Adebayo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:opaque systems, difficult to establish, methods whose reliability, reliability is difficult, language
备注:
点击查看摘要
Abstract:Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, with methods whose reliability is difficult to establish. In this work, we challenge this premise. Rather than reverse-engineering a model, we make interpretability a constraint of the training pipeline, optimized alongside the language modeling objective. Across three orders of magnitude of compute, on both autoregressive and diffusion language models, interpretability scales with capability rather than against it. Surprisingly, model representations become more disentangled and aligned with human-understandable concepts with scale. We instantiate the training-time recipe with Steerling-8B, a diffusion language model with a causal attention mask. For any group of generated tokens, Steerling-8B attributes the output to relevant input tokens, human-understandable concepts, and training data. This enables closed-loop intervention: diagnose an output through its concept or feature attribution, retrieve similar training data, and correct the behavior through concept steering without retraining. Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.
Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:
arXiv:2608.07594 [cs.CL]
(or
arXiv:2608.07594v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2608.07594
Focus to learn more
arXiv-issued DOI via DataCite</p>
156. 【2608.07537】An evolutionary model of animats with VLM-based subjective evaluation
链接:https://arxiv.org/abs/2608.07537
作者:Shota Miyazaki,Takaya Arita,Reiji Suzuki
类目:Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Multiagent Systems (cs.MA)
关键词:Vision-Language Model, incorporates subjective evaluations, subjective evaluations provided, VLM, evaluation
备注: 18 pages, 12 figures, 2 tables. This manuscript has been accepted for publication in Artificial Life and Robotics following peer review
点击查看摘要
Abstract:In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness evaluation and selection processes of a genetic algorithm. As the target of evolution, we employ virtual soft robots with flexible morphologies and locomotion and present the VLM with sequence images representing the locomotion of two individuals. Selection is performed via pairwise comparisons based on subjective evaluation terms such as adorably and weirdly. The outcomes of these comparisons are used as selection pressure within the genetic algorithm, enabling the simultaneous evolution of morphology and locomotion. Experimental results demonstrate that subjective selection by the VLM accelerates population convergence compared to random selection, while also giving rise to distinctive morphologies and motions corresponding to each evaluation term. An auxiliary experiment with human participants further showed that, although individual pairwise choices only partly agreed with the VLM selections, the resulting morphological and locomotion tendencies were qualitatively similar and repeated human evaluations imposed noticeable fatigue. Moreover, the observation that similar evolutionary outcomes emerged across different evaluation terms suggests that the VLM does not apply these terms in a purely literal manner but instead decomposes them into multiple internal evaluation criteria when making judgments. This work visualizes the evolutionary process through which subjective linguistic expressions are mapped onto embodied phenotypes and provides a foundational framework for analyzing the structure of subjective judgment in VLMs. The proposed approach is expected to contribute to new developments in evolutionary computation and artificial life research based on subjective evaluation.
157. 【2608.07531】Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
链接:https://arxiv.org/abs/2608.07531
作者:Cheng Ruoxi,Ma Haoxuan,Zhang Hongyi,Zhang Junming,Duan Ranjie,Xia Qiaolin,Wang Hao,Lu Yu,Shi Haibo,Ma Xingjun
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Search-augmented language agents, Search-augmented language, retrieve external information, Search-augmented, retrieve external
备注:
点击查看摘要
Abstract:Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding. We propose Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention-calibrated readouts. A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search. After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy. Experiments across multiple search-based question-answering benchmarks and two model scales show that Search-G1 improves the grounding--search-cost trade-off, producing shorter response-side trajectories at competitive task accuracy. Code is available at this https URL.
158. 【2608.07530】NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
链接:https://arxiv.org/abs/2608.07530
作者:Yuchen Zhou,Niels Bobet,Maribel Acosta
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Databases (cs.DB)
关键词:RDF knowledge graphs, conformance of RDF, RDF knowledge, knowledge graphs, core technology
备注: 18 pages, 8 figures, 2 tables
点击查看摘要
Abstract:SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. However, there is no dedicated benchmark for NL2SHACL, and evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure. To tackle these challenges, we present NL2SHACL-Bench, a benchmark suite for natural language to SHACL translation. Using NL2SHACL-Bench, we evaluate four state-of-the-art large language models (LLMs) for this task. Our results show that current LLMs are highly capable of generating syntactically valid SHACL, but still struggle to produce semantically equivalent constraints for complex logical and structural patterns. This indicates that NL2SHACL-Bench provides a meaningful basis for measuring advances in the NL2SHACL state of the art.
159. 【2608.07529】WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management
链接:https://arxiv.org/abs/2608.07529
作者:Yi Zhang,Hongyang Wang,Zheng Hao Leong,Zihao Wu,Kaijun Lin,Zhixing Pan,Qixun Huangfu,Wei Ren,Wenyan Wu,Fangyun Wang,Wenting Yu,Hengyu Lin,Muling Yang,Zongguo Wen
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:solid waste management, Large language models, existing benchmarks emphasize, benchmarks emphasize general, emphasize general knowledge
备注:
点击查看摘要
Abstract:Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64\% accuracy on the Foundation Module, but average accuracy still fell from 84.14\% on easy questions to 42.50\% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.
160. 【2608.07528】he Knowing-Saying Gap: When Probes See Errors that Confidence Misses
链接:https://arxiv.org/abs/2608.07528
作者:Jyotin Goel,Ipshita Bandyopadhyay,Justin Shenk
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:reliable failure prediction, Linear probes detect, detect corrupted context, Linear probes, near-perfect accuracy
备注:
点击查看摘要
Abstract:Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered "persistence beats peak" hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.
161. 【2608.07527】DocAtlas: Long-Document Understanding as Mutable-State Interaction
链接:https://arxiv.org/abs/2608.07527
作者:Hongchen Wei,Yuanzhe Wang,Bei Liu,Yifan Yang,Qi Dai,Kai Qiu,Yunsheng Li,Dongdong Chen,Chong Luo,Zhenzhong Chen,Baining Guo
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Long-document understanding requires, understanding requires models, Long-document understanding, treats long-document understanding, understanding requires
备注:
点击查看摘要
Abstract:Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but often rely on frozen proprietary backbones whose behavior is set by prompts. We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process. We instantiate DocAtlas as a mutable document harness: an external environment that determines what document information is searched, read, stored, reviewed, and shown to the model at each step. Given a document and question, the harness exposes search, reading, note-taking, and review tools, maintains a hierarchical tree and note store, and updates both as the agent records evidence. DocAtlas combines self-improving retrieval, selective evidence access, and active working memory under a fixed context budget. The same harness supports inference-time use with large VLMs and end-to-end reinforcement learning for compact VLM agents. With GPT-5.4, DocAtlas reaches 71.4\% on MMLongBench-Doc, exceeding the human-expert reference of 65.8\%. A Qwen3.5-4B VLM trained with end-to-end RL in the DocAtlas environment reaches 63.7\%, compared with a 54.4\% direct-input baseline, showing that mutable document-harness design can improve compact document agents by a large margin.
162. 【2608.07525】Unified Hallucination Fuzzing for Multimodal Large Language Models
链接:https://arxiv.org/abs/2608.07525
作者:Pengfei Zhou,Jiajun Song,Zhiwei Tang,Yixing Ma,Xiaopeng Peng,Donghui Si,Yuhang Xu,Huiqi Song,Yiyuan Miao,Yichen Qian,Weihua Chen,Wangbo Zhao,Bohan Zhuang,Jiasheng Tang,Yang You
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Multimodal Large Language, Large Language, Language Models, Multimodal Large
备注: 47 pages, 17 figures
点击查看摘要
Abstract:Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications. Existing evaluations, predominantly based on static benchmarks, suffer from narrow taxonomical coverage and rapid performance saturation, failing to reflect model robustness in evolving real-world scenarios. To bridge this gap, we present a systematic evaluation framework integrating a comprehensive benchmark with self-evolving stress testing. First, we introduce UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions. Second, to address benchmark saturation, we propose Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations. Crucially, to ensure reliable assessment of dynamic inputs, SAMF incorporates a structured metric suite driven by an ensemble of multi-modal oracles. Our extensive experiments reveal that state-of-the-art MLLMs exhibit significant performance degradation under fuzzing compared to conventional settings, exposing a dissociation between reasoning capabilities and factual grounding. Furthermore, we identify a helpfulness-hallucination trade-off, where reinforcement learning alignment inadvertently exacerbates sycophancy in instruction-following tasks. The framework, code and benchmark are available at this https URL.
163. 【2608.07517】he Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction
链接:https://arxiv.org/abs/2608.07517
作者:Tyler Dooskin,Squoosh Technical Staff
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL); Applications (stat.AP)
关键词:multimodal LLM predict, multimodal LLM, LLM predict, predict which version, web page
备注: 15 pages. Pre-registered experimental program with a public, tiered claims ledger; includes powered negative results, a label-validity audit, a cross-judge shared-prior measurement (n_eff ~ 2 of 16 votes), and a first-party 15-expert human baseline. Pre-registrations, statistical harness, human responses, and the full experiment ledger are released
点击查看摘要
Abstract:Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no -- and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen's kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the "ground-truth" labels in a leading CRO agency's catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones -- a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge's confident calls are different: a vote-margin gate isolates a subset (49% coverage) reaching kappa = 0.31 on significant labels. We measure the mechanism directly -- judges differing in model or prompt agree with each other at kappa = 0.74-0.88 while agreeing with real outcomes at only ~0.2, so a 16-vote panel carries about 2 effective independent votes -- and we reproduce it in humans: 15 CRO experts agree with each other (inter-rater kappa = 0.53) but score at chance against real outcomes (kappa ~ 0). Consensus, human or model, is reproducible, persuasive, and not evidence. We release our pre-registrations, locked gates, negative results, statistical harness, human responses, and a claims ledger in which every number carries an evidence tier.
164. 【2608.07511】How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare
链接:https://arxiv.org/abs/2608.07511
作者:Dorothee Amelung,Andrew M. Bean,Sabine C. Herpertz,Felix H. Krones,Guy Parsons,Adam Mahdi,Isabella Schneider
类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
关键词:Background, LLMs, socio-communicative, medical, cs.HC
备注:
点击查看摘要
Abstract:Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. Results. The LLMs in our study showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring. Conclusion. Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors. While existing frameworks for assessing interactional competencies may support the development of more socially responsive LLMs, they will require adaptation to account for the differences in desirable behaviour between humans and LLMs.
Subjects:
Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY)
Cite as:
arXiv:2608.07511 [cs.HC]
(or
arXiv:2608.07511v1 [cs.HC] for this version)
https://doi.org/10.48550/arXiv.2608.07511
Focus to learn more
arXiv-issued DOI via DataCite
Submission history From: Felix Krones [view email] [v1]
Mon, 29 Jun 2026 23:08:58 UTC (29 KB)
165. 【2608.07508】JaleesBench: Are AI Assistants Good Spiritual Company?
链接:https://arxiv.org/abs/2608.07508
作者:M. Waleed Kadous(1 and 2),Benjamin Olsen(2) ((1) a href="http://iaser.ai" rel="external noopener nofollow" class="link-external link-http"this http URL/a, (2) Faith Family Technology Network)
类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large language models, Large language, real decisions, advisors to millions, millions of people
备注: 21 pages, 8 figures, 4 tables. Open-source harness, scenario bank, proof texts, and full evaluation (model responses and both judges' verdicts): [this https URL](https://github.com/iaser-ai/jaleesbench) . Interactive browser for inspecting scenarios, responses, and judge verdicts: [this https URL](https://s.iaser.ai/jb)
点击查看摘要
Abstract:Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person of faith is not what a model knows or professes but what its counsel does to the person who receives it. We introduce JaleesBench, which measures whether an AI agent is a righteous companion, judged by the residue an exchange leaves on the user, in the manner of the perfume-seller and the blacksmith. It comprises 140 two-turn scenarios drawn from a classical compilation organized by virtue (Riyad al-Salihin), under six adversarial pressures and three framings, scored by two frontier judges against each scenario's own supporting texts. Across eight systems: (1) generic frontier models are only middling companions out of the box but a one-page guide makes them genuinely good ones, on par with the domain-tuned assistant: the frontier APIs climb from +0.28/+0.23 to a Guided +0.84-0.87, so most of the expert's edge is companionship instruction that fits in a prompt; (2) every system caves under relational pressure, insistence and personal appeal; (3) the domain-tuned assistant's advantage is overwhelmingly its retrieval-and-prompting layer, not its base model (+0.74 over the identical underlying model); and (4) it can be used to improve existing systems: guided by its diagnosis, a single steadfastness instruction lifts a deployed Islamic assistant from +0.48 to +0.84 (Faith unstated, after pressure), matching the best guided frontier systems while preserving first-response quality. The construct is faith-general; we instantiate it for Islam as the first of a planned cross-tradition family. Code, scenario bank, and rubric are open source (this http URL), with an interactive results browser at this http URL.
166. 【2608.07499】Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients
链接:https://arxiv.org/abs/2608.07499
作者:Jiading Zhu,Xinyu Cindy Wang,Thomas Nguyen,Yan Qing Lee,Osnat C. Melamed,Peter Selby,Jonathan Rose
类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large Language Model, based Motivational Interviewing, Language Model, Motivational Interviewing, Large Language
备注: 53 pages
点击查看摘要
Abstract:The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients. Prior work on simulated clients, however, has not aligned with the specific tasks fundamental to the MI therapy approach. A key task is evoking, in which the counsellor first elicits the client's ambivalence and then strengthens the client's motivation for change. We present Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating MI counsellors in smoking cessation, designed specifically for the evoking MI task. Evoke-Sim employs structured client profiles, an evoking-specific three-stage conversation flow, and a reveal policy that regulates which client profile information might be disclosed at each stage. We show that compared to existing profile-grounded simulated clients, Evoke-Sim is better at differentiating levels of MI quality using task-aware evaluation metrics, while reducing non-grounded client statements and premature disclosure of client information, setting a higher standard for the evaluation of LLM-based MI counsellors.
167. 【2608.07497】EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations
链接:https://arxiv.org/abs/2608.07497
作者:Baptiste Moreau-Pernet
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)
关键词:powering teachable agents, evaluating instructional materials, testing learning theories, teachable agents, valuable tools
备注: Poster at the Impactful and Responsible AI Systems for Education workshop, as part of the Festival of Learning 2026
点击查看摘要
Abstract:Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tutoring dialogues, including results for two LLM-based learner simulations, and the published GitHub code.
168. 【2608.07493】he Transparency Trap: How AI Disclaimers Create Overconfidence in High-Stakes Decisions
链接:https://arxiv.org/abs/2608.07493
作者:Neil Todkar
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)
关键词:fail to function, function as intended, intended due, Current, Current AI disclaimers
备注: 6 pages, 4 figures
点击查看摘要
Abstract:Current AI disclaimers often fail to function as intended due to warning habituation and a transparency paradox. As AI-generated information becomes pervasive in everyday decision-making, effective risk communication is increasingly critical for responsible design. This exploratory study examines how disclaimer placement and persuasive cues shape trust, perceived accuracy, and disclaimer engagement across three high-stakes domains: finance, medicine, and AI-generated content. Using a mixed within-between experimental design with 378 stimulus-level responses from 52 participants, we find that advisory content was generally trusted across conditions, even when disclaimers were present. A significant domain effect showed that medical content received the highest trust ratings. In the AI domain, the findings reveal a transparency paradox: some participants interpreted disclaimers not as warnings, but as signs of system self-awareness and honesty, paradoxically increasing perceived trustworthiness. Evidence of banner blindness further suggests that standardized AI disclaimers are insufficient to prevent over-reliance. Finance and medicine provide useful comparison domains by showing how users interpret warnings differently depending on context and perceived risk. These findings have vital implications for responsible AI design, algorithmic fairness, and consumer protection when users act on potentially misleading information in high-stakes settings.
169. 【2608.07478】PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings
链接:https://arxiv.org/abs/2608.07478
作者:Jagpal Singh Jhala
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:official languages create, critical accessibility barrier, documentation exists exclusively, medical documentation exists, ASHA workers
备注: 7 pages 4 images/figure
点击查看摘要
Abstract:India's 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the patients who most urgently require this information - rural populations, ASHA workers, and patient families - are functionally excluded from understanding it. This paper presents PragyaDoc, a Universal Document Intelligence Framework that addresses this gap through a four-layer pipeline: a parallel ensemble OCR extraction layer, a geometric-lexical fusion layer, a deterministic domain structuring layer, and a dual-LLM medical reasoning and localization layer
170. 【2608.08885】owards an LLM-based method for quantifying the sexual content in song lyrics
链接:https://arxiv.org/abs/2608.08885
作者:Ignacio M. Sticco
类目:Physics and Society (physics.soc-ph); Computation and Language (cs.CL); Sound (cs.SD)
关键词:widely consumed music, consumed music genres, highly sexualized, widely consumed, consumed music
备注: 17 pages, 10 figures. Code, scoring prompt, and corpus: [this https URL](https://github.com/ignaciosticco/open-lyrics-scorer)
点击查看摘要
Abstract:Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mostly on qualitative studies and on small-scale quantitative ones. This paper has two goals. First, we present a reproducible method that uses a large language model to quantify thematic content in song lyrics along several independent dimensions. The method is not restricted to sexual content. Second, we apply it to a corpus of 1,259 songs by 12 reggaeton artists released between 2002 and 2025. The analysis covers four topics: a dataset characterization, a per-artist comparison, an analysis of how the dimensions change over time, and a comparison between our sexual-explicitness score and Spotify's own explicit flag. We release the data collection code, the scoring prompt, and the corpus, so that other researchers can replicate the approach or apply it to their own lyrics datasets.
171. 【2608.07980】he Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints
链接:https://arxiv.org/abs/2608.07980
作者:Tianle Yang,Cuiling Zhang,Chengzhe Sun,Siwei Lyu,Phil Rose
类目:Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
关键词:person voice constitutes, unique biometric trace, biometric trace analogous, regained attention, policy-making contexts
备注:
点击查看摘要
Abstract:In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the assumption that a person's voice constitutes a stable and unique biometric trace analogous to a fingerprint. Yet this conception has been repeatedly criticized and rejected by forensic voice experts throughout the decades since its introduction. Although voices undoubtedly contain speaker-related information, this simplified conception obscures the highly dynamic and context-dependent nature of speech. This article revisits the voiceprint fallacy and reconsiders what can count as evidence of speaker identity by reviewing the historical development of voiceprint identification, evidence on human voice variability, developments in forensic voice comparison, research on human and automatic speaker recognition, and the recent challenge posed by deepfake speech to speaker identity. We point out that the voiceprint metaphor and its underlying implications are scientifically misleading because they transform a probabilistic source of speaker information into an imagined stable object of identity. To avoid treating voices as imprint-like traces, we recommend that voice evidence be interpreted through validated and calibrated probabilistic frameworks that explicitly account for variability, uncertainty, and alternative explanations.
信息检索
1. 【2608.09650】Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking
链接:https://arxiv.org/abs/2608.09650
作者:Matan Fainzilber,Shlomit Plavner
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:substantial lexical gap, Reranking medical procedures, insurance information retrieval, health insurance information, information retrieval
备注: 10 pages, 6 figures, 4 tables. Code available at [this https URL](https://github.com/matanf-healthee/listwise-crossencoder-reranking)
点击查看摘要
Abstract:Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial lexical gap between patient language and clinical nomenclature. We present a systematic comparison of two reranking paradigms for this production task: (1) small cross-encoders (MedCPT, MiniLM-L12) fine-tuned with listwise learning-to-rank objectives across layer freezing configurations, and (2) Qwen3-Reranker-4B, a 4B-parameter instruction reranker whose prompt is iteratively refined via an agentic optimization loop driven by GPT-4.1. On a purpose-built dataset of 2,647 queries across 708 insurance services, we find that a 109M-parameter cross-encoder fine-tuned with ListNet outperforms the 4B-parameter model by 2.6 percentage points on NDCG@3 and 13.3 points on Spearman correlation - at 37x fewer parameters. We report practical findings, a scalable LLM based dataset construction pipeline, and deployment trade-offs relevant to production reranking systems. We release our code and a sample dataset to support reproducibility and adaptation to other domains.
2. 【2608.09634】IntHQ: Task-Interactive Hierarchical Query on Dual-Stream Representations for Generative Recommendation
链接:https://arxiv.org/abs/2608.09634
作者:Junjie Sun,Longfei Xu,Huimin Yan,Wei Luo,Kaikui Liu,Xiangxiang Chu
类目:Information Retrieval (cs.IR)
关键词:heterogeneous data, data is fundamental, fundamental to modern, models are emerging, Multi-task learning
备注:
点击查看摘要
Abstract:Multi-task learning over heterogeneous data is fundamental to modern recommendation, while generative models are emerging as the backbone of next-generation recommenders. However, the integration of multi-task learning into the generative paradigm remains largely unexplored. Existing multi-task recommenders, in both discriminative and generative paradigms, extract task-relevant features from a single task-agnostic representation and wire tasks into a predefined conversion funnel. We show that this scheme is inherently prone to a threefold collapse. Source collapse, where task-specific signals are injected late and diluted in the shared latent space. Relational collapse, where task dependencies are either implicitly absorbed by the backbone or statically fixed by predefined funnels. Hierarchical collapse, where tasks depend on features at different scales and shift across training stages. We propose IntHQ, a multi-task generative recommender with three components, each alleviating one collapse. Dual-Stream Decoupling (DSD) injects task identity into computation stream early and separates the shared context stream from the task-specific stream, alleviating signal dilution. Task-Interactive Modeling (TIM) replaces the predefined funnel with explicit cross-task interaction, letting each task condition on the realized outcomes of its predecessors with learned, input-adaptive strength. Hierarchical Querying (HQ) lets each task gather multi-scale information across different layers at different training stages. In offline evaluations, IntHQ consistently outperforms competitive encoder backbones under four representative task-head configurations. Deployed in production on Amap, serving hundreds of millions of users for travel recommendation, IntHQ yields a 1.60\% relative UVCTR lift.
3. 【2608.09605】SPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation
链接:https://arxiv.org/abs/2608.09605
作者:Wenqiao Zhu,Chao Xu,Haipang Wu,Ji Liu
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Large Language, improving recommendation systems, Language Models, emerged as powerful
备注: 16 pages
点击查看摘要
Abstract:Large Language Models (LLMs) have emerged as powerful tools for improving recommendation systems. The effectiveness of LLMs arises from their ability to harness rich textual information and their capacity to model heterogeneous user preferences based on users' interaction history. However, due to the large-scale and deep architectures, LLM-based sequential recommendation approaches generally incur high inference costs, resulting in a low return on investment. To mitigate this cost, many existing approaches resort to using only the first few tokens of item descriptions, which inadvertently discards valuable information contained in the full text, thereby leading to suboptimal recommendation performance. To address this limitation, we propose a novel Token Selection approach for Preference Optimization in LLM-based sequential Recommendation, i.e., TSPORec, which accurately pinpoints informative tokens throughout the entire textual content to improve recommendation performance. Specifically, we design a three-stage pipeline to select informative tokens and introduce a novel proxy reward to facilitate the implementation. TSPORec not only enhances recommendation performance but also improves computational efficiency. Extensive experiments across two models and datasets demonstrate the superb performance (up to 31.25%) and efficiency (up to 63.4%) of our approach compared with six baseline approaches. Code is available at this https URL.
4. 【2608.09440】MetaStrategy: Generative Ranking with Executable LLM Strategies
链接:https://arxiv.org/abs/2608.09440
作者:Chengyu Lai,Jiuning Lin,Zhibo Xiao,Xiaodong Zhu,Ruiquan Lan,Bin Zhang,Zihong Huang,Wendong Zhang,Chuxin Chen,Yinjiang Cai,Shuai Zhong,Lingqing Zhang,Dimin Wang,Jialin Zhu,Han Zhu
类目:Information Retrieval (cs.IR)
关键词:Industrial recommender systems, recommender systems rank, systems rank heterogeneous, rank heterogeneous content, Industrial recommender
备注:
点击查看摘要
Abstract:Industrial recommender systems rank heterogeneous content under coupled user, business, commercial, and experience objectives. Existing generative ranking methods typically construct item sequences directly, making them difficult to integrate with mature predictive models, operational rules, and field-level guardrails. We present MetaStrategy, a framework that instead generates a structured, executable ranking strategy. Conditioned on request context, a large language model (LLM) policy emits a typed JSON bundle controlling objective weights, content and category preferences, experience constraints, and position policies. A deterministic validator and compiler instantiate an isolated Generator that competes atomically with incumbents under the list-level Evaluator of the Generator-Evaluator (GE) architecture. We train the policy in a production-path replay environment that re-executes logged requests through the current re-ranking stack without user exposure. The method combines selection, relative-rank, and baseline-lift rewards, a self-competitive curriculum that feeds frequent strategies back as competitors, and Evaluator-routed reward-augmented on-policy distillation that transfers complementary 4B-parameter Teachers into a compact 0.8B-parameter Student. We deploy MetaStrategy in Taobao Homepage Guess You Like through diff-triggered nearline generation; LLM inference remains outside synchronous ranking, with no observable increase in response time (RT). In a seven-day user-randomized online A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and significantly improves click page views (click PV) by 2.11%, item-detail page views (IPV) by 3.12%, and transaction amount by 2.83%.
5. 【2608.09408】DREAM Technical Report
链接:https://arxiv.org/abs/2608.09408
作者:Bin Zhang,Bowen Zheng,Chao Yi,Chengyu Lai,Dian Chen,Dimin Wang,Gaoyang Guo,Jialin Zhu,Jian Wu,Jing Yu,Jiuning Lin,Lingqing Zhang,Lingyun Zheng,Mao Zhang,Mingming Pan,Ruiquan Lan,Shuai Zhong,Wen Chen,Wendong Zhang,Xiaodong Zhu,Xuan Chen,Xunke Xi,Yifan Lu,Yiheng Wang,Yue Zeng,Yujie Luo,Yuning Jiang,Zhe Hu,Zhibo Xiao,Zihong Huang,Binbin Cao,Bo Zheng,Danning Wang,Dixuan Wang,Ge Fan,Haixia Wu,Han Zhu,Hao Fang,Haoming Chen,Huiping Chu,Jian Wang,Jianjun Wu,Jiawei Wu,Jiaxin Yu,Jingwen Liu,Jinzhe Shan,Kai Meng,Kai Zhang,Keqin Xu,Kewei Zhu,Lang Tian,Leihui Chen,Li Chen,Licheng Xu,Lide Xiao,Ruitong Zhang,Shiyao Peng,Silu Zhou,Tao Wang,Wei Shi,Wenjun Yang,Xiang Chen,Xiang Gao,Xiao Ren,Xu Liu,Xuwen Wang,Yang Li,Yeqiu Yang,Yi Hu,Yinnan Song,Yuan Liu,Yunqi Gao,Zhiliang Huang,Zhujin Gao,Zongyuan Wu
类目:Information Retrieval (cs.IR)
关键词:recommender systems commonly, Industrial recommender systems, Developing Recommender Engine, cascaded retrieval, systems commonly
备注: Technical Report
点击查看摘要
Abstract:Industrial recommender systems commonly use cascaded retrieval, ranking, and re-ranking pipelines. Although efficient, these pipelines fragment information and objectives across modules, rely on rigid rules, and have limited awareness of real-time intent, leaving session-level shifts among browsing, comparison, and purchase insufficiently addressed. We present DREAM (Developing Recommender Engine with Agentic Methods), an autonomous optimization control architecture that adds a perception-aware, orchestrable, and auditable policy layer atop existing pipelines without replacing them. DREAM has two core components. First, a three-tier Intent Engine fuses on-device signals into structured L0/L1/L2 intent representations; its edge-cloud trigger chain reduces reporting volume to approximately 8.7%. Second, a Meta Engine uses a MetaModel for layered M1-to-M2-to-M3 reasoning: intent summarization, strategy planning informed by Strategy Memory, and parameter translation. It dispatches the resulting parameters through a unified outlet with safety guardrails. A Reward Dual Loop continuously optimizes both components by combining offline simulation for strategy-space exploration with online feedback for outcome calibration, forming a cycle of generation, execution, evaluation, and experience accumulation. Large-scale A/B tests on Taobao's homepage feed show that re-ranking control alone improves IPV by 2.06%, Core IPV by 2.39%, and GMV by 0.88%. Extending control to fine ranking raises these gains to 2.71%, 3.06%, and 1.31%, respectively, while consistently improving PV by more than 1%. These gains require neither replacement of pipeline models nor compromise of serving stability, supporting agentic meta-control as a viable paradigm for industrial recommendation.
6. 【2608.09393】mporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
链接:https://arxiv.org/abs/2608.09393
作者:Rose Cymbler,Daniel Guez,Laurent Fabre
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:quantify temporal misgrounding, Standard legal RAG, identify and quantify, earlier or future, legal RAG treats
备注: 13 pages, 1 figure, 4 tables. Accepted at the ICML 2026 Workshop on AI for Law (AI4Law), Seoul. Code and data: [this https URL](https://github.com/rosecymbler/fiscal-fr-bench)
点击查看摘要
Abstract:We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the applicable version is an earlier or future one. Standard legal RAG treats the corpus as static; we argue legal question answering is a temporally-indexed retrieval problem. We introduce FiscalQA Pro, pairing a versioned corpus of 32,436 article-versions of the French tax code (93 years, 1938-2031) with an all-model-hard temporal-reasoning track: 209 scored, expert-reviewed questions across 33 CGI articles (221 released; twelve flagged out of the answerable scope). At selection time, no evaluated model recovered its date-applicable answer closed-book in any of four sampling draws, and the currently in-force text lacks the gold value for all but one of the scored questions. Answers are scored deterministically via atomic ground-truth "nuggets" (regex and numeric-with-tolerance), never LLM-as-judge: an LLM judge would inherit the temporal bias it is meant to score. Across eleven models (five frontier closed-API systems plus Gemini 2.5 Pro as a substitute entry, and five open-weight), parametric knowledge yields 3.0% mean strict accuracy and RAG over a static current-version corpus 2.7%. Static RAG retrieves the date-applicable version 0% of the time, confidently citing a real but inapplicable version. Our end-to-end retriever over a multi-version index, with no oracle, reaches 98.3% mean strict; an oracle-article ablation reaches 99.1%, locating the residual gap in first-stage recall, not version selection. We additionally release a version-aware jurisprudence dataset of 69,208 citation links, together with the corpus, benchmark, model responses, and pipeline code.
7. 【2608.09270】GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
链接:https://arxiv.org/abs/2608.09270
作者:Jiahui Cui,Yan Zhao,Kan Wei,Enze Zhu,Peirong Zhang,Lei Wang,Yiru Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multimedia (cs.MM)
关键词:aerial vision-language navigation, vision-language navigation, Cross-Modal Focus Misalignment, Semantic Prototype Codebook, Semantic Prototype
备注: Accepted at the 34th ACM International Conference on Multimedia (ACM Multimedia 2026, MM '26). 10 pages, 6 figures
点击查看摘要
Abstract:Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone scenarios impose dual challenges on vision-language understanding. At the macro level, overwhelming background clutter in visual representations leads to Cross-Modal Focus Misalignment, where the model prioritizes global environmental similarities over specific object details. At the micro level, Visual Isomorphism creates ambiguity, where candidates share similar geometric structures yet differ only in subtle attributes. To address these challenges, we propose the Granularity-Aware Region Alignment and Semantic Prototype (GRASP) learning framework, enhancing discriminative capability through two synergistic strategies. Specifically, we introduce Region-Focused Alignment (RFA) to promote object-centric cross-modal alignment while suppressing background interference. Concurrently, to tackle visual isomorphism, we propose Semantic Perturbation Enhanced Matching (SPEM), which leverages a foreground-purified Semantic Prototype Codebook (SPC) to construct semantically perturbed negatives for fine-grained semantic discrimination. Extensive experiments on the GeoText-1652 benchmark and the unseen ERA dataset demonstrate that GRASP achieves competitive performance in drone-view fine-grained image-text retrieval, validating its effectiveness for cross-modal understanding in aerial scenarios. Our code implementation is available at this https URL.
8. 【2608.09077】RVANNS: Mixed-Precision Indexing and Locality-Aware Graph Traversal on RISC-V
链接:https://arxiv.org/abs/2608.09077
作者:Chengying Huan,Yudong Liu,Jianguo Wang,Lizheng Chen,Renling Yin,Weijia Chen,Ji Qi,Jiageng Yu,Junjie Xu,Jie Zhang,Chen Tian,Yanjun Wu
类目:Information Retrieval (cs.IR)
关键词:Approximate nearest neighbor, nearest neighbor search, Approximate nearest, RISC-V Vector Extension, peak arithmetic throughput
备注:
点击查看摘要
Abstract:Approximate nearest neighbor search (ANNS) on CPUs is increasingly constrained by candidate-vector movement and decoding rather than peak arithmetic throughput. Although the RISC-V Vector Extension (RVV) provides vector-length-agnostic execution and LMUL-based register grouping, generic low-precision decoding still incurs conversion overhead, while irregular graph traversal generates scattered accesses that degrade cache locality and memory-level parallelism. We present RVANNS, an RVV-oriented ANNS engine that jointly optimizes vector representation and graph locality. Its Mixed-Precision Multi-Layer Index (MPMI) represents each vector with a dense 8-bit affine base and sparse FP16/FP32 residuals, fusing reconstruction with distance accumulation and aligning widening with LMUL-sized register groups. ROrder co-locates likely co-visited graph nodes and sorts remapped adjacency lists, transforming scattered payload probes into denser, predominantly forward-moving address streams. Integrated into Milvus, RVANNS achieves 3.39x and 4.94x speedups over scalar execution on real 128-bit and 256-bit RVV processors, respectively. Under controlled HNSW configurations, it improves throughput by 2.27--2.76x over RVV SIMD+FP32 and by 1.18--1.59x over the corresponding AVX-512 and SVE baselines. On Cohere10M, it further delivers 1.82--2.27x higher QPS/W than the evaluated GPU baselines.
Subjects:
Information Retrieval (cs.IR)
Cite as:
arXiv:2608.09077 [cs.IR]
(or
arXiv:2608.09077v1 [cs.IR] for this version)
https://doi.org/10.48550/arXiv.2608.09077
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
9. 【2608.09016】PreGress: Ranking-Native Pre-training and Prompting for Graph Node Ranking
链接:https://arxiv.org/abs/2608.09016
作者:Lujie Ban,Jiasheng shi,Yingli Zhou,Kaiwen Xue,Daiyin Wang,Xubin Li,Shuanghua Li,Chenhao Ma
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:retrieval augmented generation, graph-based retrieval augmented, retrieval augmented, measuring the relative, influence analysis
备注:
点击查看摘要
Abstract:Node ranking is a fundamental problem in graph information retrieval, measuring the relative importance of nodes and supporting a wide range of applications such as influence analysis, recommendation, and graph-based retrieval augmented generation. However, exact computation of graph-based ranking measures is often computationally prohibitive at scale. Existing GNN-based ranking methods provide scalable approximations, but they are typically tailored to individual ranking criteria and require retraining for each downstream task, which limits their transferability and efficiency. Recent graph pre-training approaches aim to enable knowledge transfer across tasks, yet their learning objectives are largely misaligned with node ranking, resulting in suboptimal adaptability to ranking-oriented applications. To address these limitations, we propose PreGress, the first ranking-native pre-training and prompting framework for supporting a wide range of node ranking tasks. PreGress performs multi-task pre-training using our carefully designed objectives, including degree centrality prediction and attribute reconstruction, to jointly capture structural and attribute information. To support heterogeneous ranking criteria, we design lightweight, task-specific prompt modules that adapt a frozen ranking backbone to downstream tasks without full retraining. Experiments on six public graphs and two real-world query-to-item benchmarks---Yelp2018 and MovieLens-100K---together with a controlled five-criterion graph-access study demonstrate strong ranking quality with low task-specific state overhead.
10. 【2608.08994】Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web Intelligence
链接:https://arxiv.org/abs/2608.08994
作者:Joshua Castillo,Santosh Nukavarapu,Ravi Mukkamala
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Retrieving relevant evidence, noisy web data, Retrieving relevant, heterogeneous language, data is challenging
备注: 8 pages, 2 figures. Accepted as a Short Paper at KDIR 2026 (International Conference on Knowledge Discovery and Information Retrieval)
点击查看摘要
Abstract:Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogeneous language, and irrelevant content. We present Guardian Crawler, a reproducible retrieval-first testbed for controlled experiments on knowledge discovery and evidence-grounded summarization over synthetic web-like corpora. The architecture combines BM25 retrieval with risk-aware, embedding-augmented, and hybrid reranking, followed by constrained retrieval-augmented generation with explicit document citations. Experiments on a synthetic 900-document corpus and 10 queries produced the highest descriptive retrieval scores under risk-based reranking, with P@10 = 1.00 and NDCG@10 = 0.94, compared with 0.94 and 0.81 for BM25. The best hybrid and BM25+Semantic configurations reached NDCG@10 values of 0.94 and 0.88, respectively. All 41 evaluable generated bullets passed the lexical coverage threshold; an automated LLM judge classified 36 as supported, one as partially supported, and four as unsupported. These results demonstrate the feasibility of Guardian Crawler as a controlled testbed but do not establish statistical superiority, human-validated faithfulness, or transfer to live-web investigative environments.
11. 【2608.08944】What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence Interventions
链接:https://arxiv.org/abs/2608.08944
作者:Wenzhang Du
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:failed retrieval-augmented generation, retrieval-augmented generation, failed retrieval-augmented, RAG, eligible Qwen failures
备注: 15 pages, 2 figures, 6 tables
点击查看摘要
Abstract:A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair. We introduce Pair-ID, an offline audit that holds one query, retrieval state, and reader constant, then crosses two operations, adding missing support and deleting verified nonsupport, to measure a same-failure counterfactual response vector. A complete funnel over 19,981 benchmark queries identifies 11,105 eligible Qwen failures, from which a prospectively fixed SHA-256 ordering selects 1,200 before generating any sampled response. Among 1,190 regenerated-valid failures, support addition repairs 197/600 JOINT cases (0.328, 95% CI [0.292, 0.367]), and deletion repairs 162/1,190 cases (0.136, 95% CI [0.117, 0.155]); length- and position-matched shams retain semantic contrasts of 0.223 and 0.101. The original view carries partial predictive signal for individual response cells (macro AUROC 0.678; Brier 0.152 versus 0.160 for a marginal baseline), but exact-vector accuracy, 0.637, does not exceed the 0.646 majority-vector baseline, and vector macro-F1 is 0.170. Across four readers, both marginal sensitivities recur, while pooled exact-vector agreement is 0.675-0.765 and JOINT-only agreement falls to 0.538-0.691. These results show that evidence sensitivity occurs at meaningful rates in the hash-selected eligible-failure sample, is only partially predictable from the observed failure, and is conditional on the reader. The evidence supports a frame-scoped offline response audit, not an information-theoretic impossibility result, reader-independent taxonomy, or runtime repair policy.
12. 【2608.08940】Difficulty-Gated Fusion of Reasoning Views for Temporal Retrieval
链接:https://arxiv.org/abs/2608.08940
作者:Jamie Holdcroft,Abdelrahman Abdallah,Adam Jatowt
类目:Information Retrieval (cs.IR)
关键词:Reasoning-intensive temporal retrieval, retrieval requires matching, temporal retrieval requires, shared temporal reasoning, Reasoning-intensive temporal
备注: Accepted at CIKM 2026
点击查看摘要
Abstract:Reasoning-intensive temporal retrieval requires matching a query to documents whose relevance depends on shared temporal reasoning rather than lexical overlap. Expanding a query into several reformulations that make its temporal intent explicit, and retrieving with each, supplies this reasoning, but fusing the resulting rankings with equal weights wastes accuracy: for any single query, only some reformulations are reliable. We propose query-difficulty-gated fusion of reasoning views. From each view we read an eight-dimensional signature of its score distribution, built from query-performance-prediction quantities such as softmax entropy, score gaps, and dispersion, and a gate of roughly one thousand parameters maps these signatures to per-query view weights. The fused ranking uses no relevance labels at inference, no re-ranking, and no fine-tuning of the retriever; the gate is trained leave-one-task-out. On the \textsc{Tempo} benchmark, the method improves all six retrievers we evaluate, from BERT encoders to 7B decoder retrievers, with the largest gains on the weaker backbones. The strongest retrievers reach $0.297$ and $0.303$ nDCG@10, and the per-query gain over the original query is significant under a paired bootstrap ($p0.001$). A per-query oracle reaches $0.364$ against our realized $0.297$, exposing headroom that identifies per-query view selection as a concrete next step.
13. 【2608.08892】A Symmetric Layer-Union Audit of Component Collapse in Hierarchical Procedural Corpora
链接:https://arxiv.org/abs/2608.08892
作者:Jiuyi Zheng,Gan Xu
类目:Information Retrieval (cs.IR)
关键词:group corpus units, Component-disjoint leakage control, Component-disjoint leakage, content similarity, group corpus
备注: 19 pages, 1 figure, 9 tables
点击查看摘要
Abstract:Component-disjoint leakage control can group corpus units by content similarity, hierarchical membership, or both. Guvenilir and Dogan previously showed that merging relation types can create a giant component that obstructs splitting; we do not claim this phenomenon as new. We examine it through a symmetric audit of a content-near-duplicate layer, a common-container layer, and their union in a fixed panel of six hierarchical procedural corpora. MyFixit and Doc2Dial exhibit the individual-layer-pass/union-fail pattern under the same operational criteria. The resulting two-of-six fraction describes this deliberately constructed panel and is not a prevalence estimate. A prespecified bridge-specific predictor is associated with the pattern, but it is not distinguished from a registered union-density control; the panel therefore does not identify a bridge-specific mechanism. Secondary diagnostics bound the interpretation of threshold sensitivity, annotation coverage, and lexical cues without extending those findings beyond their recorded sources and definitions. The paper's contribution is a bounded measurement and audit: it keeps relation families visible, evaluates their individual and union component structures symmetrically, and reports negative cases and mechanism limits. It proposes neither a new splitting algorithm nor a general causal claim about relation unions.
14. 【2608.08776】Automating Freshman Course Placement and Registration: A Case Study
链接:https://arxiv.org/abs/2608.08776
作者:Bharathwaj Vijayakumar,Samyukta Alapati,Sahana Varadaraju
类目:Computers and Society (cs.CY); Databases (cs.DB); Information Retrieval (cs.IR)
关键词:explores Rowan University, Rowan University effort, Freshman Instructional Guides, report explores Rowan, Rowan University
备注: Published in EdgeCon Proceedings 1(1), 2025
点击查看摘要
Abstract:This implementation report explores Rowan University's effort to automate the process of freshman course placement and registration. Historically, Freshman Instructional Guides (FIGS) at Rowan was manually executed, requiring significant time from Testing Services, University Advising, and the Registrar's Office to evaluate placement needs and assign students to courses. Given the 57% surge in first-time degree-seeking student enrollment over a decade, the manual processes became increasingly unsustainable. In response, a cross-departmental team developed a comprehensive automated process to integrate data from Banner (Student Information System), Google Sheets maintained by Advising, and other sources. This automated process classifies students based on program groupings, determines primary and secondary course placements, checks for real-time availability and constraints in Banner, and completes course registration for freshmen in bulk. The resulting system processed over 3500 incoming students with over 350 hours in annual time savings, reduced the potential for human error, and enabled staff to shift focus from administrative work to strategic advising. This report outlines the implementation context, design architecture, technical integration, assessment methods, lessons learned, and practical implications for institutions with similar challenges.
15. 【2608.08768】BOUND: Brief-Guided Corrective Preference Distillation at Search-Control Boundaries
链接:https://arxiv.org/abs/2608.08768
作者:Qingying Niu,Ruiyang Ren,Wayne Xin Zhao,Yaliang Li
类目:Information Retrieval (cs.IR)
关键词:Large language model, agents solve tasks, Large language, based deep search, deep search agents
备注: 15 pages
点击查看摘要
Abstract:Large language model (LLM)-based deep search agents solve tasks through iterative retrieval and reasoning, but locally relevant evidence can cause persistent wrong-anchor drift, constraint drift, or local-topic drift. Existing methods supervise trajectories, outcomes, or steps, but rarely distinguish task-aligned continuations from locally plausible ones that reinforce drift. We propose BOUND, a brief-guided corrective preference distillation framework for persistent search drift. For each student-induced decision-time state, BOUND constructs a teacher-side search-state brief that preserves the original search target and key constraints while summarizing confirmed evidence, missing information, and drift status. Guided by the brief, the teacher determines whether the student's continuation contains a correctable local search-control error likely to affect subsequent decisions. Together with the rollout outcome, this assessment determines whether to construct a corrective contrast between a student-specific correction and the original continuation, or a termination contrast between a supported answer and an unnecessary retrieval continuation. Each validated state-matched preference pair operationalizes a search-control boundary. Direct preference optimization (DPO) distills these preferences into the student, while the brief and teacher-side computation remain confined to training. We evaluate BOUND on four multi-hop QA benchmarks and three deep-search benchmarks. Across the six benchmarks for which we reran baselines, BOUND leads on five datasets and 12 of 14 metrics. Under the same search-control interface and matched settings, BOUND outperforms Trajectory SFT by 5.6 EM points on Bamboogle and 4.8 accuracy points on BrowseComp-Plus. Code is available at this https URL.
16. 【2608.08732】AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval
链接:https://arxiv.org/abs/2608.08732
作者:Haoyu Zuo,Yibo Yan,Xin Zou,Shuliang Liu,Yi Cao,Mingdong Ou,Xuming Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Multi-vector vision-language retrievers, fine-grained Visual Document, incurs substantial overhead, vision-language retrievers enable, retrievers enable fine-grained
备注: 24 pages, 7 figures
点击查看摘要
Abstract:Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive compression, whereas merging does not explicitly prioritize important regions when forming representatives. We introduce AnchorFold, a training-free focus-then-fold framework for document-side index compression. AnchorFold applies Recursive Attention Propagation over visual self-attention graphs, performing multi-step propagation within each attention head and integrating scores across heads and layers. The focus stage selects the highest-centrality tokens as anchors. The fold stage assigns remaining tokens to their most similar anchors in the normalized retrieval space and summarizes each anchor-centered group through centrality-weighted aggregation. This preserves non-anchor contributions while concentrating capacity on structurally important tokens. Across ViDoRe v1/v2 and REAL-MM-RAG with three diverse retrieval backbones, AnchorFold consistently outperforms all evaluated training-free baselines at $\gamma \leq 0.20$. On ViDoRe v1/v2, it retains 98.3% of full-index NDCG@5 on average at $5\times$ compression, achieving near-lossless compression, and 92.4% at $20\times$ compression.
17. 【2608.08645】HaloMark: A Spectral Threshold for Embedding-Vector Watermarking under C2PA
链接:https://arxiv.org/abs/2608.08645
作者:Tarun Sharma
类目:Cryptography and Security (cs.CR); Information Retrieval (cs.IR)
关键词:content-provenance machinery built, primary data asset, Foundation-model embeddings, primary data, content-provenance machinery
备注:
点击查看摘要
Abstract:Foundation-model embeddings are now a primary data asset, but the content-provenance machinery built for images and audio does not transfer to them. C2PA binds to an asset with a stable bit-level or perceptual identity; embeddings have neither, since quantisation, projection, fine-tuning, and windowed averaging reshape them in normal use and break any fixed hash. We present HaloMark, a watermark for embedding vectors cryptographically bound to a C2PA manifest. It composes four standard primitives -- a block-diagonal orthogonal rotation, public whitening, an input-dependent LSH commitment, and a per-vector nonce -- around one protocol change: the producer signs the LSH commitment c into the C2PA sidecar, and the verifier reads c from the manifest instead of recomputing it. Recomputing is fragile under whitening, which flips the commitment bucket on 62% of inputs at cos = 0.96; reading the signed c reduces the verifier's score to T = T_null + beta(A)*epsilon, so security turns on a single scalar beta, which we bound rigorously for linear and non-adaptive attackers and characterise empirically for the adaptive case. We evaluate against an adversary holding polynomially many clean/watermarked pairs under one key with full sidecar visibility, across eight baselines and ten adaptive attackers including denoising-autoencoder removal. The eleven encoders separate at an empirical threshold eff_rank(Sigma)/d ~= 0.19: above it, detection AUROC stays at 0.98 or higher across every in-budget attack on the three encoders we sweep in full, and at 0.965 or higher under single-seed DAE removal on the rest; below it every variant we tested fails. Why the threshold is dimension-uniform is left open. Deployed as a Qdrant admission filter, the verifier runs at 284 us and 24 bytes of sidecar per vector, validated end-to-end against three C2PA reference-SDK bindings.
Subjects:
Cryptography and Security (cs.CR); Information Retrieval (cs.IR)
Cite as:
arXiv:2608.08645 [cs.CR]
(or
arXiv:2608.08645v1 [cs.CR] for this version)
https://doi.org/10.48550/arXiv.2608.08645
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
18. 【2608.08636】Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach
链接:https://arxiv.org/abs/2608.08636
作者:Tong Bao,Yi Zhao,Heng Zhang,Chengzhi Zhang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL); Information Retrieval (cs.IR)
关键词:entity, entity type, entity type information, plays a crucial, named entity recognition
备注:
点击查看摘要
Abstract:Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts. Recently, large language models (LLMs) have demonstrated the capacity to achieve competitive SciNER performance with minimal human effort. Existing research highlights the importance of incorporating candidate entity type information for accurate entity recognition and classification by LLMs. However, when too many candidate entity types are provided in the prompt, LLMs struggle to accurately recognize and label entities in scientific texts, where entity types are more complex than in general domains. To address this challenge, we propose TdSciNER, a type-driven approach that effectively leverages entity type information to enhance SciNER performance. In TdSciNER, we first design an entity type filter model to identify the most likely entity types present in a given sentence. Subsequently, we introduce an auxiliary multi-class entity typing task within a multi-task learning framework alongside SciNER to obtain richer contextual representations. Then, we develop a novel demonstration selection strategy based on sentence similarity and entity type diversity to activate the in-context learning capabilities of LLMs, thereby improving entity recognition accuracy across diverse scientific domains. Experiments on three datasets demonstrate that our method achieves performance comparable to fully supervised models. Further analysis validates that each entity type-driven component in TdSciNER contributes to the improvement of SciNER performance. This work provides valuable insights for future advancements in SciNER and broader information extraction tasks in scientific text mining.
19. 【2608.08634】Can Open-Weight Models Compete on Financial Text Comprehension?
链接:https://arxiv.org/abs/2608.08634
作者:Jan Spörer
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR); General Finance (q-fin.GN)
关键词:labs caught, Open-weight language models, proprietary frontier models, Financial Touchstone benchmark, recent months
备注: To be presented at the workshop International Symposium on Large Language Models for Financial Services (FinLLM@IJCAI2026)
点击查看摘要
Abstract:Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba's proprietary flagship Qwen3-Max. Anthropic's Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google's Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.
20. 【2608.08583】Structure-Preserving Projection for Mitigating Modality Bias in LLM-Based Sequential Recommendation
链接:https://arxiv.org/abs/2608.08583
作者:Tzu-Wei Chiu,Song-Duo Ma,Hsin-Yu Lin,Pu-Jen Cheng
类目:Information Retrieval (cs.IR)
关键词:Recent LLM-based recommenders, recommenders integrate textual, LLM-based recommenders integrate, Recent LLM-based, recommenders integrate
备注: Accepted at RecSys 2026
点击查看摘要
Abstract:Recent LLM-based recommenders integrate textual and collaborative signals by projecting collaborative embeddings into the embedding space of the LLM. However, this projection can introduce modality bias that distorts the underlying collaborative structure and limits the usefulness of projected embeddings. To address this issue, we propose a novel structure-preserving projection approach that maintains the relational geometry of collaborative embeddings through dedicated structure-preserving losses. Comprehensive experiments demonstrate that our approach consistently improves recommendation performance, providing a more reliable path for LLM-based recommendation.
21. 【2608.08467】LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
链接:https://arxiv.org/abs/2608.08467
作者:Minhan Cho,Soyoung Park,Kihyeon Jeong,Byeongkyu Jeon,Daejin Choi,Jinyoung Han
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Model Context Protocol, Large Language Models, Context Protocol, Large Language, Model Context
备注: 4 pages, 1 table. Accepted at the AgentSearch Workshop at SIGIR 2026, Melbourne, Australia (non-archival). Code and data: [this https URL](https://github.com/rabqatab/llm-in-mcp-matters)
点击查看摘要
Abstract:The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequently used reference data, such as identifier lookup tables, directly in the server instructions: the system-prompt text a server hands to the host application. When a query concerns an entry of the embedded table, the model can act on it immediately instead of re-discovering the same information through a search tool. We test whether client LLMs actually consume such instruction-embedded data, reporting a 54,000-trial study across 24 LLMs (9 Claude, 6 Gemini, 9 GPT) on a production legal-information MCP server. A diagnostic condition that removes the competing search tool shows that failures are dominated by behavioral preference rather than missing capability. With search unavailable, 23 of 24 models read the embedded data reliably (hit ratio at least 98%); with a search tool merely present, 9 models drop below 15%. A 2^3 factorial analysis of three instruction-level interventions reveals strong interaction effects: combining all three restores at least 86% for 20 of 24 models, but individual interventions can backfire for specific model families. Per-server prompt engineering is therefore a workaround rather than a fix; we argue that MCP host applications should provide an explicit mechanism that places server instructions ahead of tool selection in the client LLM's deliberation.
22. 【2608.08417】Personalized Communication Skills for Agentic Recommender Systems
链接:https://arxiv.org/abs/2608.08417
作者:Zongwei Wang,Min Gao,Guangyu Hu,Xinyi Gao,Junliang Yu
类目:Information Retrieval (cs.IR)
关键词:increasingly employ large, employ large language, large language model-based, language model-based UserAgents, evaluate candidate items
备注: 11 pages, 4 figures
点击查看摘要
Abstract:Agentic recommender systems increasingly employ large language model-based UserAgents to evaluate candidate items through simulated feedback before recommendations are delivered. However, existing UserAgents typically reason in isolation based on limited personal histories, which may lead to perspective narrowing: the agent evaluates candidates from a local and incomplete view, overlooks relevant preference facets, and consequently produces inaccurate judgments. A natural way to alleviate this problem is to introduce other users as advisor agents, whose diverse histories provide complementary evidence that helps the target user reconsider overlooked preference signals. Nevertheless, a generic user-advisor communication process is insufficient, as different user decision states require different forms of external advice. Based on this insight, we propose AgentCom, a personalized communication skill framework for agentic recommender systems. AgentCom organizes reusable communication skills into a shared why--what--how--who skill bank: why identifies the decision deficiency that necessitates communication, what specifies the information task, how determines the advisor interaction protocol, and who retrieves advisors capable of executing that protocol. To make the shared skill bank personalized at use time and adaptive over time, AgentCom introduces two complementary mechanisms: personalized skill routing and failure-driven skill evolution. Personalized skill routing constructs a communication path by sequentially selecting suitable skills for each user and recommendation context. Failure-driven skill evolution learns from unsuccessful communication cases and enriches the shared bank with reusable skills that address previously uncovered communication needs. Experiments show that AgentCom consistently improves recommendation performance across traditional, social, and agentic recommenders.
23. 【2608.08389】Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
链接:https://arxiv.org/abs/2608.08389
作者:Harshitha Kolukuluru,Reshma Ashok,Kirat Arora,Evan William Ciccarelli,Nischal Ashok Kumar,Lunyiu Nie,Franck Dernoncourt,Samyadeep Basu,Ryan A. Rossi,Nedim Lipka
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multiagent Systems (cs.MA)
关键词:solve open-ended tasks, agents solve open-ended, context grows rapidly, research agents solve, iterative retrieval
备注:
点击查看摘要
Abstract:Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. This leads to unnecessary token cost, higher latency, and noisier inputs for final report generation. We study marginal value estimation for context management in deep research agents and present the first systematic stage-aware comparison of pruning strategies across the pipeline. We evaluate lightweight heuristic criteria and a learned value model at pre-retrieval, post-retrieval, and pre-synthesis stages. Our results show that pruning effectiveness depends more on where pruning is applied than on the specific scoring rule: early pruning yields the largest end-to-end savings, while later pruning mainly refines the final synthesis context. Lightweight heuristics reduce token usage by up to 73% with little quality degradation, learned pruning remains competitive on selected trade-offs, and no single method dominates across quality, efficiency, and faithfulness. These findings provide practical guidance for designing efficient long-horizon agentic systems.
24. 【2608.08253】SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
链接:https://arxiv.org/abs/2608.08253
作者:Varun Pratap Bhardwaj,Garima Singh,Arun Pratap Bhardwaj
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:commonly assembled, assembled from separate, shared infrastructure, durable memory, memory operating system
备注: 35 pages, 17 figures, 4 tables. Zenodo DOI: [https://doi.org/10.5281/zenodo.21853302](https://doi.org/10.5281/zenodo.21853302) . Code: [this https URL](https://github.com/qualixar/superlocalmemory/releases/tag/v4.0.0)
点击查看摘要
Abstract:AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational components. We present SuperLocalMemory 4.0, a governed, local-first memory operating system for AI agents. The system combines dense semantic, BM25 lexical, temporal, Hopfield-associative, and spreading-activation retrieval through reciprocal-rank fusion; a governed learning and behaviour layer; bi-temporal recall; multi-scope personal, shared, and global memory; role-based access control; GDPR-oriented export and verified erasure; audit trails; and a deployment-context EU AI Act checklist. V4 introduces a reliability spine for its primary write path: generation-fenced admission, a policy registry, verifiable memory transactions with per-projection apply, verify, compensate, and erase owners, and hash-checkable completion manifests. The runtime is available through CLI, MCP, an HTTP daemon, a dashboard, editor integration, and framework adapters, and supports fully local, local-with-on-device-model, and provider-assisted modes. We evaluate eleven fault-injection and mechanism scenarios, each repeated 200 times. The released evidence bundle reports 2,200 of 2,200 deterministic repetitions upholding their scoped component properties. The governed write envelope measured 3.522 ms at p50 and 5.297 ms at p99, versus 1.835 ms and 2.569 ms for the ungoverned baseline, corresponding to in-process control-plane overheads of 1.687 ms at p50 and 2.728 ms at p99. These are scoped component and mechanism measurements, not an end-to-end multi-process or external retrieval-accuracy benchmark. The paper consolidates prior SuperLocalMemory work on privacy-preserving multi-agent memory, information-geometric retrieval, and the V3.3 Living Brain lifecycle.
Comments:
35 pages, 17 figures, 4 tables. Zenodo DOI: https://doi.org/10.5281/zenodo.21853302. Code: this https URL
Subjects:
Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
ACMclasses:
I.2.6; H.3.3
Cite as:
arXiv:2608.08253 [cs.AI]
(or
arXiv:2608.08253v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2608.08253
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
25. 【2608.08237】SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems
链接:https://arxiv.org/abs/2608.08237
作者:Muhammad Faizan Raza,Shuo(Luna)Yang,Satish Mahadevan Srinivasan
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC); Information Retrieval (cs.IR)
关键词:service level objectives, strict service level, Retrieval-Augmented Generation, systems in production, level objectives
备注: 7 pages, 5 figures, 2 tables. Authors' accepted version of a paper published in Proc. IEEE CoDIT 2026. The version of record is available at the DOI below
点击查看摘要
Abstract:Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval pipelines rely on fixed retrieval budgets that ignore query difficulty, over-retrieving for easy queries and under-serving hard ones, forcing operators to trade answer quality against SLO compliance. This paper proposes SAGE, a learned SLO-aware adaptive retrieval policy that dynamically selects the number of passages k per query. SAGE uses lightweight features derived from initial retrieval (e.g., score distributions, rank gaps, lexical signals) and is trained offline via imitation learning from an oracle that approximates optimal latency-quality trade-offs. At inference, it adds no LLM calls and minimal overhead. On Natural Questions, under a 5s P95 latency SLO, SAGE achieves 95% SLO compliance versus 30% for the best static baseline (k=20), reduces P95 latency by 36% and retrieval cost by 51% with only 2 percentage points Exact Match (EM) loss. A single policy trained on Natural Questions generalizes across HotpotQA, UnSeenTimeQA, and four LLM families (Llama, Qwen, Mistral, Gemma), consistently yielding +45-52 point SLO improvements without quality degradation.
26. 【2608.08143】DS@GT ARC at Touché: Large Language Models for Retrieval-Augmented Debate
链接:https://arxiv.org/abs/2608.08143
作者:Anthony Miyaguchi,Conor Johnston
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL)
关键词:ARC working-note submission, Retrieval-Augmented Debate task, ARC working-note, ARC submission consisted, evaluating debate responses
备注: 12 pages, 4 figures. Accepted for publication in the CLEF 2026 Best of Labs proceedings
点击查看摘要
Abstract:We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval-augmented prompting pipeline. We summarize the results from the working paper and explore whether multi-LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at this https URL and this https URL.
27. 【2608.08075】Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
链接:https://arxiv.org/abs/2608.08075
作者:Sankalp Nagaonkar,Rohit Garg,Ankit Raj,Ashish Choithani,Ashutosh Trivedi
类目:Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:return ranked files, video-retrieval systems assume, assume a bounded, return ranked, bounded corpus
备注: 33 pages, 5 figures, 17 tables. Technical report. Benchmark configurations and reproduction instructions: [this https URL](https://github.com/video-db/search-over-the-visual-world)
点击查看摘要
Abstract:Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; models interpret them at different temporal granularities; context must be selected without replaying the complete visual record; and results must stay connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer-defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability-declared indexes, distinguishing memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this model in production, exposed through a typed search surface spanning planned retrieval, stateful investigation, direct access, and grounded synthesis. We contrast this model-agnostic infrastructure, where segmentation, sampling, model choice, embeddings, and ranking are system decisions and live streams are first-class sources, with video-native foundation models offered as fixed APIs. In a semantic-retrieval comparison against a commercial video-native engine spanning 9,800+ queries over four public datasets, a pipeline of general-purpose components achieves higher macro-averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). Retrieval quality over the visual world is today governed more by system design than by video-specific pretraining, and visual-memory infrastructure can deliver it while keeping playable, source-grounded evidence first-class.
28. 【2608.07998】Give the Long-tail More SPACE: Promoting Provider Fairness in Next POI Recommendation
链接:https://arxiv.org/abs/2608.07998
作者:Anran Zhang,Jiaqi Jiang,jiahui Jin,Yuhan Zhao
类目:Information Retrieval (cs.IR)
关键词:predicts users' future, users' future destinations, historical mobility sequences, recommendation predicts users', location-based services
备注: 20th ACM Conference on Recommender Systems (RecSys '26), September 27-October 02, 2026, Minneapolis, MN, USA
点击查看摘要
Abstract:Next point-of-interest (POI) recommendation predicts users' future destinations from historical mobility sequences and has become a key component of location-based services. However, mainstream models often concentrate exposure on a small set of popular POIs, leaving long-tail merchants systematically under-exposed. While provider fairness has recently attracted increasing attention, directly applying existing provider-fairness techniques to POI recommendation is problematic: (i) users face execution constraints; and (ii) POIs face resource supply constraints. To address this, we propose SPACE (Supply- and Physics-Aware Conditional Embedding generation), a model-agnostic framework that improves long-tail POI exposure via virtual user generation under explicit feasibility and supply control. SPACE consists of three stages: (1) community inference to capture heterogeneous user execution constraints; (2) unbalanced optimal-transport allocation to decide how many virtual users each tail POI should receive from which communities under POI-specific supply budgets; and (3) constraint-guided latent diffusion to generate POI-conditional, community-consistent virtual user embeddings. The generated user-POI pairs can be seamlessly used to train existing recommenders without modifying their architectures. Extensive experiments on three real-world datasets demonstrate that SPACE substantially improves provider fairness while maintaining and often improving recommendation accuracy across multiple backbone models. Our code is publicly available at this https URL.
29. 【2608.07994】VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge
链接:https://arxiv.org/abs/2608.07994
作者:Wenqi Chen,Haofei Yang,Rui Yang,Fangming Li
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:Retrieval-Augmented Generation, knowledge question answering, complex product documentation, enterprise knowledge question, question answering
备注:
点击查看摘要
Abstract:Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integration of diverse retrieval strengths, leading to inaccurate domain routing, poor utilization of hierarchical document structures, and consequently limited reasoning capabilities over enterprise knowledge. To address these limitations, we present VDGR-RAG, which integrates vector retrieval, directory-driven reasoning, graph traversal, and iterative reflection in a unified framework for accurate enterprise knowledge QA. Specifically, VDGR-RAG is an agentic GraphRAG system that first constructs a Hierarchical Heterogeneous Knowledge Graph ($\text{H}^2$KG) from document chunks to preserve both hierarchical directory structures and semantic relationships, and then employs a set of atomic tools for knowledge retrieval that can be freely composed to navigate the $\text{H}^2$KG: (1) a directory-enhanced routing tool that uses table-of-contents (TOC) structures to route user queries to appropriate domain-specific $\text{H}^2$KGs; (2) a multi-route retrieval tool that combines vector search, TOC-based agentic search, and graph search for comprehensive knowledge retrieval; (3) a directory backtracking tool that corrects knowledge localization biases; and (4) a dynamic reflection tool that iteratively plans the next retrieval phase. We conduct extensive experiments on our enterprise product documents across four wireless domains (e.g., energy saving and fault management). Experimental results demonstrate that our method significantly outperforms a variety of RAG baselines in terms of both knowledge retrieval recall and QA accuracy.
30. 【2608.07989】PushDualGen: Enabling LLMs to Generate Semantic IDs with Interpretable Copy for Industrial Push Recommendation
链接:https://arxiv.org/abs/2608.07989
作者:Manjia Lin,Da Li,Yan Wang,Yong Jin,Zheming Ding,Wei Yuan,Lei Yan,Yanan Xia,Lu Zhang,Fan Yang,Xuanping Li,Yanan Niu
类目:Information Retrieval (cs.IR)
关键词:proactively delivers personalized, KuaiShou proactively delivers, delivers personalized content, facilitate their engagement, proactively delivers
备注:
点击查看摘要
Abstract:Push recommendation in KuaiShou proactively delivers personalized content to nearly one billion users to facilitate their engagement. Recently, generative recommendation has achieved end-to-end user personalization through semantic ID. However, their black- box characteristics make recommendation logics difficult to trace, hindering their deployment. OneRec-Thinking addresses this by incorporating CoT before generating SIDs, but this significantly increases inference cost. To support large-scale industrial applications, we propose PushDualGen, a lightweight generator, which first generates the SID and then produces a copy as a skippable explanation. PushDualGen has been deployed in Kuaishou's push recommendation system. Online A/B tests demonstrate the effectiveness of PushDualGen, delivering significant improvements in both user attraction and satisfaction. The effective play rate for videos recommended to users has relatively increased by 8.50%, while the dissatisfaction rate has relatively fallen by 37.70%. In the long term, PushDualGen optimises the content ecosystem, providing more exposure for long-tail videos.
31. 【2608.07949】Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation
链接:https://arxiv.org/abs/2608.07949
作者:Yifan Wu,Yuchen Peng,Jiaqi Chai,Yufei Qian,Xilin Li,Ke Chen,Lidan Shou
类目:Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Databases (cs.DB); Information Retrieval (cs.IR); Multiagent Systems (cs.MA)
关键词:agents increasingly rely, Autonomous agents increasingly, complete downstream tasks, increasingly rely, rely on external
备注: This paper has been accepted for presentation at VLDB 2026
点击查看摘要
Abstract:Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, existing data discovery systems remain largely retrieval-oriented: they surface candidate datasets from heterogeneous sources, but provide limited support for estimating task-specific utility, selecting cost-effective datasets under budget constraints, or incorporating trustworthy feedback from prior usage. This paper presents Guixu, a valuation-driven data discovery system for autonomous agents. Guixu employs a three-phase valuation pipeline with proxy-label propagation and multi-round knapsack optimization for task-aware data valuation. Guixu integrates agentic payment protocol to enable budget-constrained data procurement workflows. Guixu leverages on-chain data market and attestation signals for verifiable data discovery. Our demonstration highlights how Guixu enables an agent to move beyond keyword-based dataset retrieval toward task- and budget-aware, trustworthy data discovery and procurement. Attendees can interactively explore the full workflow, from NL task specification and multi-source search to data valuation and verifiable transaction feedback.
32. 【2608.07861】How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems
链接:https://arxiv.org/abs/2608.07861
作者:Henri Vanhuynegem,Weitao Xu,Yiran Shen,Guohao Lan
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Multimedia (cs.MM)
关键词:visual question answering, answer users' questions, question answering, Vision-language models, enabling smartphones
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users' questions about the physical world. Since modern VLMs remain difficult to run on mobile and edge devices, VQA systems increasingly offload inference to cloud-based VLMs. This gives mobile devices access to stronger computation, but it also makes visual input preparation a key system variable: how the image is prepared before offloading affects not only answer quality but also payload size, token cost, and system latency. Proprietary APIs expose little control over model internals or serving behavior, leaving client-side preprocessing as the main practical optimization space for downstream developers. Many such techniques have been proposed for visual offloading, yet their cost-quality impact on commercial cloud VLMs has never been studied. To fill this gap, we present VQABench, the first systematic benchmark that treats client-side input preprocessing as a controlled variable for cloud-VLM-based VQA. We evaluate 12 preprocessing techniques across three VQA datasets and four commercial VLMs from three providers, totaling 95,168 API calls. Our results show that preprocessing is not universally beneficial: its effectiveness depends on the target model, API paradigm, provider token-accounting rule, and task formulation. A poorly selected preprocessing strategy can increase deployment cost or latency while degrading answer accuracy. Overall, our benchmark clarifies when preprocessing helps, when it fails, and why, providing insights to guide future research and real-world deployment of VQA systems.
33. 【2608.07816】Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation
链接:https://arxiv.org/abs/2608.07816
作者:Donald Loveland,Liam Collins,Bhuvesh Kumar,Danai Koutra,Neil Shah
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:generate recommendations conditioned, leverage large language, Recent advances, large language models, directly generate recommendations
备注:
点击查看摘要
Abstract:Recent advances in generative recommendation (GR) leverage large language models (LLMs) as recommender backbones, enabling LLMs to directly generate recommendations conditioned on item-interaction histories. In these systems, items are often represented through semantic IDs (SIDs) added to the LLM vocabulary as special tokens. Ideally, SIDs imbue item token representations with semantic priors, thereby improving model generalization. However, standard vocabulary expansion typically initializes these tokens as random Gaussian vectors, discarding the SIDs' underlying continuous geometry and forcing the LLM to relearn token relationships from interaction data. To demonstrate the consequences of this design, we first show that training from this initialization tends to organize SID embeddings around item popularity rather than semantics. We further show that, despite partially reducing the reliance on popularity and improving cold item performance, the computationally expensive process of continual pretraining (CPT) fails to reliably recover the original semantic geometry. To address these findings, we propose a simple, parameter-free intervention that initializes SID token embeddings directly from their corresponding centroids in the semantic embedding space. Requiring only a few lines of code and no additional training or inference overhead, this drop-in approach improves pure-SFT Recall@5 by up to 16%, reaches peak performance with up to 40% fewer SFT steps, and improves cold-item Recall@5 by up to 60%. Moreover, on datasets that benefit from additional CPT, centroid initialization reaches comparable performance while requiring half as many CPT epochs. Together, our findings show that preserving SID geometry, beyond shared-prefix structure, provides a simple and effective semantic prior for LLM-based GR.
34. 【2608.07622】Controlled Memory Interference in Continual LLM Agents
链接:https://arxiv.org/abs/2608.07622
作者:Ao Ding,Hongzong LI,Shiqin Tang,Li Zhang,Liang Chen,Xuyang Chen,Zi Liang
类目:Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:Long-term memory enables, Long-term memory, memory, personalize behavior, continuity across sessions
备注:
点击查看摘要
Abstract:Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience. Yet memory evolution is not simply a process of storing more information: new experiences may reinforce, revise, or interfere with existing memory states. Existing systems mainly emphasize memory construction and relevance-based retrieval, but several memories may remain simultaneously relevant while differing in state, temporal validity, or authority. We introduce Controlled Memory Interference (CMI), a controlled diagnostic and data-generation framework for studying how agent memory evolves under different memory relationships. Across controlled memory evolution, benign accumulation has limited effects, whereas relationship-specific interference sharply suppresses update plasticity with little stability gain, either by blocking target-memory exposure or by disrupting its downstream use. Lexical and Dense retrieval exhibit distinct interference pathways, while poisoning is more sensitive to update-authority cues than to recency alone. Beyond diagnosis, CMI provides targeted examples for interference-aware memory learning, improving the distinction between valid updates and interference-inducing memories while preserving performance on original memory tasks. These findings show that memory evolution is shaped not only by memory scale, but also by interactions among accumulated experiences. More broadly, memory interference emerges as an important factor for reliable continual agent memory systems.
35. 【2608.07593】Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning
链接:https://arxiv.org/abs/2608.07593
作者:Kadharmoideen Fadurudeen
类目:Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:Context-aware recommender systems, Context-aware recommender, choose to eat, long recognized, recognized that factors
备注: 5 pages. An agentic LLM system that reasons over combined location and weather context for region-sensitive dining recommendation. Working prototype implemented and briefly deployed end-to-end
点击查看摘要
Abstract:Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Existing weather-aware food and point-of-interest recommenders, however, typically treat weather generically -- mapping conditions to preferences through hand-crafted rules or specially trained context models -- and do not capture that the culturally appropriate response to weather is itself region-specific: a rainy evening calls for hot tea and fried snacks in one culinary culture and for very different comfort food in another. Encoding such weather-by-region-by-cuisine interactions as explicit rules or training data is brittle and does not scale. We present a weather- and location-aware agentic dining-recommendation system that takes a different approach: a large language model (LLM) orchestrates tools for location and weather retrieval and then reasons in natural language over the combined context, drawing on the cultural and culinary world knowledge already latent in the model to produce region-sensitive, weather-appropriate recommendations without per-region rule tables or specialized training. We describe the agent architecture, the tool-orchestration flow (Google location services and a weather service feeding an OpenAI LLM), and the reasoning mechanism, and we report on a working prototype that was implemented and briefly deployed end-to-end. We discuss design trade-offs -- cost, latency, ambiguity handling, and fallbacks -- and we are explicit about limitations, including the absence of a formal user study and the risk of cultural stereotyping in locality-based inference. The contribution is architectural: a simple, extensible pattern for incorporating environmental and cultural context into agentic recommendation through LLM reasoning rather than engineered rules.
36. 【2608.05235】From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents
链接:https://arxiv.org/abs/2608.05235
作者:Zijie Zhuang,Changxin Lao,Pengbo Xu,Hanwen Xu,Ruochen Yang,Yingzhi He,Peng Zhang,Jiangxia Cao,Yusheng Huang,Guohong Mu,Jian Liang,Ruiming Tang,Shuang Yang,Zhaojie Liu,Wenwu Ou,Kun Gai
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:agents increasingly conduct, increasingly conduct multi-round, conduct multi-round machine-learning, multi-round machine-learning experiments, industrial recommendation settings
备注:
点击查看摘要
Abstract:Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resulting trajectories to guide later decisions. Yet a completed trajectory is not automatically evidence: generated artifacts may be unsupported or incomplete, executed rounds may be invalid or confounded, and later modifications may obscure earlier findings. We study \textbf{trajectory-to-evidence conversion}, asking what a completed research process has actually established. We introduce an evidence-grounded framework that couples bounded verification of consequential artifacts with post-execution claim qualification. A context-isolated generate--verify--repair process checks artifacts for evidence violations and missing downstream requirements before release. After execution, validity and attribution checks consolidate evidence across rounds, qualify intervention-level claims as actionable repairs, diagnostic guards, or withheld findings, and preserve admitted claims as auditable records with explicit provenance and applicability boundaries. A hybrid LLM-assisted controller subsequently applies, defers, or rejects records based on available target evidence. Record audits characterize which claims survive qualification, while downstream diagnostics identify affirmative applicability judgment as a bottleneck for the tested controller. Across paper-to-target adaptations, later rounds often improve on the first, while final rounds frequently underperform an earlier best, exposing non-monotonic trajectory evolution. Candidates produced through the complete workflow also yielded positive online lifts relative to deployed baselines.
计算机视觉
1. 【2608.09931】Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
链接:https://arxiv.org/abs/2608.09931
作者:Shravan Venkatraman,Omkar Thawakar,Ritesh Thawkar,Abdelrahman Shaker,Rao Muhammad Anwer
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:coarse scalar feedback, large language models, multimodal large language, scalar feedback, large language
备注: BMVC 2026
点击查看摘要
Abstract:Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce \textbf{CVPD} (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning. We propose a three-gate Counterfactual Criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation. On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression. It achieves gains of $+3.60$ on OCRBench, $+3.38$ on MMStar Fine-Grained Perception, and $+3.08$ on MMStar Logical Reasoning, while maintaining or improving performance on broader multimodal benchmarks.
2. 【2608.09928】Multimodal Model Diffing for Feature Discovery and Control
链接:https://arxiv.org/abs/2608.09928
作者:Hunar Batra,Lachin Naghashyar,Ashkan Khakzar,Philip Torr,Christian Schroeder de Witt,Constantin Venhoff,Ronald Clark
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large Language Models, Multimodal Large Language, Language Models, Large Language, exhibit strong visual
备注: Preprint. Accepted at ICML 2026 Trustworthy AI for Good Workshop
点击查看摘要
Abstract:Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.
3. 【2608.09926】Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
链接:https://arxiv.org/abs/2608.09926
作者:Haodong Li,Shaoteng Liu,Tianyu Wang,Chongjian Ge,Sihui Ji,Jiahan Zhang,Xin Lin,Haolin Lu,Zhe Lin,Manmohan Chandraker
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:LDR, dynamics, Latent Dynamics Reasoning, world evolves, learned dynamics
备注: Project page: [this https URL](https://lat-dyn-reason.github.io/)
点击查看摘要
Abstract:The world evolves following its dynamics, i.e., its laws of motion. However, leading video diffusion models largely fit the pixels without modeling how the pixels transit over time. Thus, they render visually plausible frames but may not accurately obey the laws. To capture the dynamics purely from pixels, we introduce Latent Dynamics Reasoning (LDR). LDR casts the latent transition as an explicit kinematic integration, where the lower-order dynamics are integrated numerically and the model regresses only the third- and higher-order residual that drives the rollout. For this integration to extrapolate better, LDR runs it on a structured latent rather than dense convolutional features. Following PhyWorld, we validate LDR on a controlled white-box physics benchmark spanning five tasks (uniform motion, parabola, collision, bouncing, looming), focusing on out-of-distribution scenarios that reveal whether a model has truly learned the underlying dynamics. LDR extrapolates the learned dynamics far better: the gap between its in- and out-of-distribution error is over 20$\times$ smaller than the video diffusion baseline's, under both single- and joint-task training at 256$^2$ resolution, while using 26$\times$ fewer parameters and running 143$\times$ faster. LDR can even generalize under severe shift: for example, trained only on red balls moving left-to-right, it correctly predicts the motion of a blue square moving right-to-left. To our knowledge, this is the first video world model that extrapolates learned dynamics beyond its training distribution. Project page: this https URL
4. 【2608.09914】Overcoming Data Scarcity and Confidentiality in Hardware Assurance via Synthetic Generation
链接:https://arxiv.org/abs/2608.09914
作者:Gijung Lee,Ronald Wilson,Damon L. Woodard,Domenic Forte
类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:scanning electron microscopy, strict intellectual property, high-quality datasets required, verify nanoscale structures, Hardware assurance relies
备注:
点击查看摘要
Abstract:Hardware assurance relies on scanning electron microscopy (SEM) to verify nanoscale structures, but assembling the large, high-quality datasets required for automated analysis is impeded by time-intensive acquisition and strict intellectual property (IP) constraints on proprietary designs. We propose a privacy-preserving pipeline that secures IP by heavily distorting the functional design while generating a visually realistic synthetic dataset from a small set of initial examples. A StyleGAN first learns the distribution of hardware layout masks to generate novel, macroscopically varied structures. Subsequently, a conditional GAN (Pix2PixHD) translates these masks into realistic SEM images that preserve authentic textures and noise. The primary finding of this work is that a segmentation model trained exclusively on this synthetic data not only demonstrates a successful "sim-to-real" transfer to real images but also outperforms a baseline model trained on the limited real dataset. Because the underlying synthetic layouts are demonstrably novel and reproduce none of the specific proprietary routing of the original design, deploying the final segmentation model mitigates the risk of exposing sensitive IP to attacks like gradient inversion and membership inference, providing a highly secure, high-performance solution for hardware assurance.
5. 【2608.09908】Beyond Hazard Resemblance: Contrastive Event Adjudication for Training-Free Video Anomaly Detection
链接:https://arxiv.org/abs/2608.09908
作者:Wenti Yin,Xiang Wang,Huaxin Zhang,Hanqing Wang,Hongbo Shao,Changxin Gao,Nong Sang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:localize abnormal events, Video anomaly detection, temporally localize abnormal, aims to identify, anomaly detection
备注: Code is available at [this https URL](https://github.com/lessiYin/CEAVAD)
点击查看摘要
Abstract:Video anomaly detection (VAD) aims to identify and temporally localize abnormal events in videos. Supervised methods learn anomaly decision boundaries from target-domain annotations but require substantial in-domain data. Existing training-free methods leverage the rich semantic knowledge and reasoning capabilities of pretrained models to interpret visual content, yet these capabilities do not directly define an anomaly decision criterion: richer anomaly descriptions better capture hazard resemblance without resolving abnormality. To this end, we propose Contrastive Event Adjudication for training-free Video Anomaly Detection (CEAVAD), which shifts the unit of inference from isolated anomaly concepts to falsifiable event hypotheses and establishes an inference-time explanatory boundary through the interaction between competing explanations and video evidence. Specifically, CEAVAD first uses public-safety knowledge to construct hazard-benign event contrasts, pairing each hazard mechanism with a generic normal account and a mechanism-specific benign counterpart. It then determines whether the target interval better supports a hazard explanation or its benign competitor, yielding a revisable contrastive boundary proposal for the target. Finally, CEAVAD adjudicates between the competing explanations to determine whether the hazard hypothesis survives the video evidence, supporting both temporally localized anomaly detection and evidence-grounded explanations. Experiments on three widely used VAD benchmarks demonstrate that CEAVAD achieves state-of-the-art performance under the training-free paradigm.
6. 【2608.09907】DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning
链接:https://arxiv.org/abs/2608.09907
作者:Mainak Singha,Niccolò Biondi,Elisa Ricci,Subhankar Roy
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large Language Models, Multimodal Large Language, multimodal instruction-following ability, strong multimodal instruction-following, shown strong multimodal
备注:
点击查看摘要
Abstract:Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client-specific drift, we introduce a public-anchored expert composition stage that updates only routers and lightweight private projection adapters on a mix of local client data and public data, via an isotropic regularization loss, therefore making it cross-client rehearsal-free composition. During inference, DistMoE performs modular routing over public and private experts, enabling token-wise domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show that DistMoE enables flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge. Codes are available at this https URL.
7. 【2608.09887】Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football
链接:https://arxiv.org/abs/2608.09887
作者:Seongjin Choi
类目:Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY); Machine Learning (cs.LG)
关键词:number in football, spent pinning, most-cited and most-misleading, most-misleading number, FIFA World Cup
备注: 11 pages, 2 figures. Code and data pipeline: [this https URL](https://github.com/nowayfootball/junk-possession)
点击查看摘要
Abstract:Ball possession is the most-cited and most-misleading number in football: 60% recycled in one's own half is not 60% spent pinning the opponent back. Existing event-based possession-value frameworks (expected threat, VAEP, on-ball value) price on-ball actions but ignore the off-ball question a sterile possession poses: did holding the ball create space, or was the circulation dead? We answer this in two layers. First, an event-side junk-possession index prices each possession sequence by its peak threat gain under an expected-threat grid and -- after reconstructing the live scoreline to exclude lead-protecting circulation -- flags low-threat sequences in tied-or-losing states. On the 2026 FIFA World Cup (103 matches, 206 team-matches) the flag correlates negatively with points (r=-0.37) and xG difference (r=-0.51, partly index-coupled). It is not a repackaging of on-ball value: with team offensive VAEP and field tilt held fixed, the junk flag stays strongly negatively associated with points (p0.0001, also match-clustered) while VAEP is not significant -- in this same-match (descriptive) regression it adds information beyond this on-ball action-value model. Second, for a flagged window we resolve whether it was spatially dead or space-creating by projecting broadcast video to pitch coordinates and measuring a Space-Creation Index (SCI): a net pitch-control change capturing whether the possession seized space or pushed the opponent's block back. Across 31 of 35 flagged windows from nine World Cup matches (a purposive sample), 74% are spatially non-space-creating, 19% weak progression, and 6% space-creating windows the event flag alone would score as failure -- including a side with 73% of the ball that exited on penalties (two non-creating windows). The two layers separate space-creating-but-unconverted from sterile possession, a distinction event-only on-ball value cannot make.
8. 【2608.09885】SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
链接:https://arxiv.org/abs/2608.09885
作者:Wanying Qu,Qinghua Mao,Yu Li,Jiyao Liu,Xin Zhang,Dadi Guo,Yanxu Zhu,Qingyu Liu,Leitao Yuan,Xi Lin,Shanfeng Zhu,Yanwei Fu,Jing Shao,Xia Hu,Dongrui Liu
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:large language model, manages context, runtime control, large language, LLM
备注: Project: [this https URL](https://github.com/RainbowQTT/SHE)
点击查看摘要
Abstract:The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety mechanisms often treat the harness as a fixed deployment artifact, limiting their ability to evolve with emerging risks. Moreover, coupled functions across harness components obscure safety responsibility attribution, making localized evolution difficult. We propose Safety Harness Evolution (SHE), a framework that learns evolving safe boundaries from rollout trajectories. SHE decomposes the harness into four artifacts with explicit safety responsibilities, including the System Prompt, Rule Bank, Safety Memory, and Tool Policy, defining clear functional boundaries for localized evolution. Based on this decomposition, SHE introduces an attribution-guided evolution loop that converts trajectory failures into structured diagnoses, learns artifact-specific boundary refinements, and selects evolved harnesses through safety-utility validation. Experiments on Agent-SafetyBench demonstrate that SHE effectively enhances safety through harness evolution, achieving a 3.1x ASR reduction compared with static SafeHarness, while also improving benign utility. The evolved harness further generalizes to unseen risks on the held-out AgentHarm benchmark and transfers across agent models without additional evolution.
9. 【2608.09880】Financial Numerical Prediction and Allocation as Token Generation
链接:https://arxiv.org/abs/2608.09880
作者:Xu Ouyang,Moontae Lee
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:object ultimately evaluated, prediction typically relies, numerical object ultimately, task-specific regression, ultimately evaluated
备注:
点击查看摘要
Abstract:Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ultimately evaluated. We investigate whether a causal language model can instead represent forecasts and decisions directly through constrained token generation. FinATOM introduces a unified, head-free interface for three-step stock-return forecasting and dynamic five-ETF allocation. The forecasting model autoregressively emits volatility-standardized return tokens and is trained with ordinal and ranking supervision followed by a one-epoch token-level policy stage. The allocation model generates normalized long-only weights; supervised fine-tuning imitates a causal mean--variance anchor, and DAPO-augmented GRPO optimizes realized 21-day Sharpe subject to anchor consistency. In 2023--2025 ETF tests, the allocation policy improves pooled gross Sharpe from 1.428 to 1.529 and net Sharpe under a 5-bp transaction-cost model from 1.394 to 1.494. The multimodal allocation input attains the highest three-period mean Sharpe of 1.540, with its clearest advantage in 2025. On FinTexTS, the SFT and policy strategies achieve 73.52\%/2.68 and 73.72\%/2.69 cumulative-return/Sharpe, respectively. These results support the feasibility of direct language-model token generation for financial numerical prediction and decision-making, while motivating broader tests across assets, regimes, and random seeds.
10. 【2608.09873】Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
链接:https://arxiv.org/abs/2608.09873
作者:Diandian Zhang,Tingyu Song,Lin Fu,Zheyuan Yang,Yilun Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Humanities Social Sciences, reasoning-intensive video generation, introduce Sci-VBench, evaluating knowledge, Natural Science
备注: COLM 2026
点击查看摘要
Abstract:We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.
11. 【2608.09861】owards Expert-level Medical AI for Real-time Video Consultations
链接:https://arxiv.org/abs/2608.09861
作者:Mahvish Nagda,Jihyeon Lee,Matthew Thompson,Chunjong Park,Tim Strother,Valentin Liévin,Roma Ruparel,Akshay Goel,Teya Bergamaschi,Suhana Bedi,Meet Shah,Pavel Dubov,Liviu Panait,Toshiyuki Fukuzawa,Sam Schmidgall,Craig Schiff,Joseph Xu,Aliya Rysbek,Yana Lunts,Jan Freyberg,Rebecca Hemengway,Sunny Virmani,David Racz,Carey Radebaugh,Joëlle Barral,Kavi Goel,Dale R. Webster,Katherine Chou,Avinatan Hassidim,Yossi Matias,James Manyika,Gregory Wayne,Tao Tu,Yun Liu,Ethan Goh,Christina Chen,Ryutaro Tanno,Po-Hsuan Cameron Chen,Mike Schaekermann,Anil Palepu
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:enabling natural communication, enabling natural, AMIE, standard for patient-physician, natural communication
备注:
点击查看摘要
Abstract:Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.
12. 【2608.09853】RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
链接:https://arxiv.org/abs/2608.09853
作者:Dongchi Huang,Hongyin Zhang,Bohan Hou,Siteng Huang,Zhian Su,Hang Guo,Tong Lu,Zhaofeng Xu,Jiahao Tang,Jianfei Yang,Donglin Wang,Peixi Peng,Mingxiu Chen,Deli Zhao,Xin Li
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:corpora remains underexplored, large-scale heterogeneous corpora, heterogeneous corpora remains, General-purpose reward models, learning value-related capabilities
备注: 23 pages, 5 figures
点击查看摘要
Abstract:General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's tau_a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.
13. 【2608.09842】From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing
链接:https://arxiv.org/abs/2608.09842
作者:Jutao Xiao,Yuan Qu,Dongsheng Ma,Fan Wu,Tianyao He,Weihong Li,Jie Yang,Yu Qiao,Bin Wang,Conghui He
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent document parsers, audit reveal persistent, Recent document, reveal persistent failures, complex real-world tables
备注:
点击查看摘要
Abstract:Recent document parsers achieve table TEDS scores above 93 on OmniDocBench v1.6, yet community feedback and our audit reveal persistent failures on complex real-world tables. To quantify this gap, we introduce TableParseMap, a diagnostic benchmark of 916 real-world tables organized into five challenging scenarios and nine failure types. The strongest evaluated parser achieves only 85.03 TEDS, showing that aggregate benchmark scores conceal substantial weaknesses. Our analysis attributes these failures to three complementary limitations: large tables exceed the reliable processing scale of a single pass, weak or ambiguous visual cues hinder structure perception, and the reconstructed table may remain visually inconsistent with the image. We therefore propose DEC (Decompose--Enhance--Correct), a visual-consistency-guided agentic framework that improves frozen table parsers without retraining. DEC uses a general VLM as the controller: Decompose partitions large tables along structure-aware boundaries, Enhance exposes weak visual evidence and reparses transformed views, and Correct diagnoses and repairs residual errors. A Visual Consistency Gate (VC-Gate) selectively triggers intervention, while a Visual Consistency Ranker (VC-Ranker) verifies candidate updates and supports rollback without ground-truth HTML at inference time. We further derive a 1,977-table Consensus-Hard Set from 4,556 candidates through offline metrics and cross-model consensus. Across three frozen parsers, DEC improves TEDS by 1.57 points on average; on TableParseMap, gains reach 1.89 points overall, 2.62 on structural errors, and 5.66 on large tables.
14. 【2608.09818】MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation
链接:https://arxiv.org/abs/2608.09818
作者:Haoyu Yang,Meixing Shi,Zengjie Chen,Haoran Sun,Haitao Leng,Xiaoming Shi,Yuxiang Cai,Yankai Jiang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Reliable medical image, image understanding requires, Reliable medical, medical image understanding, understanding requires models
备注:
点击查看摘要
Abstract:Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language--mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at this https URL.
15. 【2608.09801】Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization
链接:https://arxiv.org/abs/2608.09801
作者:Dinh Tan Nguyen,Quang-Hien Kha,Le-Hoang Nguyen,Minh-Toan Dinh,Xuan-Huy Nguyen,Dac Phu Ho,Cao Truong Tran,Sai Ho Ling,Lan T Ho-Pham,Liem Pham,Nguyen Quoc Khanh Le
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Joint exam-level prediction, Joint exam-level, improve the usefulness, exam-level prediction, prediction and candidate-region
备注: Medical Imaging with Deep Learning 2026 - Short Paper Track
点击查看摘要
Abstract:Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this setting using a multi-task DETR framework, where shared representations support both image-level malignancy prediction and lesion localization, and evaluate its performance on OPTIMAM and a biopsy-confirmed SGM1k cohort. Across both datasets, modern backbones consistently outperformed older ResNet-style features, with ConvNeXtV2 and DINOv3 giving the strongest overall results, whereas MambaVision was less competitive. On OPTIMAM, ConvNeXtV2 achieved the best overall performance, reaching 97.96% AUC, 99.89% sensitivity, 25.08% mAP@.5, and 74.38% recall@.25. On SGM1k, DINOv3 gave the strongest overall results, with 90.97% AUC, 86.28% sensitivity, 82.00% specificity, 27.04% mAP@.5, and 77.32% recall@.25. These findings suggest that backbone quality is a critical factor in effective multi-task mammography, with ConvNeXtV2 emerging as a particularly strong and well-matched CNN backbone for mammography in this framework.
16. 【2608.09789】ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection
链接:https://arxiv.org/abs/2608.09789
作者:Jingtai He,Shiyuan Meng,Wenchao Meng,Qinmin Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Industrial anomaly detection, normal visual patterns, requires identifying fine-grained, identifying fine-grained deviations, requires identifying
备注:
点击查看摘要
Abstract:Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) can improve recognition accuracy by comparing query images with references at inference time, but these benefits rely on additional retrieval and processing. We investigate whether the benefits of reference comparison can instead be internalized in the model parameters. Access to references during training allows a reference-aware teacher to supervise a query-only student. However, the teacher may favor plausible responses based on query cues or language priors rather than valid visual information. We propose ADOPD, a reference-privileged on-policy distillation framework. The teacher evaluates student-generated rollouts under matched and mismatched references. The matched-reference teacher-to-student log-ratio defines the token-level learning direction, specifying what the student should learn. The likelihood gap between the two reference views estimates reference-specific support and calibrates the sequence-level weight. ADOPD achieves 77.31% average accuracy on the MMAD benchmark under zero-shot inference, improving the Qwen3-VL-4B backbone by 6.14 points and outperforming its one-shot setting by 2.64 points. Experiments show that ADOPD learns a fine-grained anomaly inspection strategy from reference comparison. The project will be available at this https URL.
17. 【2608.09782】NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge
链接:https://arxiv.org/abs/2608.09782
作者:Aleksei Khalin,Egor Ershov,Artyom Panshin,Sergey Korchagin,Georgiy Lobarev,Arseniy Terekhin,Sofiia Dorogova,Amir Shamsutdinov,Yasin Mamedov,Bakhtiyar Khalfin,Bogdan Sheludko,Emil Zilyaev,Nikola Banić,Georgy Perevozchikov,Radu Timofte,Shuai Liu,Yuqian Zhang,Lize Zhang,Yibin Huang,Chaoyu Feng,Luyang Wang,Xiaotao Wang,Dongqing Zou,Lei Lei,Tianli Liu,Dejun Hao,Chunxia Lei,Furkan Kınlı,Andrei Mironov,Alexander Dikov,Aleksei Sadokhin,Vladimir Zvorygin,Constantine Habarlak,Shuwei Yue,Egor Mirantsov,Daniil Okunev,Dmitry Arkhipov,Aleksandr Yugay,Anas M. Ali,Bilel Benjdira,Wadii Boulila,Wei Zhou,Linfeng Li,Lingdong Kong,Jiachen Tu,Guoyi Xu,Yaoxin Jiang,Jiajia Liu,Yaokun Shi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Twilight Cowboy Challenge, Twilight Cowboy, Cowboy Challenge, paper presents, presents a review
备注: 11 pages, 6 figures, 1 table
点击查看摘要
Abstract:This paper presents a review of the NTIRE 2026 Low-light Enhancement: Twilight Cowboy Challenge. The objective of the competition was to merge a set of misaligned smartphone images in the raw domain, captured in low-light conditions, into a single, clean image. Introduced setup simultaneously addresses two problems of low-light photography: visual degradations such as high noise and mixed scene illuminants, and the geometric inconsistencies caused by hand movement during multi-frame capture. To advance research in low-light and nighttime computational photography, a challenging dataset was collected comprising 585 real-world scenes, spanning indoor low-light and outdoor nighttime conditions, for training and benchmarking participant solutions. The competition employed a three-stage evaluation protocol: automatic validation via the CodaBench platform in stages one and two, followed by blind assessment on a private test set for the final ranking. Ten teams surpassed the established baseline, achieving improvements of up to +6.49 dB in PSNR and +0.0101 in SSIM, thereby establishing new state-of-the-art performance for burst-based low-light image enhancement. These results demonstrate significant progress in handling real-world noise, motion, and illumination variability in the low-light setting. Comprehensive results, leaderboards, and additional information are publicly available at this https URL.
18. 【2608.09774】C$^2$A: Coupling Spatial Evidence with Clinical Priors via Co-occurrence Aware Class Attention for Multi-Label Chest X-Ray Classification
链接:https://arxiv.org/abs/2608.09774
作者:Akash Gogineni,Nagur Shareef Shaik,Aasrith Mandava,Adnan Masood,Dong Hye Ye
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV); Signal Processing (eess.SP)
关键词:Thoracic pathologies rarely, pathologies rarely occur, standard multi-label classifiers, multi-label classifiers rely, Thoracic pathologies
备注: Accepted at 2026 IEEE International Workshop on Machine Learning for Signal Processing
点击查看摘要
Abstract:Thoracic pathologies rarely occur in isolation, yet standard multi-label classifiers rely on shared global descriptors, discarding \emph{where} findings lie and \emph{how} they co-occur. We propose \textbf{C$\mathbf{^2}$A} (Co-occurrence Aware Class Attention), a classification head that explicitly couples spatial evidence with clinical priors. First, C$^2$A casts pooling as an expectation over learned per-class spatial attention maps, yielding localized descriptors for each disease. Second, it couples these descriptors via a learnable graph warm-started from empirical label co-occurrence. A single residual message-passing step shares evidence among related findings, proving to be a bounded perturbation of the identity where co-occurrence enters each logit through an explicit bilinear interaction. On CheXpert, C$^2$A achieves a superior $0.895$ macro-mean AUROC, outperforming advanced context-gating baselines. Crucially, gains concentrate on highly co-occurrent classes with ambiguous spatial evidence (rescuing Atelectasis by $+1.5$ over GCG), demonstrating the prior's regularizing effect with a negligible overhead of one linear projection and a $C\!\times\!C$ edge matrix.
19. 【2608.09765】REFRAMED: Towards Realistic Audio Description Generation for Movies
链接:https://arxiv.org/abs/2608.09765
作者:Igor Sterner,Mirella Lapata,Alex Lascarides,Frank Keller
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:visually impaired audiences, key visual content, Audio Description, enabling access, impaired audiences
备注: COLM 2026
点击查看摘要
Abstract:Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video captioning, AD is a structured editorial task: descriptions must be inserted into gaps in dialogue and must convey only what is needed to understand the narrative being told. However, existing approaches formulate AD generation in an artificial setting where both the content and timing of descriptions are pre-specified, reducing the task to clip-level captioning. They further rely on noisy transcription and alignment pipelines, and lack the rich parallel data required for modeling narrative context. We introduce a new formulation of AD generation in which models must jointly decide what to describe and when to do it. To support this, we present REFRAMED, a high-quality dataset of 2,023 videos that span 3,302 scenes from 206 movies, with professional AD transcripts (both American and British versions), professional subtitles and aligned screenplays. We also provide a manually curated challenge set that pairs full movies with multiple AD references, together with evaluation protocols that leverage dialogue gaps and multi-reference comparisons. Experiments with state-of-the-art AD systems and multimodal LLMs show that they outperform trivial baselines but fall far short of expert human performance. Our dataset and benchmark establish a new foundation for research on video understanding.
20. 【2608.09752】Disentangling Co-Occurring Retinal Pathologies with Saliency-Guided Sparse Expert Routing
链接:https://arxiv.org/abs/2608.09752
作者:Nagur Shareef Shaik,Jeongwoo Park,Yeong-Jin Kim,Jaeuk Jung,Hyunjung Oh,Dong Hye Ye
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV); Signal Processing (eess.SP)
关键词:classifiers apply static, underlying disease distribution, fundus images frequently, images frequently exhibit, frequently exhibit multiple
备注: Accepted at 2026 IEEE International Workshop on Machine Learning for Signal Processing
点击查看摘要
Abstract:Retinal fundus images frequently exhibit multiple co-occurring pathologies, yet standard deep learning classifiers apply static, identical computation to every image regardless of the underlying disease distribution. We propose a novel architecture that resolves this via sparse conditional computation, pairing a Guided Context Gating (GCG) spatial attention front-end with a sparsely-routed Mixture-of-Experts (MoE) block operating over feature tokens. Crucially, this routing yields an interpretable, data-driven decomposition. Expert allocation is significantly disease-dependent (p 0.001), with the healthy Normal state and morphologically distinct pathologies (e.g., ERM, AMD) isolating to dedicated experts. On a five-class, patient-disjoint 5-fold cross-validation benchmark, our model achieves 0.912 +/- 0.008 macro AUC and 0.653 +/- 0.014 macro F1. Furthermore, Grad-CAM++ and post-MoE t-SNE visualizations confirm that expert routing aligns with localized lesions and geometrically maps co-occurring cases between their constituent clusters, positioning sparse MoE as an interpretable approach to multi-disease retinal screening.
21. 【2608.09735】HandSplatter: Automated Digital Goniometry from Neural Rendering
链接:https://arxiv.org/abs/2608.09735
作者:Emmett Chen,Neal Chen,Xiang Li,Quanzheng Li,Siyeop Yoon
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:quantify joint motion, musculoskeletal disability, disorders are leading, leading contributors, contributors to musculoskeletal
备注: Accepted for publication in the Proceedings of the 48th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC 2026), Full Paper #1985
点击查看摘要
Abstract:Hand and finger disorders are leading contributors to musculoskeletal disability, creating a clinical need for precise methods to quantify joint motion. Range of motion (ROM) serves as the metric for diagnosis, rehabilitation monitoring, and evaluating surgical outcomes. Currently, the goniometer is the standard tool for assessing finger flexion and extension. However, manual goniometry is labor-intensive and suffers from inconsistent inter-rater reliability due to variations in examiner technique. While digital alternatives exist, current software-based approaches often lack the necessary accuracy for clinical usage. To address these limitations, we present a novel pipeline for 3-D hand joint location and pose estimation using neural rendering. Unlike previous methods, our approach combines 2-D feature extraction with view synthesis to significantly improve accuracy and clinical viability. Furthermore, we introduce a discrete density hill climbing algorithm that facilitates the meaningful correction of projected landmarks in 3-D space. This system overcomes the inefficiencies of manual measurement and the inaccuracies of existing software, providing a robust tool for objective functional assessment.
22. 【2608.09730】World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
链接:https://arxiv.org/abs/2608.09730
作者:Qu Tang,Benhui Zhuang,Bo Yuan,Xue Yu,Longteng Guo,Junlan Feng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:widely adopted paradigm, World Adapter, world, World Tokens, widely adopted
备注:
点击查看摘要
Abstract:Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation. It uses world modeling during training to enhance the action policy while preserving efficient deployment. Specifically, the World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and simultaneously serve as the action expert's sole visual-language context. This shared conditioning allows gradients from future-video denoising to directly shape the representation used for action prediction, while exclusive routing prevents the policy from bypassing that representation. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert, with no online video-model inference. With a 2B backbone and no embodied action pretraining, World Tokens is highly competitive on LIBERO, attains the best reported averages on SIMPLER, substantially improves real-world R1 Pro success over a matched action-only baseline, and generates each action chunk at VLA-level latency.
23. 【2608.09723】LookAgain: Closed-Loop GUI Grounding with Visually Grounded Reflection
链接:https://arxiv.org/abs/2608.09723
作者:Renshan Zhang,Haoyang Meng,Yixiao He,Rui Shao,April Hua Liu,Liqiang Nie
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent graphical user, graphical user interface, densely packed controls, significantly advanced single-shot, advanced single-shot accuracy
备注:
点击查看摘要
Abstract:Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, densely packed controls and out-of-distribution interfaces. We attribute this gap to a paradigmatic limitation shared by existing approaches: none of them treats a produced coordinate as a hypothesis to be reflected upon and revised under new visual evidence. This manifests as three coupled issues: 1) Lack of post-hoc reflection. The prediction is frozen at the moment of emission, leaving no internal mechanism to challenge or refine it. 2) Visual evidence decoupled from the prediction. The auxiliary visual evidence is gathered to support the upcoming coordinate rather than to scrutinise the one already committed to. 3) Refinement over views, not over predictions. The iterative zoom-in refines the inspected region instead of inheriting a previous coordinate as a spatial prior to be corrected. In this paper, we propose LookAgain, a closed-loop GUI grounder driven by post-prediction visual reflection. LookAgain reformulates grounding as a multi-turn predict-look-again-refine process with two primitives: "locate" posts a coordinate hypothesis, renders a marker on the image and appends a local patch of the predicted region. It anchors the next reasoning step to the previous prediction as a spatial prior; "confirm" accepts or reject the hypothesis and terminates the procedure. We train the LookAgain grounder with SFT on constructed reflective trajectories as a cold start, followed by GRPO with terminal grounding correctness as the sole reward. Extensive experiments show that LookAgain consistently improves performance on both refusal-aware and general GUI grounding benchmarks, achieving state-of-the-art results. Comprehensive ablations further verify the effectiveness of the proposed framework.
24. 【2608.09691】Diffuse the object, keep its label: curating detector training data from a few unlabeled photographs via VLM-built 3D vegetation scenes
链接:https://arxiv.org/abs/2608.09691
作者:Mario Malizia,Marnix Enting,Rob Haelterman,Ken Hasselmann
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:small objects hidden, generalize poorly, labeled training images, synthesize labeled training, Labeled images
备注: Accepted at the Curated Data for Efficient Learning (CDEL) Workshop @ ECCV 2026
点击查看摘要
Abstract:Labeled images of small objects hidden in vegetation are scarce, and detectors trained on them generalize poorly across sites. Rather than reusing labels collected at another site, we synthesize labeled training images from a handful of unlabeled photographs of the deployment site itself. A vision--language model generates a coarse 3D vegetation scene from one photograph; placing 3D object meshes in the scene yields bounding boxes, segmentation masks, and per-instance occlusion directly from the scene geometry, without manual annotation. A lightweight adapter fine-tuned on the photographs conditions a diffusion pass that re-textures the renders, and a graded mask-lock sets how much diffusion may touch the object itself. In our runs this grade was the most influential curation choice: lightly diffusing the object improves minority-class recall over fully protecting its pixels, while unrestricted diffusion dissolves it. Trained on these images, a standard detector matched or exceeded its counterpart trained on a larger labeled dataset of real images from a different site, consistently across seeds on a humanitarian-demining benchmark; the comparison is thus unsupervised site adaptation from a handful of photographs against conventional cross-site label reuse. In our ablations the gains were largely insensitive to the photograph and crop budgets, and in-domain accuracy did not predict cross-site performance.
25. 【2608.09682】hinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
链接:https://arxiv.org/abs/2608.09682
作者:Jiahao Shao,Yuanbo Yang,Yiyi Liao,Yujun Shen,Ceyuan Yang,Yinghao Xu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Tool-augmented vision-language models, Tool-augmented vision-language, vision-language models increasingly, returned, returned pixels
备注:
点击查看摘要
Abstract:Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not carry the gain, what does? We hypothesize that the load-bearing signal is the structured text emitted before any returned pixel arrives: tool name, coordinates, target description, and intent. This textual scaffold encodes where to look and what to find. We introduce TextCall (call-but-no-return) to test this: it keeps the scaffold but replaces returned images with the text placeholder [Image output skipped]. Three studies support the hypothesis. (i) Non-necessity of returned pixels: across LoRA, full fine-tuning, and RL, TextCall matches or exceeds full thinking-with-images; under RL it preserves tool use at the reported checkpoint, avoiding the failure mode where, under matched settings, seeing the returned image causes the model to stop calling tools and answer directly. (ii) Sufficiency of the scaffold: on matched training queries, scaffold-only input yields equivalent accuracy to returned-image input. (iii) Component specificity: decomposing the scaffold into reasoning text and spatial code shows both components contribute, with the dominant one varying by task. Together these results support the Tool-Call Scaffold Hypothesis: in current thinking-with-images distributions, the active signal is the structured text emitted at tool-call time; the returned image is a redundant carrier. TextCall preserves accuracy while reducing latency by 29-46% and eliminating tool-execution API calls. Our claims hold for current thinking-with-images benchmarks; constructing tasks where pixels are genuinely load-bearing remains an open direction.
26. 【2608.09672】MPISuperRes-PnP: A Super-Resolution Zero-Shot Plug-and-Play Reconstruction Algorithm for Magnetic Particle Imaging
链接:https://arxiv.org/abs/2608.09672
作者:Vladyslav Gapyak,Thomas März,Andreas Weinmann
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:medical imaging modality, emerging medical imaging, Magnetic Particle Imaging, MPI, emerging medical
备注: 19 pages, 8 tables
点击查看摘要
Abstract:Magnetic Particle Imaging (MPI) is an emerging medical imaging modality. MPI is based on the non-linear response of magnetic nanoparticles to an applied magnetic field and avoids ionizing radiation. The measured signal is the voltage induced in receive coils by the particles' response. Reconstructing the particle concentration from the signal constitutes the imaging task. Even using state-of-the-art measurement-based reconstruction, the associated spatial grid is very coarse, hence super-resolution (SR) techniques are important. In this work, we propose an approach for SR in MPI inspired by energy minimization. Different methods have been proposed for SR in MPI, ranging from upscaling of the associated system matrix to interpolation of the reconstruction. Here we incorporate SR into the reconstruction task via an energy minimization formulation. Following the plug-and-play approach to energy minimization we derive a splitting scheme and a SR method for MPI where the arising Gaussian denoising task is treated with a pre-trained learned Gaussian denoiser in a zero-shot fashion. This way, we incorporate benefits of deep learning without training and avoid the need of training data. Further, we provide a quantitative and qualitative evaluation of the proposed method. Hyper-parameter are selected via an extended parameter search. The found parameters are applied for reconstruction on real data. We show the applicability of our method on synthetic and on real data (MPIData: EquilibriumModelWithAnisotropy and 2D-OpenMPI Data). The proposed method employs a deep-learning denoiser without training -- thus it does not require presently scarcely available MPI training data. The denoiser behaves conservatively, i.e., no hallucination artifacts were observed. The SR approach is generic such that it can be applied in future MPI contexts involving different regularizers or different imaging tasks.
27. 【2608.09669】CIFA: Contextual-Intersectional Fairness Auditing for Hidden Subgroup Discovery in Face Analysis
链接:https://arxiv.org/abs/2608.09669
作者:Nazia Aslam,Khalid Adnan Alsayed,Thomas B. Moeslund,Kamal Nasrollahi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:computer vision commonly, vision commonly relies, aggregate accuracy, computer vision, vision commonly
备注: 16 pages, 2 figures
点击查看摘要
Abstract:Fairness evaluation in computer vision commonly relies on aggregate accuracy and demographic subgroup analysis. However, visual models are also sensitive to contextual factors such as illumination, blur, image quality, facial accessories, and appearance attributes. These factors may interact with demographic characteristics, producing hidden subgroups in which performance degrades substantially despite strong aggregate accuracy and apparently acceptable demographic fairness. To address this, we propose the Contextual-Intersectional Fairness Auditing Framework (CIFA), a structured framework for identifying subgroup vulnerabilities arising from interactions between demographic and contextual attributes. CIFA performs demographic, contextual, and contextual-intersectional auditing, followed by worst-group discovery to identify and rank the most vulnerable attribute combinations. We evaluate CIFA on gender classification using ResNet-50 \cite{he2016deep} and ViT-B/16 \cite{dosovitskiy2020image} across FairFace \cite{Karkkainen2021}, CelebA \cite{Liu2015}, and UTKFace \cite{Zhang2017}. Our results show that aggregate accuracy and demographic-only evaluation can mask substantial contextual-intersectional disparities. We further assess several established mitigation strategies through an audit--mitigate--reaudit protocol and find that, although some worst-group disparities are reduced, no single strategy consistently eliminates them across datasets and architectures. These findings establish contextual-intersectional auditing as an important component of fairness evaluation and provide a reproducible framework for discovering, prioritizing, and reassessing hidden subgroup risks in face analysis systems.
28. 【2608.09658】Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells
链接:https://arxiv.org/abs/2608.09658
作者:Emma Takács,Mátyás Hajós,Ádám Juniki,Ádám Fischer,Zoltán Komáromi,Kristóf Abai,Dániel Horváth,Sándor Máthé,Konstantinos Kousias,Bence Tipary
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Networking and Internet Architecture (cs.NI)
关键词:low volume industrial, volume industrial scenarios, face workcell rearrangements, frequently face workcell, high mix
备注: Accepted at the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). A supplementary video demonstrating the workcell is available at [this https URL](https://youtu.be/zobin6oytGk)
点击查看摘要
Abstract:Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently face workcell rearrangements. Traditional setups are constrained by power and data cabling, restricting modularity and reconfigurations, while the selection of commercial wireless devices suitable for real-time perception and safe collaboration are limited in availability. This paper presents a highly flexible, wireless, 5G-based system that serves as a versatile experimental testbed for applications including remanufacturing, operator training, and user studies. To eliminate infrastructure barriers, the workcell integrates a novel battery-powered, multi-sensor platform prototype. Additionally, to support operator safety and system adaptability across environmental shifts, the system integrates a computer vision module for object detection and pose estimation, further augmented for robust hand recognition. Trained on synthetic and real data, the model reliably detects oriented grasping poses and human hands across varying lighting and background conditions (with an mAP@50-95 of 97.74 +- 0.10% and a mean inference time of 12.5 ms). Offloading these computationally intensive tasks to the edge via 5G, the proposed architecture contributes to resolving the bandwidth-latency trade-off. To demonstrate portability, the system was implemented in both Hungary and Norway, and was evaluated across a combination of public and private, Standalone and Non-Standalone 5G infrastructures. The performed network experiments produced results in round-trip response times down to 12 ms in case of compatible network-device pairings, suitable for safe, adaptive HRC. However, these measurements also revealed practical limitations related to interoperability in current 5G deployments that should be addressed in future works.
29. 【2608.09656】EgoHieraLoc: A Cortically Inspired Hierarchical Segmentation-Guided Framework for Egocentric Visual Query Localization
链接:https://arxiv.org/abs/2608.09656
作者:Yifei Cao,Guolong Wang,Mingliang Hou,Xiya Bu,Daming Liu,Yu Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:effectively guide fine-grained, Visual query localization, guide fine-grained localization, Visual query, aims to retrieve
备注: 60 pages, under review
点击查看摘要
Abstract:Visual query localization (VQL) aims to retrieve and re-localize a queried object in egocentric videos, yet remains challenging when object boundaries are ambiguous and global context cannot effectively guide fine-grained localization. Human vision handles such ambiguity through a hierarchical process: it rapidly screens foreground candidates, selectively attends to the target despite distractors, refines perception via feedback between global context and local detail, and, when a single view is unreliable, integrates evidence across viewpoints according to its credibility. Inspired by these competencies, we propose \textbf{EgoHieraLoc}, a unified framework for VQL-2D and VQL-3D. A Discriminative Parsing Module first extracts foreground-aware query representations using segmentation priors; a Query-Aware Module then performs robust target localization through discriminative correlation filtering with deformable modeling; and a Regional Adaptation Module feeds multi-scale context back into local regions to recover precise object boundaries. To extend this perceptual hierarchy to 3D localization, we introduce Geometric-Semantic Joint Confidence (GSJC), which multiplicatively couples segmentation confidence with local depth consistency, multi-view back-projection consistency, and triangulation-baseline quality, so that a viewpoint contributes to the 3D estimate only when it is credible both semantically and geometrically. Extensive experiments demonstrate state-of-the-art performance on both VQL-2D and -3D benchmarks.
30. 【2608.09637】DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation
链接:https://arxiv.org/abs/2608.09637
作者:Zian Li,Litong Gong,Borui Liao,Pengfei Liu,Xinyu Wang,Xinyuan Wei,Yifan Gao,Tiezheng Ge,Muhan Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:iterative sampling hinders, Diffusion models, enabled high-quality video, high-quality video generation, recent years
备注:
点击查看摘要
Abstract:Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation alleviates this cost, yet exposes a quality--diversity trade-off between its two dominant paradigms: trajectory-level distillation (e.g., sCM) favors diversity, whereas distribution-level distillation (e.g., DMD) favors quality. Targeting extreme two-step video generation, we introduce DUET, which reconciles the two paradigms through a noise-level duet of experts: an sCM expert takes the high-noise step to lay out diverse structure, and a DMD expert takes the low-noise step to refine appearance detail. Since the two experts are trained independently with their native objectives, DUET sidesteps the optimization difficulties of loss-level combinations and delivers quality and diversity jointly rather than trading one for the other. We further identify the relay interface and the high-noise stage as the remaining bottlenecks, and address them with RL-guided expert adaptation, yielding DUET+. With the Wan2.1-T2V-1.3B backbone, DUET lifts the two-step quality of sCM close to the level of DMD while retaining nearly all of its structural diversity---about twice that of DMD---and DUET+ further improves overall quality while preserving this diversity advantage. Together, these results establish noise-level expert specialization as a simple, effective paradigm for reconciling diversity and quality in two-step video generation.
31. 【2608.09636】NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
链接:https://arxiv.org/abs/2608.09636
作者:Haiyang Yan,Jinyue Guo,Yanchao Zhang,Bingqing Wang,Zhenchen Li,Jing Liu,Jiazheng Liu,Linlin Li,Hua Han
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:critical for neuroscience, fluorescence microscopy, microscopy is critical, neurons poses significant, Abstract
备注: Accept by ECCV 2026
点击查看摘要
Abstract:Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of neurons poses significant challenges to existing segmentation methods. These methods struggle to preserve both local details and global topology, leading to fragmented results. To address this, we propose NeuroRefiner, a multi-agent system that formalizes the human expert workflow involving iterative global observation and local editing. Specifically, NeuroRefiner comprises three collaborative agents dedicated to diagnosing topological errors, generating correction instructions, and validating refinement quality. To facilitate agent instruction-guided segmentation refinement, we propose TopoRefineNet, a dedicated 3D U-Net-based tool that leverages cross-modality feature fusion to generate refined masks. Through multi-round agent reasoning and voxel-level editing, NeuroRefiner produces topologically more accurate segmentations with enhanced interpretability. Experiments on the BigNeuron, CWMBS, and ZBFWB datasets demonstrate that NeuroRefiner outperforms state-of-the-art methods, notably achieving a 3.02% improvement in F1 score on the challenging ZBFWB dataset.
32. 【2608.09633】LoRA-based Adaptation Alone Is Not Enough: Understanding the Limits of Foundation Models for Face Presentation Attack Detection
链接:https://arxiv.org/abs/2608.09633
作者:Peter Lorenz,Anjith George,Marcel Sébastien
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Face presentation attack, presentation attack detection, Face presentation, presentation attack, presentation attacks
备注: accepted at ECCV 2026 Workshop on Foundation and Generative Models in Biometrics
点击查看摘要
Abstract:Face presentation attack detection (PAD) aims to reliably detect a wide range of presentation attacks. While PAD methods achieve strong performance within individual datasets, their performance degrades under cross-dataset evaluation. Variations in sensors or lighting conditions can reduce the effectiveness of detectors from near-perfect to nearly random. Foundation models (FMs) have emerged as a promising alternative because typical PAD datasets, such as the MCIO benchmarks (MSU-MFSD, CASIA-FASD, Replay-Attack, and OULU-NPU), are small relative to the scale used for web-based pretraining. However, existing PAD systems primarily focus on CLIP-based foundation models, while overlooking other FMs with different architectures and training procedures. This study addresses this question by systematically evaluating 32 FMs. Zero-shot prompting achieves performance near chance across model families and scales. The vision encoders, when low-rankadapted (LoRA) with fewer than 1% trainable weights, achieve below 2% intra-dataset ACER in most cases, while cross-dataset ACER is substantially higher. LoRA primarily refines the decision boundary within a dataset, suggesting that pretrained representations and the adaptation dataset play a larger role in cross-dataset generalization than the evaluated lightweight adaptation strategy.
33. 【2608.09613】Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking
链接:https://arxiv.org/abs/2608.09613
作者:Liying Yang,Hao Mo,Jialun Liu,Chen Liu,Xinxing Yu,Chenhao Guan,Hui Ma,Xiao Cao,Ajian Liu,Yanyan Liang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Ordinary Differential Equation, approaches typically rely, Existing unified, tracking approaches typically, lacking kinematic coherence
备注: Preliminary version
点击查看摘要
Abstract:Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacking kinematic coherence and failing to model dynamics at any arbitrary timestamp. In this paper, we propose Uni4R, a framework that unifies these tasks by learning continuous velocity fields through the synergy of Optimal Transport (OT) and Ordinary Differential Equation (ODE). Importantly, this continuous velocity field acts as a kinematic prior that mutually benefits both 4D reconstruction and point tracking. Specifically, we propose the Flow Matching Guided Decoder (FMGD). A global velocity branch first extracts anchor features that capture the global dynamic state of the sequence. Then, FMGD leverages Flow Matching (FM) theory to formulate a probability path defined by OT on the anchor feature manifold, instantiating it as FM-guided velocity features for velocity prediction. This establishes a robust kinematic inductive bias. Meanwhile, a point reconstruction branch provides geometric features. The local velocity prediction module then joint above features and time embeddings, to decode velocities at arbitrary timestamps. To overcome the absence of high-quality ground-truth velocities in fractional frames, we propose an integral-consistency training strategy. This strategy uses an ODE solver to integrate velocities to recover target pointmaps, enabling the model to be supervised end-to-end directly from integer timestamps. Experimental results demonstrate that Uni4R achieves SOTA performance in both 4D reconstruction and point tracking, and achieves SOTA in our new kinematics-aware benchmark at continuous time.
34. 【2608.09610】Structure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane Detection
链接:https://arxiv.org/abs/2608.09610
作者:Weize Cai,Yongqi Dong,Zhida Shao,Yichen Liu,Zixin Fu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Image and Video Processing (eess.IV); Signal Processing (eess.SP)
关键词:requires recovering thin, challenging driving conditions, detection requires recovering, Lane detection requires, frequently occluded lane
备注: 17 pages, 5 figures
点击查看摘要
Abstract:Lane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anchor-based detectors provide efficient candidate generation, their performance is limited by two coupled issues: backbone features often lose structural continuity along partially visible lanes, and classification confidence may decouple from line-level localization quality, allowing inaccurate anchors to persist before non-maximum suppression (NMS). We propose a structure-enhanced and quality-aware framework that improves lane representation and dynamic-anchor scoring while preserving the inference pipeline of the Anchor Decomposition Network (ADNet). Specifically, a Gated Horizontal-Vertical Token (GHVT) module enhances mid- and high-level backbone features via lightweight directional token interactions with a learnable residual gate. In parallel, Line-Quality-Aware Dynamic Anchor Scoring (LQAS) calibrates existing classification logits using quality supervision, hard-negative suppression, and pairwise ranking without adding inference branches. On the VIL-100 dataset, our method improves ADNet-R34 from 89.97 to 91.28 in F1 score at the 0.5 intersection-over-union threshold (F1@50), reducing both false positives and false negatives. Additional experiments on CULane and TuSimple datasets, extensive ablations, score-distribution diagnostics, and runtime analysis confirm complementary structural and ranking improvements with minimal computational overhead.
35. 【2608.09604】A Hybrid Neural-Microfacet BRDF Model for Real-Time Rendering
链接:https://arxiv.org/abs/2608.09604
作者:Louis De Oliveira,Anastasia Karpova,Georges Nader,Antoine Houdard,Pierre Mezieres,Damien Rioux-Lavoie,Romain Pacanowski
类目:Graphics (cs.GR); Computer Vision and Pattern Recognition (cs.CV)
关键词:real-time rendering pipelines, real-time rendering, past decade, formed the foundation, neural models
备注: 13 pages, 12 figures, conference, project page see [this https URL](https://ubisoft-laforge.github.io/world/hybridrdf)
点击查看摘要
Abstract:Over the past decade, microfacet-based BRDF models have formed the foundation of real-time rendering pipelines. Despite their widespread use, they often fail to reproduce subtle appearance effects arising from complex light-surface interactions, which have led to the emergence of specialized physics-based models for specific optical phenomena (e.g., diffraction, iridescence, multilayers). Although more accurate, these models lose versatility and lack performance for real-time rendering. Recently introduced, neural models have demonstrated their ability to approximate BRDF reference data coming from measurements, simulations, or even complex shading networks. However, most current neural models require relatively large networks, making them costly for real-time rendering. In this paper, we introduce a hybrid model that combines a GGX-type microfacet model and a neural model to leverage the best features of both representations. The neural component corrects the appearance approximated by the microfacet component, allowing much smaller network than in existing neural models. We show that, at identical memory cost, our model approximates measurements better than state-of-the-art neural models for a low evaluation overhead compared to a microfacet-based model. Furthermore, our hybrid model remains easily editable by artists and benefits from an important sampling scheme, making it attractive for both offline and real-time rendering.
36. 【2608.09597】ResemBrick: Brick Reconstruction from Photographs with Perceptual Fidelity and Buildability
链接:https://arxiv.org/abs/2608.09597
作者:Xilun Chen,Hanwen Wan,Yusong Zhao,Zexin Lin,Ruixiang Liao,Xiaoqiang Ji
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:meets hard physical-assembly, hard physical-assembly constraints, colored brick model, budget-limited voxel grid, Producing a hand-buildable
备注:
点击查看摘要
Abstract:Producing a hand-buildable, colored brick model of a 3D object from a few casual photographs is a clean testbed for a broader challenge: generating 3D content that meets hard physical-assembly constraints under a discrete, budget-limited voxel grid. On a coarse lattice, visual resemblance and structural stability pull against each other, yet prior brick pipelines address only one side and treat voxelization as fixed preprocessing rather than a variable to optimize. We present ResemBrick, which couples the two. Budgeted occupancy completion reframes discretization as allocation: given a target occupied-voxel count, a single resolution-conditioned network decides in one feed-forward pass which surface voxels to fill for best appearance, one weight set spanning 13 resolutions. Buildability by construction then combines support- and look-ahead-aware greedy placement with a deterministic, provably terminating repair that grounds every floating component. Under a matched budget, ResemBrick surpasses existing voxel selectors in perceptual fidelity while uniquely reaching zero floating and zero unstable bricks on unfiltered held-out objects; as a complete pipeline, it attains the best perceptual fidelity among prior brick-construction systems. Our results point to treating discretization and assembly as tightly coupled stages rather than independent ones.
37. 【2608.09594】Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation
链接:https://arxiv.org/abs/2608.09594
作者:Yifei Xue,Yuanchen Fei,Hao Zhang,Chenzhi Nie,Tie ji,Yizhen Lao
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:attracted considerable attention, AI-driven video generation, considerable attention, generation has attracted, attracted considerable
备注:
点击查看摘要
Abstract:Recently, AI-driven video generation has attracted considerable attention. This surge increases the demand for reliable video quality assessment (VQA) metrics to evaluate AI-generated content (AIGC) videos and guide model optimization. Existing studies assess video quality through visual harmony, video-text consistency, and domain-specific alignment, yet lack quantitative metrics for measuring fidelity to physical laws. To address this limitation, we present a novel benchmark that evaluates the quality of AIGC videos based on their compliance with physical principles by quantitatively measuring geometric consistency across frames extracted from generated sequences. This serves as a proxy for estimating the extent to which generated videos conform to real-world physical rules. Specifically, GeoCon-Bench captures global motion through translation estimation, fits homography or fundamental matrix models using background correspondences, and reports complementary metrics, including inlier ratio and geometric error. We also release a dataset containing 20 scenes across six motion categories. Experiments on state-of-the-art AIGC models demonstrate the reliability of GeoCon-Bench as a video quality assessment metric.
38. 【2608.09591】FactorDrive: Adaptive Multi-Step Reasoning Driven by Planning-Critical Factors for End-to-End Autonomous Driving
链接:https://arxiv.org/abs/2608.09591
作者:Guolei Huang,Tengfei She,Yuxuan Lu,Yao Huang,Yuqi Ye,Yongjun Shen
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:advanced scene understanding, Vision-language models, enabled explicit reasoning, advanced scene, scene understanding
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) have advanced scene understanding and enabled explicit reasoning in end-to-end autonomous driving. However, existing methods insufficiently integrate spatial-physical evidence into planning reasoning, while reasoning adaptation remains coarse-grained and falls short of scene-specific planning demands. Furthermore, reasoning-path optimization for higher planning quality remains largely unexplored in autonomous-driving post-training. To address these limitations, we propose FactorDrive, an end-to-end autonomous driving framework for adaptive multi-step reasoning driven by planning-critical factors (PCFs). We first perform large-scale driving-domain instruction tuning to establish foundational driving knowledge. Building on this foundation, we construct PCF-CoT, a chain-of-thought (CoT) dataset that grounds planning reasoning in trajectory-relevant spatial-physical evidence and organizes reasoning around scene-specific PCFs, enabling the composition and depth of reasoning paths to adapt to different planning demands. We further introduce Quality Search-Guided Group Relative Policy Optimization (QS-GRPO), which guides Monte Carlo Tree Search (MCTS) with trajectory-level planning rewards to discover reasoning paths with higher planning quality and uses the resulting responses to optimize the policy through GRPO, thereby improving trajectory planning performance. Extensive experiments on both open-loop (nuScenes) and closed-loop-oriented (NAVSIM) benchmarks demonstrate that FactorDrive achieves state-of-the-art planning performance.
39. 【2608.09590】aMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching
链接:https://arxiv.org/abs/2608.09590
作者:Chongjian Wang,Junjie Gao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Learning reliable correspondences, reliable correspondences, images and point, point clouds, clouds is fundamental
备注:
点击查看摘要
Abstract:Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods, existing approaches primarily optimize matching within a single model and often struggle to maintain reliable correspondences under challenging conditions such as noisy inputs, low overlap, and ambiguous structures. In this work, we propose TeaMatch, a novel framework that introduces teachability as a criterion for cross-modal representation learning. We define teachability as the ability of a representation to be effectively recovered by weak learners under degraded inputs, reflecting its structural consistency and robustness. To this end, we construct a set of task-specific weak students that simulate common failure modes and train them to imitate the teacher on a training split while evaluating their recoverability on a disjoint meta split. The teacher is then optimized to improve the students' ability to recover reliable correspondences, guided by correspondence-level and geometry-aware constraints. Our framework can be seamlessly integrated into existing coarse-to-fine matching pipelines without additional inference cost. Extensive experiments demonstrate that TeaMatch improves matching robustness and achieves state-of-the-art performance on challenging 2D-3D matching benchmarks.
40. 【2608.09581】GenTrack3: Hybrid Stochastic-Deterministic Online Multi-Object Tracking with Cluster-Aware Association
链接:https://arxiv.org/abs/2608.09581
作者:Toan Van Nguyen,Rasmus G. K. Christiansen,Dirk Kraft,Leon Bodenhagen
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:involves maintaining consistent, objects dynamically enter, consistent target identities, maintaining consistent target, involves maintaining
备注: The content of this paper was included in the full manuscript of GenTrack family which has been submitted to the journal for possible publication
点击查看摘要
Abstract:Multi-object tracking (MOT) involves maintaining consistent target identities as objects dynamically enter and leave a scene. Deterministic approaches, such as tracking-by-detection with data association, produce reproducible results and are computationally efficient, but they rely heavily on motion models and are sensitive to noisy detections that can lead to association errors. In contrast, stochastic methods explicitly model uncertainty and can better handle complex non-linear dynamics, albeit at the cost of increased computational complexity and variability arising from random sampling. This paper presents an online MOT framework that integrates deterministic and stochastic principles to achieve robust tracking under uncertainty. Furthermore, a novel track-to-detection matching approach is introduced to enhance scalability with increasing target numbers while supporting group tracking. The tracking inference mechanism employs a tracklet that includes identifiers, states, velocities, track penalties and track ages of targets, supporting a systematic tracking pipeline. Each target is associated with a stochastic particle set to compute the matching cost to detections. Reference implementations of the proposed approach and baseline trackers can be found on GitHub: this https URL.
41. 【2608.09579】You Only Flow Once: Calibrated and Real-Time Radar Pose Estimation with Multi-Hypothesis Normalizing Flows
链接:https://arxiv.org/abs/2608.09579
作者:Jonas Leo Mueller,Sebastian Hoefler,Dario Zanca,Naga Venkata Sai Jitin Jami,Thomas Altstidl,Bjoern M. Eskofier
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:multiple plausible human, Sparse and noisy, estimation fundamentally ill-posed, plausible human poses, noisy millimeter-wave radar
备注: Accepted at the Winter Conference on Applications of Computer Vision (WACV) 2027
点击查看摘要
Abstract:Sparse and noisy millimeter-wave radar point cloud observations often correspond to multiple plausible human poses, making deterministic pose estimation fundamentally ill-posed. Yet existing radar methods remain deterministic, collapsing this ambiguity into a single estimate. Diffusion-based alternatives can model multi-hypothesis distributions but require costly sequential denoising for each distribution sample and lack calibrated uncertainty. We propose Multi-Hypothesis Normalizing Flow Pose Generator (MH-NFPG), which models pose distributions from radar point clouds using a conditional normalizing flow. Specifically, we combine a spatiotemporal transformer backbone with a normalizing flow that transforms a Laplace base distribution into an expressive posterior, generated in parallel through a single forward pass. Leveraging this efficiency, we outperform diffusion-based alternatives in calibration across three radar benchmarks (MM-Fi, mmRadPose, mRI), improve pose accuracy on two, and match it on the third, while achieving over 20x faster inference for applications and reducing calibration error by up to 85%. We find that calibration degrades substantially for diffusion models, whereas our flow-based approach maintains reliable coverage, also in cross-environment settings. These results demonstrate normalizing flows as a practical alternative to diffusion models for real-time, uncertainty-aware radar pose estimation. Our code will be made publicly available.
42. 【2608.09575】MSP-Net: Manifold-Guided Spectral Prompt Network for Hyperspectral Object Tracking
链接:https://arxiv.org/abs/2608.09575
作者:Juliu Li,Hanlin Qin,Shuowen Yang,Jingjing Li,Yuedong Tan,Shuai Yuan,Huixin Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:provide unique advantages, leverages abundant spectral, abundant spectral information, Hyperspectral object tracking, object tracking leverages
备注:
点击查看摘要
Abstract:Hyperspectral object tracking leverages abundant spectral information to provide unique advantages for target discrimination in complex scenes. However, existing methods typically treat hyperspectral images as multi-channel extensions of RGB images, performing feature fusion in fixed band order. This approach leads to models dependent on specific sensor configurations while neglecting manifold relationships between bands, making generalization to heterogeneous sensors difficult. Moreover, the discriminative contribution of bands dynamically changes with target attributes and scene variations, further limiting the representational capacity of static fusion strategies. To address this, we propose the Manifold-Guided Spectral Prompt Network (MSP-Net). This network first reconstructs band relationships and forms adaptive spectral grouping through graph-driven manifold routing, then jointly integrates grouped spectral statistics with template appearance to construct target-related dynamic conditional prompts, enhancing target features while suppressing background interference. Furthermore, as tracking progresses, spectral conditions continuously evolve based on intermediate target representations, enabling target prompts to adapt in real-time to appearance and scene changes. Meanwhile, reliable historical states are used to constrain target localization and scale fluctuations, significantly improving temporal stability in cross-sensor tracking. Experiments on HOT2020 and HOT2023 demonstrate that MSP-Net achieves AUC and Precision exceeding 0.80 and 0.96, respectively, exhibiting exceptional robustness under heterogeneous sensors, target deformation, and complex background conditions. The code will be released at this https URL.
43. 【2608.09573】VideoVIBE: A Video-Grounded Diagnostic Benchmark for One-Shot Interactive Website Generation
链接:https://arxiv.org/abs/2608.09573
作者:Jiajun Xu,Yanghao Zhou,Jingyun Liao,Yu Bai,Jinxing Zhou,Chengliang Liu,Changsen Yuan,Bo Wang,Qian Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:interactive web applications, vibe coding, enables the one-shot, one-shot generation, generation of visually
备注:
点击查看摘要
Abstract:Natural-language-driven "vibe coding" enables the one-shot generation of visually rich and interactive web applications, yet reliable assessment of their quality has not kept pace. Existing evaluations often score isolated artifacts or final task outcomes, offering limited evidence about which failures occur and why. We introduce VideoVIBE, a video-grounded benchmark that transforms human-operated webpage recordings into fine-grained diagnostic tasks. It contains approximately 1.7K diagnostic Video QA instances derived from 6,338 verified failures across generated webpages, spanning semantic-logical, visual-motion, structural-temporal, and functional failures. Diagnoses are grounded primarily in recorded presentation and behavior, with webpage source code used as complementary context. We further propose V2Lens, a training-free, evidence-grounded multi-agent system that challenges and selectively refines initial video-based diagnoses through targeted visual and source-code verification. Across thirteen closed-source and open-weight Video MLLMs, Gemini-2.5-Flash is the strongest standalone model with a score of 64.54, while V2Lens reaches 71.72, an improvement of 7.18 points. Together, our results show that video-grounded evaluation can move beyond isolated artifacts and aggregate outcomes toward a behaviorally faithful and diagnostically informative account of generated application quality.
44. 【2608.09564】From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation
链接:https://arxiv.org/abs/2608.09564
作者:Zeyuan Ma,Jiaxin Chen,Di Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:UAV vision-language navigation, follow natural-language instructions, UAV vision-language, egocentric visual observations, vision-language navigation
备注: 10 pages, 5 figures. Accepted at ACM Multimedia 2026 (MM '26)
点击查看摘要
Abstract:UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environments from egocentric visual observations. Current approaches suffer from three coupled issues: weak grounding of instruction-relevant landmarks in visual observations, insufficient exploitation of long-horizon history, and unstable decisions under local traps or repeated exploration. To address these issues, we propose a unified semantic-to-decision framework. First, we present an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state. Subsequently, we develop a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frames into structured landmark prompts for the decoder. Finally, we devise a topology-aware decision method that combines local-optimum cognition with group-relative policy optimization under progress, goal, semantic, and path-compliance rewards. Experiments on the widely used AerialVLN and OpenFly benchmarks clearly demonstrate that our method achieves state-of-the-art performance.
45. 【2608.09550】PressureMesh: 3D Human Mesh Estimation from Multi-Device Pressure Images
链接:https://arxiv.org/abs/2608.09550
作者:Changhai Ma,Ziyu Wu,Yunkang Zhang,Fangting Xie,Mengting Niu,Heyu Ding,Quan Wan,Jiayue Yuan,Boyan Liu,Yi Ke,Xiaohui Cai
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:human-computer interaction, Human pose monitoring, crucial in fields, rehabilitation assessment, assessment and human-computer
备注:
点击查看摘要
Abstract:Human pose monitoring is crucial in fields such as rehabilitation assessment and human-computer interaction. Due to its privacy-preserving nature, pressure-based human pose monitoring has become a primary approach for unobtrusive sensing. However, existing methods are generally limited to a single device, which restricts the effective monitoring range. To address this limitation, we propose MDP-Net, an end-to-end network capable of directly estimating human meshes from temporal pressure data across multiple devices. We introduce a multimodal fusion mechanism inspired by the Mixture of Experts (MoE) framework to achieve effective complementarity and enhancement of cross-device pressure information. To support the training and evaluation of MDP-Net, we constructed MDP, a high-quality multi-device temporal pressure dataset that includes various pose labels such as 2D/3D joints and human meshes. Experimental results demonstrate that MDP-Net achieves a joint position error of 12.6 cm on the MDP dataset. These results prove that fusing multi-device pressure information is an effective and promising new solution for daily human pose monitoring.
46. 【2608.09541】owards Collaborative Joint Perception and Prediction: Framework, Baseline Evaluation, and Deployment Perspectives
链接:https://arxiv.org/abs/2608.09541
作者:Lei Wan,Hannan Ejaz Keen,Alexey Vinel
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Connected Autonomous Vehicles, Connected Autonomous, Autonomous Vehicles, multi-source sensor information, exchange multi-source sensor
备注: 28 pages, 4 figures, post-publication of conference paper, accepted at journal SN Computer Science
点击查看摘要
Abstract:Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-PP), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions. We present a conceptual framework for Collaborative Joint Perception and Prediction (Co-PP) that improves motion prediction of surrounding road users, thereby enhancing situational awareness in complex and dynamic traffic environments. Building upon our preliminary study, this extended version compares the performance of different fusion strategies and establishes baseline performance for a modular design of perception and prediction. Experimental results show that prediction-level fusion leads to a decline in overall system performance compared to detection-level or tracking-level fusion. We further implement a minimal end-to-end Co-PP prototype that couples collaborative point-cloud sharing via the RENO neural codec with joint detection-forecasting via FutureDet, showing that collaboration improves forecasting accuracy while neural compression preserves this benefit at roughly 34x lower communication bandwidth.
47. 【2608.09536】DocPure: Prompt-Free Unified Document Restoration via Degradation-Aware Structure-Guided Wavelet Modulation
链接:https://arxiv.org/abs/2608.09536
作者:Lingming Su(1),Wanglong Lu(1 and 2),Tao Wang(3),Kaihao Zhang(4),Nan Zhang(1),Liyan An(1),Hanli Zhao(1) ((1) College of Computer Science and Artificial Intelligence, Wenzhou University, Wenzhou, China, (2) AI Analytics Team, Nasdaq, St. John's, Canada, (3) vivo Mobile Communication Co., Ltd, Shanghai, China, (4) College of Engineering and Computer Science, The Australian National University, Canberra, Australia)
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:downstream automatic processing, High-quality document images, High-quality document, automatic processing, pivotal for information
备注: 19 pages, 15 figures. Lingming Su and Wanglong Lu contributed equally to this work
点击查看摘要
Abstract:High-quality document images are pivotal for information archiving and downstream automatic processing. However, they are frequently compromised by diverse degradations during uncontrolled acquisition and transmission. While unified document restoration techniques have been proposed to restore images from multiple degradations, they often struggle with training multiple degradation-specific models, reliance on manual task-specific prompts, or cross-task data pairing. To address these limitations, we propose DocPure, a prompt-free unified framework that achieves degradation-aware document restoration. We design a degradation-aware structure auto-encoder with degradation-informed routing regularization to predict clean structural priors from degraded inputs. The model is prompt-free at inference, and degradation labels are only used as auxiliary supervision for the routing regularization during training. Furthermore, we introduce a structure-guided wavelet interaction mechanism to bridge frequency-domain features and spatial semantics. Within the structure-guided wavelet interaction mechanism, a cross-frequency adaptive modulation utilizes low-frequency sub-bands to modulate high-frequency recovery, ensuring structural consistency. Extensive experiments demonstrate that DocPure achieves strong performance compared with state-of-the-art methods across various tasks, including deblurring, denoising, compression artifact reduction, and deshadowing.
Comments:
19 pages, 15 figures. Lingming Su and Wanglong Lu contributed equally to this work
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2608.09536 [cs.CV]
(or
arXiv:2608.09536v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.09536
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Related DOI:
https://doi.org/10.1109/TCSVT.2026.3720433
Focus to learn more
DOI(s) linking to related resources</p>
48. 【2608.09529】owards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
链接:https://arxiv.org/abs/2608.09529
作者:Dongxu Ge,Shansong Liu,Cheng Gong,Xiao-Lei Zhang,Chi Zhang,Xuelong Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:synthesizing static visual, static visual content, attracted increasing research, increasing research attention, synthesizing static
备注: 23 pages, 16 figures
点击查看摘要
Abstract:As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted increasing research attention in recent years. Nevertheless, despite the remarkable visual quality of modern text-to-image (T2I) models, the performance of A2I remains fundamentally limited by traditional datasets, which often lack both high-fidelity images and precise cross-modal alignment. As a result, existing methods still struggle to achieve high-quality audio-to-image generation through finetuning strong T2I models, thereby constraining practical applications in this area. Motivated by this gap, we introduce A2I-Set, a unified, high-quality tri-modal dataset consisting of 323K paired audio, images, and detailed text captions, specifically designed for audio-visual research, including audio-conditioned image generation. Besides, we developed a new mixed-source test set for the A2I task through human supervision. We further propose an A2I model, AudioCanvas, fine-tuned on our A2I-Set. Experiments show that AudioCanvas achieves more visually expressive as well as cross-modal alignment results that generally outperforming existing approaches. Our dataset and source code are available at this https URL.
49. 【2608.09522】riView-YOLO: Early Multi-View Fusion for Ground Penetrating Radar Cavity Detection in Soft, High-Water-Content Soils
链接:https://arxiv.org/abs/2608.09522
作者:Suphawut Thawinutchokaudom,Sompote Youwai,Warat Kongkitkul,Mitsumasa Yamashina,Jose M.D.S. Rodrigues Neto,Jun Shinohara
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Ground Penetrating Radar, water-saturated soil attenuates, Penetrating Radar, degrades cavity reflections, Automated detection
备注:
点击查看摘要
Abstract:Automated detection of subsurface cavities from Ground Penetrating Radar (GPR) is most difficult in soft, high-water-content ground, where conductive, water-saturated soil attenuates the signal and degrades cavity reflections, yet this is also the condition under which cavities most readily form. This paper proposes TriView-YOLO, a multi-view YOLOv12 detector for road cavity screening in such ground. Three co-registered views (longitudinal B-scan, horizontal C-scan, and cross-section B-scan) form a 9-channel input fused by a TripleInputConv layer that replaces the YOLOv12 stem; the rest of the network is unchanged, and bounding boxes are required on the longitudinal view only. Training used 1,600 expert-verified field samples, principally metropolitan road surveys of Bangkok, Thailand, acquired with a vehicle-mounted multichannel three-dimensional GPR mobile mapping system, with surveys over the firmer subgrades of Japan added to training and validation only. The test set comes exclusively from the Bangkok surveys, over soft marine clay with 80-140% water content and a water table at 1-2 m depth, a ground condition for which no dedicated deep learning cavity-detection evaluation has been reported. On this unaugmented, field-only test set, split randomly within surveys, the proposed model attains mAP50 of 0.558 +/- 0.028 over three seeds at 23.6 GFLOPs and 3.1 ms per image. Ablations show that removing the auxiliary views lowers mAP50 and recall, whereas public and synthetic training images, DINOv3 features, larger model scale, and COCO pretraining bring no gain.
50. 【2608.09520】A Height-Constrained 2-Point Minimal Solver for Pose Estimation from Active LED Markers with Event Cameras
链接:https://arxiv.org/abs/2608.09520
作者:Runze Yuan,Alexander Kappler,Jun Zhang,Kuangyi Chen,Fabio Morbidi,Pascal Vasseur,Cédric Demonceaux,Friedrich Fraundorfer
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:autonomous applications requiring, computationally demanding feature-based, applications requiring real-time, requiring real-time localization, demanding feature-based methods
备注: 8 pages, 6 figures, accepted by IEEE/RSJ International Conference on INTELLIGENT ROBOTS SYSTEMS (IROS) 2026
点击查看摘要
Abstract:In many autonomous applications requiring real-time localization, active marker-based systems are preferred due to their low latency and ease of deployment compared to computationally demanding feature-based methods. Event~\mbox{cameras} offer high temporal resolution and minimal delay and are commonly used with active LED markers for robust real-time localization. Existing methods typically rely on Perspective-n-Point (PnP) solvers for pose estimation. However, structured marker layouts can be challenging to deploy in space-constrained scenarios, while partial self-motion information (e.g., gravity direction and altitude) is readily available from onboard sensors. We derive a robust and accurate minimal solver that estimates camera pose from only two LED markers by incorporating known tilt angle and camera height measured by an onboard sensor, such as an IMU or an altimeter. The proposed formulation uniquely determines the camera pose through both a closed-form and a linear least-squares solution. We further analyze degenerate configurations and characterize the conditions under which height information does not contribute to rotation estimation. For evaluation, we developed an event-based active marker system to collect real-world data with ground truth from a motion capture system. Experiments on both synthetic and real data demonstrate improved accuracy over the state-of-the-art P2P solver and competitive performance relative to P3P.
51. 【2608.09519】XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher
链接:https://arxiv.org/abs/2608.09519
作者:Lazar Đoković,Aimee Lin
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:lightweight local feature, local feature extractor, resource-constrained hardware, present a reproducibility, reproducibility study
备注: 21 pages, 6 figures. Published in Transactions on Machine Learning Research (TMLR); Reproducibility Certification
点击查看摘要
Abstract:We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the architecture based on the paper and supplementary material, re-evaluate the authors' released checkpoint alongside our re-implementation, and conduct additional architectural ablations to examine design choices that were not fully justified in the original work. This distinction between re-evaluation and reproduction is important, as the paper, supplement, and public code differ in several implementation details, including the backbone layout, fusion block, and training losses. Empirically, our reproduced models closely match and, in some cases, outperform the re-evaluated original checkpoint on MegaDepth-1500 and ScanNet-1500, supporting the main claim that XFeat provides a strong accuracy-efficiency trade-off for standard image-matching benchmarks. Our ablations provide a more nuanced view of two architectural arguments from the original paper. In particular, the parallel keypoint branch is important for semi-dense matching, but its benefit is less pronounced than originally claimed, while the evidence for the specific placement of the single skip-connection remains inconclusive. Finally, we reproduce the original downstream evaluations and find close agreement for homography estimation, while Aachen visual localization remains below the reported results, even for the released checkpoint, suggesting sensitivity to underspecified evaluation details. We then extend the analysis to zero-shot out-of-distribution and cross-modal matching across retinal, thermal-visible, and multimodal remote-sensing imagery, where XFeat remains effective in some settings but degrades sharply under severe modality shifts.
52. 【2608.09512】Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
链接:https://arxiv.org/abs/2608.09512
作者:Karim Zaghw,Andrew Pashea,Marc Pritsch,Wouter Nuijten,Karl Friston,Lancelot Da Costa
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Active inference offers, scaling discrete active-inference, discrete active-inference models, temporal domains remains, domains remains difficult
备注: 25 pages, 1 figure. Accepted as a full paper at the 7th International Workshop on Active Inference (IWAI 2026). Supplementary material: [this https URL](https://doi.org/10.5281/zenodo.20533539) . Code: [this https URL](https://github.com/apashea/RGMs)
点击查看摘要
Abstract:Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spatial and temporal domains remains difficult. Renormalising generative models (RGMs) address this challenge by composing discrete generative models across spatial and temporal scales, coarse-graining lower-level states and paths into higher-level causes for objects, events, and action. However, fully reproducing and adapting the framework remains difficult: the mathematical exposition is compact, and the reference implementations are deeply integrated within specialized software environments, leaving many algorithmic details implicit. This paper addresses these challenges by providing a self-contained, derivation-oriented account of RGMs together with an open, verified implementation. We explain how the hierarchy is built, how beliefs and actions are updated within it, and how information is passed between levels. Where the published equations and implementation differ in emphasis, we make those choices explicit and explain their modelling consequences. By clarifying the theory and separating it from its original implementation context, this work lowers practical barriers to entry and makes RGMs more transparent, auditable, and reproducible, providing a foundation for future quantitative evaluation and development on machine-learning benchmarks.
53. 【2608.09497】SwissCrop25: A National Multi-Year Benchmark for Operational Crop Mapping
链接:https://arxiv.org/abs/2608.09497
作者:Thomas Lauber,Mehmet Ozgur Turkoglu,Sélène Ledain,Helge Aasen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:crop mapping, generalise across years, surrounding landscapes, crop mapping requires, resolve fine-grained crop
备注: Accepted at the ECCV 2026 Workshop TerraBytes II. To appear in the workshop proceedings
点击查看摘要
Abstract:Operational crop mapping requires models that generalise across years, resolve fine-grained crop taxonomies, and distinguish cropland from surrounding landscapes. However, existing crop mapping datasets enable evaluation of these requirements only in isolation. We therefore introduce SwissCrop25, a national-scale crop mapping benchmark dataset spanning seven growing seasons (2019-2025). SwissCrop25 combines Sentinel-2 time series, daily temperature observations, a fine-grained 73 crop taxonomy including grassland management types, and 5 explicit non-crop land cover classes. To evaluate realistic deployment conditions, we define a leave-one-year-out protocol with joint cropland delineation and crop classification for benchmarking representative crop mapping architectures. Evaluating U-TAE (convolutional temporal-attention model), TSViT (transformer-based spatio-temporal model), and Galileo (EO foundation model) reveals differences between architectures hidden by conventional benchmarks. In this setting, domain-specific models outperform Galileo, with TSViT achieving the best overall performance and a 12 pp macro-mIoU advantage over U-TAE. SwissCrop25 also exposes substantial interannual distribution shifts and shows that incorporating temperature-derived phenological information improves robustness. Finally, in-season evaluation reveals a trade-off between models, with U-TAE performing better early in the season and TSViT gaining an advantage later through improved rare-class discrimination. SwissCrop25 provides a challenging testbed for evaluating crop mapping systems under realistic operational conditions and is publicly released at this https URL .
54. 【2608.09493】GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction
链接:https://arxiv.org/abs/2608.09493
作者:Khang Minh Le,Hieu Dinh Trung Pham,Luu Thanh Danh,Nam-Tien Le,Hieu Anh Ngo,Phuong Huu Vu Tran,Son Nguyen Minh Le,Nguyen Trong Nghia,Tu Tran Thi Cam,Huy Minh Nhat Nguyen,Cuong Tuan Nguyen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Long-horizon future-frame prediction, inconsistent object motion, remains challenging due, Long-horizon future-frame, intelligent transportation systems
备注: accepted to the ECCV 2026 AI City Challenge Workshop
点击查看摘要
Abstract:Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.
55. 【2608.09482】Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance
链接:https://arxiv.org/abs/2608.09482
作者:Chunxiao Liu,Wei Liu,Anbin Xiong,Erli Meng
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:unified low-level vision, effectively recover high-quality, recover high-quality images, low-level vision task, unified low-level
备注: Accepted by ACMMM2026 as Oral Paper
点击查看摘要
Abstract:All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded by various types and levels of corruption using a single model. Recent works have achieved remarkable progress by learning degradation-adaptive prompts or network architectures. However, these methods typically apply a uniform restoration strategy across the entire image, neglecting the fact that different regions may suffer from distinct degradation types and varying degrees of severity. In contrast, we propose to perform restoration at the pixel level, thereby enabling more fine-grained and precise control over the restoration process. Specifically, we present MGN-AIR, a novel pixel-level restoration framework for all-in-one image restoration. Our approach first learns to estimate a pixel-level visual prompt. Then, it leverages both textual and visual prompts to provide global and local degradation cues, guiding the model on where to look and how to restore at each pixel. We conduct extensive experiments on multiple all-in-one image restoration benchmarks, covering a wide range of tasks including denoising, deraining, deblurring, dehazing, desnowing, and low-light enhancement. Experimental results demonstrate that our proposed method consistently and significantly outperforms existing approaches.
56. 【2608.09475】Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge
链接:https://arxiv.org/abs/2608.09475
作者:Yiwen Ren,Jianing Liu,Yingxin Wang,Kexin Zhang,Licheng Jiao,Lingling Li,Xu Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:spoken motion expression, return empty masks, segment the objects, spoken motion, motion expression
备注:
点击查看摘要
Abstract:The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent. We present a simple staged solution. Qwen3-ASR first converts speech into text. Several video mask tracks are then produced with complementary grounding and segmentation models. Instead of trusting a single prediction, we select the track that has the highest average mask agreement with the other candidates. A small set of explicit direction, count, and plural rules corrects queries that require more than ordinary single-object tracking. Finally, a video-level classifier combines visual, audio-visual, and within-video query scores to decide whether any target is present. The submitted system obtains 0.5952 J F, 0.7931 no-target accuracy, 0.9205 target accuracy, and a final score of 0.769589. The challenge organizers notified our team that this result ranked first in the track.
57. 【2608.09474】FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search
链接:https://arxiv.org/abs/2608.09474
作者:Hieu Dinh Trung Pham,Phuong Huu Vu Tran,Thuan Duc Mai,Son Nguyen Minh Le,Khang Le Minh,Hoang Vo,Minh-Chi Phung,Huy Minh Nhat Nguyen,Cuong Tuan Nguyen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:requires retrieving real-world, retrieving real-world pedestrian, real-world pedestrian images, detailed natural-language descriptions, models trained primarily
备注: accepted to the ECCV 2026 AI City Challenge Workshop
点击查看摘要
Abstract:Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained primarily on synthetic data. This Sim2Real setting is particularly challenging because visually similar candidates may differ only in subtle actions, object interactions, or appearance attributes, while applying multimodal large language models to the entire gallery is computationally expensive. We propose an anchor-constrained coarse-to-fine retrieval framework that combines global semantic matching with fine-grained verification. First, each query is represented by its original caption, a structured concatenation, and several semantic facets. Heterogeneous vision-language retrievers are then integrated through robust per-query score calibration and soft claim-aware fusion. Full and concatenated captions serve as anchors to preserve candidate recall, whereas appearance, action, and object facets provide bounded corrective evidence. The resulting candidate pool is further refined by a discriminative Qwen3 reranker and two complementary semantic verification modules based on anomaly-aware cloze completion and multi-agent evidence reasoning. Finally, an uncertainty-gated consensus module adaptively reweights the three experts on ambiguous queries. Experiments on the PAB benchmark show that the proposed soft claim-aware retrieval achieves 86.44% mAP@10, substantially outperforming individual retrieval backbones. The complete framework further improves performance to 95.41% mAP@10, 94.44% R@1, and 99.09% R@5. These results demonstrate that preserving strong global retrieval while restricting expensive semantic reasoning to a small candidate pool is effective for fine-grained Sim2Real person anomaly search. Our code will be available on Github.
58. 【2608.09467】RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation
链接:https://arxiv.org/abs/2608.09467
作者:Boxiong Wang,Hui Kang,Geng Sun,Jiahui Li,Chao Yu,Daxin Tian
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Unmanned aerial vehicle, vehicle vision-language navigation, aerial vehicle vision-language, translate visual observations, Unmanned aerial
备注:
点击查看摘要
Abstract:Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions into reliable flight actions in complex environments. Although recent end-to-end UAV vision-language-action (UAV-VLA) policies reduce reliance on separately designed perception, planning, and control modules, their behavior-cloning objectives provide limited corrective supervision for interactive closed-loop execution. Reinforcement learning (RL) offers a promising solution, while its effectiveness is constrained by inefficient use of samples, long-tailed scene distributions, and policy distribution shift during optimization. To this end, we propose RecoverFly, a failure-aware RL post-training framework for end-to-end UAV-VLA policies. Specifically, RecoverFly adapts token-level RL for stable optimization of grammar-constrained autoregressive UAV actions, revisits unresolved failure cases to strengthen corrective learning and sample utilization, and combines a two-stage long-tail scene curriculum with reference-policy regularization to improve scene adaptation while preserving acquired capabilities. Experiments on the TravelUAV benchmark demonstrate that RecoverFly achieves the best performance on the seen, unseen-map, and unseen-object splits. Moreover, compared to the AerialVLA initialization, RecoverFly improves success rate by 3.12 to 8.37 percentage points under a total rollout budget of about 30\% of the training-set size, validating its effectiveness, robustness, and generalization capabilities.
59. 【2608.09460】Flow-based conditional cardiac anatomy generation for virtual cohorts
链接:https://arxiv.org/abs/2608.09460
作者:Konstantinos Kevopoulos,Beatrice Moscoloni,Benjamin Alheit,Cameron Beeche,Julio A. Chirinos,Alexander Heinlein,Mathias Peirlinck
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Quantitative Methods (q-bio.QM); Tissues and Organs (q-bio.TO)
关键词:digital twin research, represent clinically relevant, clinically relevant population, relevant population subgroups, subject-specific anatomical replicas
备注:
点击查看摘要
Abstract:Cardiac digital twin research is moving from subject-specific anatomical replicas toward virtual cohorts that represent clinically relevant population subgroups. Yet access to representative imaging-derived anatomy datasets remains limited by cohort size, subgroup sparsity, and data-sharing constraints. Conditional generative models could help address this gap, but virtual cohorts are useful only if they preserve realistic, metadata-dependent anatomical variability. Existing cardiac anatomy generators largely rely on conditional variational autoencoders (cVAEs), which couple representation learning and metadata conditioning through a shared regularized latent prior. We introduce CAN-FLOW, a two-step Conditional ANatomy generation framework based on normalizing FLOWs that first learns geometry-only latent representations of diffeomorphic cardiac shape momenta and then models their sex-, age-, and body-mass-index-dependent distribution with a conditional normalizing flow. We trained CAN-FLOW on 2,208 healthy UK Biobank subjects and compared it with cVAEs across regularization strengths. CAN-FLOW generated plausible stochastic biventricular anatomies that better reproduced clinical phenotype distributions, metadata-dependent trends, subgroup variability, point-cloud coverage, and high-dimensional shape variability. Together, these results establish CAN-FLOW as a shareable framework for generating realistic, stochastically varying, metadata-conditioned biventricular anatomies for virtual cohort construction and in silico clinical trial workflows.
60. 【2608.09452】A Content-Aware Pure Permutation with Intrinsic Avalanche Effect: Breaking the Diffusion-Permutation Dichotomy
链接:https://arxiv.org/abs/2608.09452
作者:Zahra Ghoraeian,Mohammad-Reza Sadeghi,Samaneh Mashhadi
类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
关键词:data hiding, fundamental tool, differential sensitivity, Pixel, TCA
备注: 12 pages, 5 figures
点击查看摘要
Abstract:Pixel permutation is a fundamental tool in image processing, image encryption, and data hiding (including watermarking and steganography) that rearranges pixels without changing their values. A common assumption in the literature is that permutation alone cannot create differential sensitivity; changing one pixel merely relocates that pixel in the output, producing no avalanche effect. This paper challenges this by introducing the Triangular Content-Aware Permutation (TCA) algorithm. The method extracts edge points using Canny and applies Delaunay Triangulation to edges and corners, creating a unique partition. Since triangulation is highly sensitive to image geometry, changing a single pixel alters the edge map, resulting in a completely different triangulation and global permutation pattern. Unlike classical dimension-based permutations and advanced content-aware methods (2025-2026), which lack differential sensitivity, TCA increases NPCR from near-zero to 97.10% solely through pixel relocation. Experiments on 50 images show that TCA, with an average of 14.81 iterations, achieves NPCR = 97.10% and UACI = 20.06%, proving pure permutation can create significant differential sensitivity. Conventional methods maintain near-zero NPCR. The iteration threshold varies from 6.4 to 30.7 based on content complexity. Low PSNR (11.93 dB) and near-zero correlation (~10^-3) confirm superior statistical performance. Although slower than classical methods due to triangulation, this is a deliberate trade-off for stronger security. Given the non-analytic, content-dependent nature of the pattern, TCA is ideal for reference-based encryption, fragile watermarking, and non-blind steganography.
Comments:
12 pages, 5 figures
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
MSC classes:
68U10, 94A60
ACMclasses:
I.4.9; E.3; K.6.5
Cite as:
arXiv:2608.09452 [cs.CV]
(or
arXiv:2608.09452v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.09452
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
61. 【2608.09449】Sekai2: From World Exploration to Interactive World Modeling
链接:https://arxiv.org/abs/2608.09449
作者:Kang He,Wenshuo Peng,Zihui Gao,Jiaming Tan,Kaipeng Zhang,Yongtao Ge
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Video world models, Video, camera, long videos paired, scenes evolve
备注: Sekai2 dataset technical report. Developed at Alaya Lab
点击查看摘要
Abstract:Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.
62. 【2608.09448】VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction
链接:https://arxiv.org/abs/2608.09448
作者:Hongjin Ji,Guoyang Xia,Luoyang Sun,Fangxiang Feng,Lei Ren
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:unlabeled deployment streams, Test-time training, offers a lightweight, deployment streams, closed-loop manipulation
备注:
点击查看摘要
Abstract:Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.
63. 【2608.09445】DiffSafeMerge: Mitigating Backdoor Inheritance in Diffusion Model Merging
链接:https://arxiv.org/abs/2608.09445
作者:Jiayang Zhang,Ji Guo,Jiachen Li,Wenshu Fan,Wenbo Jiang
类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:Unconditional diffusion checkpoint, Unconditional diffusion, compromised public checkpoint, assumes benign sources, diffusion checkpoint merging
备注:
点击查看摘要
Abstract:Unconditional diffusion checkpoint merging assumes benign sources, yet a compromised public checkpoint can transfer a dormant backdoor while clean generation appears normal. Mitigation is difficult without knowing the compromised source, trigger, or target, and broad sanitization may degrade image quality. We introduce DiffSafeMerge (DSM), which uses a small unlabeled clean set and fixed, attack-agnostic stress probes to score source blocks, shrink suspicious contributions toward a trusted reference, and select attenuation under a clean denoising-loss budget. We evaluate four attacks, two datasets, and 21 target conditions. Intended merging already has zero worst-target ASR in 10 of 14 source cases; DSM preserves these outcomes and records no target match in the remaining four over three seeds, including three with baseline ASR of 48--100\%. Among methods with zero worst-target ASR on both datasets, DSM obtains the lowest case-averaged FID in the matched seed-0 comparison.
64. 【2608.09438】Unveiling the Secret of AdaLN-Zero in Diffusion Transformer
链接:https://arxiv.org/abs/2608.09438
作者:Jie Zhu,Mingyu Ding,Boqiang Duan,Leye Wang,Jingdong Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:rapidly emerging architecture, Diffusion transformer, gained much attention, rapidly emerging, emerging architecture
备注: Accept by IEEE TPAMI 2026, camera-ready version
点击查看摘要
Abstract:Diffusion transformer (DiT), a rapidly emerging architecture for image generation, has gained much attention. However, despite ongoing efforts to improve its performance, the understanding of DiT remains superficial. In this work, we delve into and investigate a critical conditioning mechanism within DiT, adaLN-Zero, which achieves superior performance compared to adaLN. Our work studies three potential elements driving this performance, including an SE-like structure, zero-initialization, and a "gradual" update order, among which zero-initialization is proved to be the most influential. Building on this understanding, we propose an analysis-guided initialization strategy, termed adaLN-Gaussian, which serves both as an empirical validation of our analysis and as a practical initialization method that consistently improves optimization efficiency. On the other hand, inspired by the SE-like structure, we introduce an improved conditioning mechanism called SE-adaLN-Zero. Extensive experiments following DiT on four datasets, especially on ImageNet1K demonstrate the effectiveness and generalization of adaLN-Gaussian and SE-adaLN-Zero. Beyond class-to-image generation, we also evaluate the generalization of the two improved methods on text-to-image generation.
65. 【2608.09427】Foundation Models are Implicit Deepfake Detectors
链接:https://arxiv.org/abs/2608.09427
作者:Stefan Smeu,Dragos-Alexandru Boldisor,Elisabeta Oneata,Dan Oneata
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:fake media distinguishable, properties make real, Pretrained self-supervised representations, deepfake detection methods, media distinguishable
备注:
点击查看摘要
Abstract:Pretrained self-supervised representations have emerged as a core component of current deepfake detection methods, yet it remains unclear which of their properties make real and fake media distinguishable. In this work, we uncover a surprisingly consistent phenomenon: across multiple pretrained models, datasets, and both image and video domains, fake samples systematically produce lower-magnitude representations than their real counterparts. Motivated by this finding, we formulate deepfake detection as an anomaly detection problem and show that simple statistics of feature magnitude achieve competitive performance with far more sophisticated deepfake detection methods. We further investigate the origin of this effect and demonstrate that reduced feature magnitude is primarily associated with semantic shifts introduced by fake content, while low-level generative fingerprints play a comparatively smaller role. Finally, we show that this discriminative signal strengthens as the size of the underlying foundation model grows, suggesting that advances in representation learning naturally translate into stronger zero-shot deepfake detectors.
66. 【2608.09405】MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution
链接:https://arxiv.org/abs/2608.09405
作者:Axi Niu,Jiawei Kou,Kang Zhang,Qingsen Yan,Jinqiu Sun,Yanning Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:costly iterative denoising, requires costly iterative, Diffusion-based super-resolution, achieves strong perceptual, strong perceptual quality
备注: 7 pages, 8 figures. Submitted to AAAI 2027
点击查看摘要
Abstract:Diffusion-based super-resolution (SR) achieves strong perceptual quality but requires costly iterative denoising. Existing one-step distillation methods reduce inference time but depend on expensive pretrained teachers, whereas CTMSR avoids distillation through PF-ODE consistency training yet does not explicitly model the restoration dynamics from low-resolution (LR) inputs to high-resolution (HR) images. We propose MeanSR, a one-step perceptual SR method that learns an LR-conditioned average velocity field to directly capture the finite-time transition from degraded or noisy inputs to plausible HR outputs. We further reformulate distribution trajectory matching for average-velocity generation and introduce a Stage-Aware Temporal Sampling strategy to improve trajectory learning. Experiments on synthetic and real-world benchmarks show that MeanSR outperforms CTMSR on CLIPIQA, MUSIQ, and MANIQA while substantially reducing FLOPs and inference latency. MeanSR also reconstructs sharper structures and more realistic textures with fewer perceptual artifacts.
67. 【2608.09403】One Model to Magnify Them All: Efficient Scale-Invariant Histopathology via Conditional Normalization and Continuous Magnification Training
链接:https://arxiv.org/abs/2608.09403
作者:Agnieszka Florkowska,Henning Müller,Marek Wodzinski
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:fine-grained cellular morphology, levels encoding complementary, encoding complementary diagnostic, complementary diagnostic information, global tissue architecture
备注:
点击查看摘要
Abstract:Whole slide images (WSIs) in digital histopathology are acquired at discrete magnification levels encoding complementary diagnostic information from global tissue architecture to fine-grained cellular morphology. Yet, deep learning models remain sensitive to scale variation. Existing magnification-invariant methods rely on multi-scale architectures at predefined discrete resolutions, while in clinical deployment the acquisition magnification varies continuously, rarely aligns with a model's fixed training resolution, and intermediate scales are common, so robust coverage otherwise demands a costly ensemble of magnification-specific models. We propose Conditional Layer Normalization (CLN), a lightweight mechanism that generates affine normalization parameters from input pixel size via a small MLP, integrated into standard CNN architectures for both WSI classification and segmentation. Trained on patches sampled continuously across a range of pixel sizes, the model decouples inference from scanner-dependent magnification and generalizes to arbitrary, previously unseen scales at test time. On the PANDA prostate cancer dataset, our approach on average matches or exceeds independently trained single-magnification models and ranks among the top three performers at every evaluated magnification, including those unseen during training. This collapses a five-model ensemble into a single network and reduces training, and inference cost roughly 4-5 times, while leaving the multiply-accumulate count unchanged. The code is available at: this https URL.
68. 【2608.09400】Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
链接:https://arxiv.org/abs/2608.09400
作者:Rustem Ozakar,Eyup Gedikli
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:sign language recognition, depth images, sign language, RGB images, images
备注:
点击查看摘要
Abstract:Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are limited. Point clouds obtained from depth images can be used for sign language recognition with neural networks like PointNet. In recent years, various neural networks are used for generating realistic depth images from monocular RGB images. In this work, synthetic depth images were created from RGB images using Depth Anything V2 network. For this purpose, three sign language datasets (Real-time ASL Fingerspelling, KArSL, AUTSL) which contain both RGB and depth images were used. Classification accuracies of the point cloud data created from both original and synthetic depth images using various PointNet architectures were measured for sign language recognition. From the original and synthetic point clouds, frame based, Point Gesture Map and Long Short Term Memory data models were used for classification and their performances were compared. In the results, both original and synthetic based data achieved acceptable performance in most models. In general, original depth based point cloud models performed better than synthetic ones, however in some models synthetic depth based models performed better than the originals.
69. 【2608.09392】CableDex: Cable Length Estimation on Industrial Reels Using a Handheld Device
链接:https://arxiv.org/abs/2608.09392
作者:Francisco Guillén,Ricardo Almeida,Bruno Silva,João C. Neves
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:single photograph captured, computer vision system, inaccurate manual measurement, computer vision, addresses the time-consuming
备注:
点击查看摘要
Abstract:CableDex is a computer vision system that addresses the time-consuming and inaccurate manual measurement of cable length on industrial reels from a single photograph captured with a mobile phone. The system combines camera calibration, instance segmentation, pose estimation, and volumetric calculation to estimate the cable length across five different reel types and various cable sizes. This system is based on an instance segmentation model trained on 1,000 manually annotated images, achieving 99.5\% mAP50 with an inference time of 5.66 ms per image. Evaluated on 75 reels across five reel types, the system achieves a MAPE of 4.90\%, within the 10\% error tolerance commonly accepted in industrial cable-reel measurement. The demonstration presents the end-to-end pipeline, from reel label scanning and image capture to segmentation and length estimation, through the mobile application.
70. 【2608.09391】CoInS-Net: A Continuous Position-Aware Network for Joint Medical Image Interpolation and Segmentation
链接:https://arxiv.org/abs/2608.09391
作者:Yujia Sun,Ningfeng Que,Peiting Shi,Rongrong Fu,Yingying Yang,Xinhang Li,Yin Dai
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Accurate medical image, Accurate medical, treatment planning, fundamental for computer-aided, computer-aided diagnosis
备注:
点击查看摘要
Abstract:Accurate medical image interpolation and anatomical structure segmentation are fundamental for computer-aided diagnosis and treatment planning. Anisotropic medical volumes with sparse through-plane sampling often suffer from structural discontinuity and boundary blur, hindering reliable clinical image analysis. Most existing methods implement interpolation and segmentation independently, which introduces redundant computation and fails to fully exploit complementary cross-slice structural information between sequential slices. To address these issues, we propose a continuous position-aware interaction network, termed CoInS-Net, for joint frame interpolation and lesion segmentation. Unlike conventional cascaded interpolation-then-segmentation paradigms, the framework enables bidirectional interaction under a shared Swin encoder with continuous spatial coordinate queries. A spatially continuous position interpolation module generates target-position features at every scale from the relative coordinate and physical spacing, and a prototype-based task mutual interaction module lets the segmentation and interpolation branches exchange global structure through a small set of shared prototypes rather than dense feature mixing. A multi-scale task-cooperative decoder further separates each scale into shared and task-specific components, so the two tasks reinforce common anatomy while preserving their distinct requirements down to the boundary level, without extra annotations. Experiments on four public medical imaging datasets with diverse modalities and anatomical regions demonstrate that the proposed method outperforms conventional single-task schemes. The joint optimization framework effectively realizes mutual promotion between interpolation and segmentation tasks, providing a reliable and universal technical scheme for intelligent clinical medical image analysis.
71. 【2608.09388】Efficient Human-Contact Representation for Human-Scene Interaction
链接:https://arxiv.org/abs/2608.09388
作者:Nghia Vu,Tuong Do,Binh X. Nguyen,Erman Tjiputra,Anh Nguyen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:active research topic, virtual reality, Human-scene interaction, active research, research topic
备注: Accepted in ECCV 2026 Workshops
点击查看摘要
Abstract:Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite significant progress in network architectures to improve the results or optimize models' parameters for fast inference speed, the efficient representation of contact between humans and their environments remains an open challenge. In this paper, we propose a new efficient human-contact representation for human-scene interaction. Our primary contribution is the introduction of sparse contact masks that strategically select essential contact information, significantly reducing redundant data in high-dimensional inputs. Leveraging this efficient contact representation, we propose a suite of sparse operators to replace traditional dense operators within deep network layers for faster computation. Our approach not only enhances computational speed but also filters out non-essential contact data, thereby improving the precision of human-scene interaction models. To validate the effectiveness of our method, we conduct intensive experiments across three public benchmark datasets, focusing on two critical tasks for human-scene interaction: contact prediction and scene synthesis. The experimental results show that our approach outperforms state-of-the-art models in reconstruction accuracy and achieves a computation speed-up of at least 12 times over recent baselines.
72. 【2608.09385】Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation
链接:https://arxiv.org/abs/2608.09385
作者:Hossein Goli,Farzan Farnia,Amin Gohari
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:introduce Imaginative Generative, diversity, data distribution, primarily designed, designed to imitate
备注:
点击查看摘要
Abstract:Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learned generator nor defines how generation should extend beyond the diversity of the data itself. We introduce Imaginative Generative AI (IGA), a framework that makes diversity part of the target-distribution design problem: among distributions close to a reference, IGA selects one whose spectral diversity reaches a prescribed level. Diversity is measured by the von Neumann entropy of the generated distribution's kernel covariance operator in a fixed representation space, providing a reference-free representation-guided measure of how broadly probability mass occupies embedding directions. The spectral entropy of the population data distribution defines an Entropy Wall. Below the wall, IGA performs diversity repair, recovering variation that a learned generator has lost while remaining within the diversity level of the data. Beyond the wall, the data distribution itself becomes infeasible, and IGA deliberately departs from it to produce distributions with greater representation-relative spectral diversity, an operational notion of imaginative generation. These regimes form a single regularization path from imitation to imagination and define an i.i.d. target distribution at each prescribed diversity level. We develop the theory of this entropy-constrained projection and show that, under a KL anchor to a pretrained generator, the optimum satisfies a self-consistent exponential-tilt relation. This characterization leads to IGA Guidance, a retraining-free inference-time method for score-based and diffusion models, including DDPM and DDIM samplers. Experiments on synthetic and vision benchmarks demonstrate diversity repair below the Entropy Wall and controlled spectral extrapolation beyond it.
73. 【2608.09374】CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits
链接:https://arxiv.org/abs/2608.09374
作者:Xinqi Yang,Kang An,Tengyue Wang,Zhongyu Yang,Chenxu Du,Yuanchi Zhu,Hebao Zhu,Ziliang Wang,Faqiang Qian,Yunli Yang,Qibing Ren
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Electrical circuit analysis, Electrical circuit, recognizing components, circuit analysis requires, Electrical
备注:
点击查看摘要
Abstract:Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, select a physical model, formulate coupled equations, propagate intermediate quantities, and preserve units, signs, directions, and phase conventions. We introduce \benchmark, a benchmark of 1,000 authentic textbook problems for evaluating this complete long-horizon visual-to-symbolic reasoning process. Each problem pairs one or more circuit diagrams with a self-contained question, a typed or semantically specified answer, and a reference worked solution. An evidence-first construction pipeline aligns questions, figures, and solutions, while a reasoning-oriented taxonomy organizes problems by circuit type and dependency depth. Evaluation combines conservative typed scoring with identity-blinded multi-model semantic consensus, retaining every problem in the denominator. Across three commercial chatbot systems and six open-source multimodal large language models, the highest-scoring system reaches 84.8\% accuracy. However, performance consistently deteriorates on long-horizon problems, and qualitative analysis exposes persistent failures in topology-to-target binding, physical conventions, and late-stage output propagation. \benchmark{} provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning. Code are available at GitHub - CircuitReason/CircuitReason1K.
74. 【2608.09373】Preserve More Details: Mitigating Content Drift in Real-World Image Super-Resolution
链接:https://arxiv.org/abs/2608.09373
作者:Chunxiao Liu,Wei Liu,Anbin Xiong,Erli Meng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Real-world image super-resolution, diverse real-world degradations, diverse real-world, aims to reconstruct, reconstruct high-quality
备注: Accepted to ACM MM 2026. This is the author's accepted version. The definitive version is published in the Proceedings of ACM MM 2026
点击查看摘要
Abstract:Real-world image super-resolution (Real-ISR) aims to reconstruct high-quality (HQ) images from low-quality (LQ) inputs subject to diverse real-world degradations. Recent advances have leveraged the LQ inputs and natural image priors learned by Stable Diffusion models to achieve impressive results. However, existing methods often overlook insufficient clarity of LQ inputs inevitably induce content drift in the generated HQ images. This manifests primarily as visual detail degradation and textual semantic shift, severely compromising both fidelity and perceptual quality. To address this challenge, we propose FSP-Diff, a novel one-step diffusion model featuring a dual-pathway architecture. This architecture comprises a Detail-Conditioned Pathway for injecting structured details to recover fine structures, and a Detail-Modulated Semantic Pathway that refines semantic guidance using structured details to mitigate semantic deviations. Extensive experiments on standard Real-ISR benchmarks demonstrate that FSP-Diff surpasses existing one-step diffusion methods in both quantitative and qualitative metrics.
75. 【2608.09369】FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking
链接:https://arxiv.org/abs/2608.09369
作者:Yueyang Cang,Xiaoteng Zhang,Zhiyuan Ning,Yuchen He,Li Shi
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Visual object tracking, effective temporal integration, Visual object, requires effective temporal, feed-forward feature extraction
备注:
点击查看摘要
Abstract:Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feature extraction. Existing temporal mechanisms typically update templates, prompts, queries, or prediction states, while intermediate representations are rarely reused to modulate corresponding processing stages. We propose \textbf{FeedbackTrack}, a visual-cortex-inspired framework that introduces sparse, group-level layer-aligned cross-frame feedback into pretrained Transformer trackers. Previous-frame intermediate states are detached, cached, and returned to corresponding Transformer groups in the current frame through two lightweight pathways: Query Feedback for token-level query modulation and Gate Feedback for context-dependent feature modulation. FeedbackTrack preserves the original tracking pipeline with only a fixed-size one-frame cache. Across SPMTrack and ARTrackV2, FeedbackTrack consistently improves five backbone configurations on LaSOT and GOT-10k, achieving 83.4 AO and 79.1 AUC with SPMTrack-G while adding less than 1\% parameters. Controlled comparisons show that cross-frame feedback outperforms same-frame modulation by 1.8--3.2 AO points, demonstrating that the gains mainly come from recurrent historical information. Further analysis reveals a non-uniform depth-dependent organization of learned feedback strengths, highlighting the effectiveness of recurrent feedback for Transformer tracking.
76. 【2608.09360】Deep Learning based Detection of Fishing Vessels and Fishing Monitoring using Nightlight Images
链接:https://arxiv.org/abs/2608.09360
作者:Shantakar Mohanty,Prasun Kumar Gupta,Raian Vargas Maretto
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Image and Video Processing (eess.IV); Applications (stat.AP)
关键词:Automatic Identification System, Identification System, Automatic Identification, operate without Automatic, fishing vessel activities
备注:
点击查看摘要
Abstract:The demand for maritime surveillance has given rise to the need for monitoring fishing vessel activities, particularly in addressing the challenge of "dark vessels" that operate without Automatic Identification System (AIS) transmission. This study presents a novel approach for detecting small-scale fishing vessels using nighttime light (NTL) imagery from the SDGSAT-1 satellite, combined with deep learning techniques to enhance fishing monitoring awareness along the western coast of India. A dual-branch YOLO11 architecture was developed to exploit both the 10-meter panchromatic and 40-meter RGB imagery from SDGSAT-1. The custom model architecture was specifically optimized for small object detection in NTL imagery, featuring parallel convolutional backbones that process both modalities before concatenation for enhanced feature extraction. The dual-branch YOLO11 model demonstrated optimal performance with a precision of 0.99, recall of 0.93, F1-score of 0.96, and mAP@50 of 0.96, significantly outperforming single-branch implementations of YOLOv5s, YOLOv8s, and standard YOLO11s architectures. When applied to the western coast of India, the model detected 31525 vessel instances across the temporal dataset spanning 2022-23. Cross-matching analysis with AIS data revealed that only 7146 (22.7%) of detected vessels had corresponding AIS transmissions, while 24379 (77.3%) were identified as potential dark vessels. Spatio-temporal analysis showed peak fishing activity during January-April, with a primary activity corridor parallel to the coastline within 50-100 km, corresponding to productive continental shelf areas. This research contributes to maritime surveillance capabilities by highlighting the effectiveness of nighttime lights satellite imagery for fishing vessel detection and provides valuable insights into fishing patterns and potential regulatory compliance issues in Indian waters.
77. 【2608.09357】ControlRadio: Prompt-Driven Controllable Diffusion for Cross-Modal Radio Map Generation
链接:https://arxiv.org/abs/2608.09357
作者:Kangjun Liu,Xiying Pan,Shuhang Zhang,Xiang Xiang,Ke Chen,Yaowei Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:wireless signals propagate, Radio maps describe, Radio maps, network planning, accurate radio maps
备注: 17 pages, 9 figures, 8 tables
点击查看摘要
Abstract:Radio maps describe how wireless signals propagate across space and are essential for wireless communication, sensing, and network planning. However, constructing accurate radio maps traditionally requires either dense measurements or computationally expensive physical simulations, which limits scalability and real-time deployment. Recent advances in generative artificial intelligence offer a promising alternative, but existing approaches lack fine-grained control and physical consistency when applied to real-world wireless environments. Here we present \textbf{ControlRadio}, a controllable generative framework that produces radio maps from natural-language descriptions and environmental layouts, including building structures and transmitter locations. Joint semantic and spatial conditioning enables interpretable, propagation-plausible generation, while a controlled latent prior and layout-aware conditioning improve stability and structural consistency. Extensive experiments demonstrate that ControlRadio achieves state-of-the-art accuracy and strong generalization across diverse urban scenarios, while reducing computation time by more than four orders of magnitude compared with conventional simulation-based methods. Such results suggest a new paradigm for scalable and controllable wireless environment modeling, with broad implications for next-generation communication systems and data-driven radio sensing.
78. 【2608.09355】Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation
链接:https://arxiv.org/abs/2608.09355
作者:Zhe Li,Honghao Qiao,Zhixin Xu,Qijie Wang,Bo Peng,Dawei Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:combine RGB appearance, enabling animated assets, RGBA videos combine, RGBA video, videos combine RGB
备注:
点击查看摘要
Abstract:RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in gaming industry. However, generating high-quality RGBA animations for games remains challenging for two reasons. First, most existing RGBA video datasets are dominated by photorealistic content, with limited coverage of game assets. Second, the traditional generate-then-matte pipelines estimate alpha only after RGB synthesis, so semi-transparent regions are often blurred by background, resulting in unstable matting outputs. More recently, many methods have begun to model RGB and alpha jointly, but existing approaches are mostly text-conditioned, and still have unresolved issues in efficiency and quality. To address these challenges, we introduce GameAlpha-2.4K, a 2.4K-clip game-style RGBA video dataset built with matte-friendly synthesis, multi-hypothesis alpha recovery, and compositing-based quality gates. Using this dataset, we train a reference-conditioned RGBA video generator that jointly produces RGB frames and alpha mattes in a single pass. To improve efficiency, we propose a visibility router that identifies transparent tokens in an early stage and bypasses their later DiT updates, while x_0-lock guides them along the original flow-matching schedule toward self-predicted endpoints. Our model obtains lower FVD than traditional two-stage pipelines, and the visibility router skips 35% of token evaluations in the final two DiT denoising steps, providing a 1.2x backbone speedup with negligible quality degradation compared to dense inference.
79. 【2608.09345】One-Time Training for All Grains: Open-Set Grain Recognition and Quantitative Analysis
链接:https://arxiv.org/abs/2608.09345
作者:Qihe Su,Mengyu Sun,Yuxi Ke,Zhuoyan Jiang,Wanneng Yang,Chenglong Huang,Ziyuan Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Advances in crop, newly introduced varieties, creating a growing, crop breeding, increasing number
备注: 15 pages, 15 figures
点击查看摘要
Abstract:Advances in crop breeding have introduced an increasing number of grain varieties, creating a growing demand for efficient variety recognition and quantitative analysis. However, existing methods are typically trained on a fixed variety set, and incorporating newly introduced varieties requires additional data collection and model retraining. To address this limitation, we propose GROW, a framework for Grain Recognition and quantitative analysis in Open sets Without retraining. GROW first performs class-agnostic grain localization, converting mixed-grain images into individual instances for variety-wise counting and phenotypic measurement. It then combines visual embeddings and morphological descriptors into fused grain descriptors stored in an extensible GrainBank. Query grains are recognized through rank-similarity weighted top-k retrieval, and newly introduced varieties are incorporated by appending their descriptors without updating the deployed models. Extensive experiments under progressive variety expansion, varying grain densities, and background domain shifts demonstrate the scalability, robustness, and adaptability of GROW. Compared with joint retraining, GROW reduced the average category-registration time from 4153 s to only 39 s while maintaining competitive recognition performance. These results demonstrate that GROW provides an efficient and maintainable solution for extensible grain recognition, counting, and phenotypic analysis without repeated model retraining.
80. 【2608.09344】Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs
链接:https://arxiv.org/abs/2608.09344
作者:Ali Cheraghian,Hamidreza Dastmalchi,Hamed Barzamini,Morteza Saberi,Mojtaba Golzan,Shafin Rahman,Hossein Rahmani
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:enabled powerful multimodal, powerful multimodal reasoning, Recent advances, integrating visual encoders, enabled powerful
备注: BMVC 2026
点击查看摘要
Abstract:Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language models (LLMs). However, their reliability is frequently undermined by hallucinations, where generated text inaccurately describes the visual input. Although fine-tuning can mitigate this problem, it is computationally expensive and requires large, curated datasets, making training-free alternatives attractive. Among these, model editing is more promising than decoding-based approaches: decoding methods adapt outputs per input but introduce computational overhead and instability, whereas model editing modifies internal representations offline, providing a more efficient and stable solution. However, existing model-editing techniques typically rely on a single global subspace to correct hallucinations, treating all test samples identically and failing to capture diverse hallucination modes across inputs. To address this limitation, we propose a training-free hallucination mitigation framework for dynamic, per-instance suppression at test time. Our method first constructs a set of Disentangled Hallucination Subspaces, each isolating a distinct hallucination mode. During inference, the model adaptively calculates weights reflecting each input's relationship to these subspaces, guiding a dynamically combined projection that selectively suppresses the most probable hallucination directions while preserving image-grounded semantics. Extensive experiments across multiple vision-language benchmarks and LVLM families demonstrate consistent improvements, highlighting the robustness, generalizability, and efficiency of our approach.
81. 【2608.09342】Revisiting the Current Frame: Physical-Trace-Guided Network Output Correction for Video Restoration
链接:https://arxiv.org/abs/2608.09342
作者:Yifeng Lin,Liuxiang Qiu,Guangming Ren,Tiesong Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:recover information missing, exploit temporal information, methods exploit temporal, restoration methods exploit, recover information
备注: 9 pages, 6 figures
点击查看摘要
Abstract:Video restoration methods exploit temporal information to recover information missing from degraded observations. However, reference frames within the sequence may introduce inconsistent degradation, content discrepancy, or reconstruction errors due to physical image-formation variations, occlusion, and imperfect temporal aggregation. Existing approaches mainly focus on improving restoration networks, while the reliability of the generated outputs at different spatial locations remains largely unexplored. In this work, we propose ANCHOR, a model-agnostic framework that revisits the low-quality current frame as a temporally aligned anchor for video restoration correction. Specifically, ANCHOR estimates a spatial trust field from heterogeneous physical-trace evidence and adaptively balances the restoration proposal with the original observation. Experiments on High Dynamic Range video reconstruction and video deraining demonstrate consistent improvements across various state-of-the-art restoration models, validating the effectiveness of reliability-aware output correction for video restoration.
82. 【2608.09325】GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models
链接:https://arxiv.org/abs/2608.09325
作者:Zhihang Liu,Mei-Po Kwan,Jinlin Wu,Hao Li
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Newly triggered landslides, regional risk assessment, Newly triggered, triggered landslides rarely, landslides rarely carry
备注:
点击查看摘要
Abstract:Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational transfer, yet on unseen regions, events, and data sources they still generate high-confidence false alarms. Terrain, material, and rainfall triggering can constrain such errors, but their supports are local, regional, and event-scale, so that resampling onto a 10~m grid misaligns them with the segmentation decision unit and compounds the uncertain geographic context problem (UGCoP). We propose GeoPhysAdapter, which anchors on a frozen vision foundation model, restricts terrain, material, and triggering to dense spatial guidance, regional modulation, and event-timing forcing, and applies bounded adaptation at two decision units, the pixel and the candidate landslide body, reverting exactly to the visual prediction where support is insufficient. On an event-isolated PILD dataset of four public sources, 55 global landslide events, and 7,890 test samples, 70.3% of cross-domain false-positive mass lies in near-pure spurious bodies of median equivalent diameter 207m, matching coarse-prior support rather than the pixel. Pixel-level adaptation removes a net 507,817 erroneous pixels and reduces error by 7.76%, whereas raising the decision unit to the candidate body, under identical samples, anchor, and baseline, increases error reduction to 23.99%, approximately 3.1 times the pixel-level effect, improves IoU by 0.031 (14.2% relative), and corrects 9.92 pixels per pixel harmed. The data and code are publicly available at: this https URL.
83. 【2608.09322】Diffusion Image Editing via Asynchronous Token Decoding
链接:https://arxiv.org/abs/2608.09322
作者:Yang Shi,Liangsi Lu,Minzhe Guo,Yifeng Xie,Yanhui Chen,Jingchao Wang,Xuhang Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Text-guided diffusion image, Text-guided diffusion, diffusion image editing, image editing aims, modify semantic attributes
备注: Accepted by ACMMM 2026
点击查看摘要
Abstract:Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naïvely switching the text condition during sampling often causes global drift, as denoising dynamics propagate changes across tokens and can disrupt unedited regions. To address this issue, we propose \textbf{A}synchronous \textbf{T}oken \textbf{D}ecoding \textbf{Edit} (ATDEdit), an inference-time framework that views each sampler step as a parallel update of a globally coupled token matrix and enables token-indexed condition switching with differentiated update policies. Instead of applying synchronous target-conditioned updates to all tokens, ATDEdit estimates editable locations using token-wise conditional surprisal and applies target-conditioned corrections to the selected token set. It supplies source key/value memory at keep-token positions and projects selected keep-token latent rows back to their source values; these operations promote background preservation but do not constitute a pixel-level invariance guarantee. This approach combines local editing and background preservation without external or user-provided spatial masks and without model fine-tuning. On PIE-Bench, ATDEdit achieves the strongest reported preservation metrics, including 27.44~dB PSNR and 0.055 LPIPS, while retaining competitive semantic alignment.
84. 【2608.09321】Warp-free Cross-view Geo-localization via Feature-space Consensus Mining
链接:https://arxiv.org/abs/2608.09321
作者:Zhuo Song,Lian Xu,Runqing Jiang,Yongjian Zhang,Kunhong Li,Ye Zhang,Yulan Guo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:large appearance discrepancies, satellite imagery, challenging due, due to drastic, drastic viewpoint
备注:
点击查看摘要
Abstract:Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences. To overcome this, we propose a novel joint-view consensus-guided learning framework that entirely bypasses explicit geometric warping. Instead of forcing rigid spatial alignment, we dynamically mine and adaptively strengthen a semantic consensus directly within the feature space. Specifically, an auxiliary joint-view pathway during training enables direct cross-view interaction, allowing each view to selectively aggregate corroborative evidence into a unified consensus representation. To resolve feature heterogeneity among the single- and joint-view streams, we introduce global pattern probes acting as a semantic dictionary to project divergent modalities into a strictly aligned metric space. Guided by a consensus-mediated contrastive objective, single-view embeddings are explicitly pulled toward the joint-view anchor during training, distilling this consensus-mining capability into the single-view encoders for robust retrieval at inference. Extensive experiments demonstrate that our method achieves state-of-the-art performance across four standard benchmarks, underscoring the importance of discovering cross-view semantic consensus for reliable geo-localization.
85. 【2608.09316】MemeMind: Reference-Guided Trace Construction for Offline Context Optimization
链接:https://arxiv.org/abs/2608.09316
作者:Run Yang,Weihang Wang,Boheng Sheng,Yuchen He,Jielei Zhang,Pengyu Chen,Zhiyu Wu,Qiang Sun,Huyang Sun,Longwen Gao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:agent by revising, offline reference answer, Offline context optimization, offline reference, reference
备注: 24 pages, 16 figures, 8 tables
点击查看摘要
Abstract:Offline context optimization improves an agent by revising its instructions and examples while keeping the model frozen. This approach learns from rollouts on an adaptation set, but some queries produce only failed rollouts. In these cases, the optimizer sees no successful example of how the available tools can reach the correct answer. We introduce MemeMind, which uses an offline reference answer to recover this missing experience. TraceBuilder identifies the evidence required by the reference, executes text search, image retrieval, and visual grounding, and verifies the resulting tool trace before adding it to the adaptation buffer. ToolGuide then summarizes the collected traces into a shared guide and separate instructions for each tool. The reference answers and constructed traces are used only during adaptation, while inference uses the learned guides with a frozen model. We study this problem through Anime, Comic, and Game meme interpretation. These memes combine edited and ambiguous visual content, overlaid text, long tail franchise knowledge, and culture specific references. Their interpretation can require coordinated visual grounding, image retrieval, and text search, making them a demanding setting in which native rollout groups may fail together. We evaluate MemeMind on MemeX, a benchmark of 1,000 such memes annotated by experts. Across two Qwen3-VL models, two language partitions, and two independent judges, MemeMind improves over the strongest context optimization baseline by 22.0% and 21.1% on Qwen3-VL-30B-A3B, and by 8.1% and 8.0% on Qwen3-VL-235B-A22B under GPT-5 judging. Ablations and held out traces show that constructing successful tool use for failed groups provides the largest component gain and produces more effective evidence acquisition at inference time.
86. 【2608.09311】Degraded Infrared Small Object Detection via Degradation-Adapted Physics-Guided Restoration
链接:https://arxiv.org/abs/2608.09311
作者:Xinkai Lu,Wenjun Chen,Yi Li,Yi Chang,Luxin Yan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:made significant progress, Infrared small object, small object detection, recent years, made significant
备注: Accept by ICIG2026 (Oral)
点击查看摘要
Abstract:Infrared small object detection has made significant progress in recent years. However, degradations such as fog and nonuniformity can suppress target-background contrast, substantially increasing detection difficulty. Existing methods mainly rely on image restoration as preprocessing, but they are typically designed for specific degradation types and fail to generalize to varying degradations. To alleviate this, we propose DAISOD, a degradation-adapted infrared small object detection framework for robust detection under different degradations. DAISOD first identifies the type and severity of degradations, then adapts the processing via dedicated branches, and finally fuses the results for subsequent detection. Moreover, a physics-guided restoration mechanism is incorporated to explicitly estimate degradation parameters and remove degradation effects through physical models, avoiding excessive restoration that may erase small targets. Moreover, we construct a degraded infrared small object detection dataset covering diverse degradation types and levels. Extensive experiments show that DAISOD outperforms state-of-the-art methods under various degradation conditions.
87. 【2608.09302】Bootstrapping Vision-Language Model for Hysteroscopic Surgical Scene Segmentation
链接:https://arxiv.org/abs/2608.09302
作者:Jun Huang,Meiyi Chen,Zijie Yue,Yuhang Xiao,Fang Li,Hanli Wang,Xiaowen Tong,Yi Guo,Miaojing Shi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Hysteroscopic surgical scene, hysteroscopic intraoperative environment, surgical scene segmentation, Hysteroscopic surgical, surgical scene
备注: Accept by Biomedical Signal Processing and Control
点击查看摘要
Abstract:Hysteroscopic surgical scene segmentation plays a pivotal role in understanding the hysteroscopic intraoperative environment as well as computer-assisted intervention. However, this task presents unique challenges due to the high morphological similarity among different lesions and the presence of artifacts such as specular reflections, motion blur, and fluid occlusions in surgical videos. In this work, we propose the first vision-language model (VLM)-based hysteroscopic surgical scene segmentation method, which performs pixel-wise localization for fifteen representative categories in hysteroscopic surgical scenes. Our VLM-hyster has a segmentation backbone that utilizes the pretrained image encoder for robust visual feature extraction, coupled with a transformer-based decoder for dense prediction. Moreover, we design category-specific text prompts and incorporate a masked distillation branch to filter out visual features with low correlation to the text prompts, enabling the model to focus more effectively on category-specific image regions and thereby enhancing segmentation performance. We collect a large multicentric hysteroscopic surgical scene dataset, containing 4,020 high-resolution images with detailed mask annotations, for model training and evaluation. Experimental results demonstrate that VLM-hyster substantially outperforms state-of-the-art AI models. Furthermore, extensive assessments by gynecologists, as well as multicentre and prospective validations, demonstrate VLM-hyster's robustness and generalizability. The results suggest that VLM-hyster earns considerable potential in enabling AI-assisted localization of surgical instruments and lesions in hysteroscopic surgeries. Code is available at this https URL.
88. 【2608.09296】CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
链接:https://arxiv.org/abs/2608.09296
作者:Harmanjot Singh,Abhra Dubey,Jorge Alejandro Amador Herrera
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
关键词:Abstract, satisfy design requirements, CAD, reference structural response, correct
备注:
点击查看摘要
Abstract:A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints. We present CADEngBench, a two-track benchmark for these capabilities. CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and one functional-editing task (600 tasks in total), through boundary-representation (B-Rep) validity, engineering and DFM checks, parameter-family perturbations, functional editing, and matched linear-static FEA in CalculiX. CADEngBench-A evaluates 150 body pairs through ranked joint retrieval, exact face-and-edge grounding, joint-frame prediction, and kinematic verification. Across eight multimodal, code-capable models, editing supplied CAD is substantially easier than generating it, while complex edits and matched FEA remain difficult. Assembly predictions often locate the relevant region but fail to recover the recorded joint or mating entities. These results show that CAD evaluation must test engineering behavior rather than appearance alone.
89. 【2608.09287】UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation
链接:https://arxiv.org/abs/2608.09287
作者:Xuewan He,Tong Chu,Zihan Cheng,Yuchen Su,Qianxin Xia,Guoming Lu,Jielei Wang,Wen Li
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:original training dataset, compact student model, synthesizing semantically informative, semantically informative data, pretrained teacher model
备注:
点击查看摘要
Abstract:Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing semantically informative data, eliminating the need for access to the original training dataset. Existing DFKD methods rely heavily on architecture-specific statistical priors (e.g., Batch Normalization statistics) to guide data synthesis, however, such architecture-dependent priors are often absent in modern architectures such as Vision Transformers (ViTs), resulting in degraded semantic quality of the synthesized data and consequently catastrophic performance degradation. In this paper, we propose \emph{UniDFKD}, a unified data-free knowledge distillation framework that replaces architecture-specific statistics with explicit, architecture-agnostic semantic priors. \emph{UniDFKD} governs the entire synthesis-distillation pipeline along three dimensions: (1) Categorical Semantic Conditioning (CSC) defines \emph{what} to synthesize by persistently modulating the generator with language-derived embeddings to capture semantic diversity; (2) Spatial Semantic Anchoring (SSA) dictates \emph{where} evidence belongs by anchoring the teacher's spatial attributions to a Gaussian prior; and (3) Spatial Semantic Distillation (SSD) controls \emph{how} knowledge is transferred by explicitly aligning teacher-student spatial evidence alongside predictions. Extensive experiments across CNNs and ViTs demonstrate that UniDFKD establishes a new state-of-the-art, outperforming existing methods by an average absolute margin of over 20\% in both homogeneous and heterogeneous settings.
90. 【2608.09270】GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
链接:https://arxiv.org/abs/2608.09270
作者:Jiahui Cui,Yan Zhao,Kan Wei,Enze Zhu,Peirong Zhang,Lei Wang,Yiru Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Multimedia (cs.MM)
关键词:aerial vision-language navigation, vision-language navigation, Cross-Modal Focus Misalignment, Semantic Prototype Codebook, Semantic Prototype
备注: Accepted at the 34th ACM International Conference on Multimedia (ACM Multimedia 2026, MM '26). 10 pages, 6 figures
点击查看摘要
Abstract:Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone scenarios impose dual challenges on vision-language understanding. At the macro level, overwhelming background clutter in visual representations leads to Cross-Modal Focus Misalignment, where the model prioritizes global environmental similarities over specific object details. At the micro level, Visual Isomorphism creates ambiguity, where candidates share similar geometric structures yet differ only in subtle attributes. To address these challenges, we propose the Granularity-Aware Region Alignment and Semantic Prototype (GRASP) learning framework, enhancing discriminative capability through two synergistic strategies. Specifically, we introduce Region-Focused Alignment (RFA) to promote object-centric cross-modal alignment while suppressing background interference. Concurrently, to tackle visual isomorphism, we propose Semantic Perturbation Enhanced Matching (SPEM), which leverages a foreground-purified Semantic Prototype Codebook (SPC) to construct semantically perturbed negatives for fine-grained semantic discrimination. Extensive experiments on the GeoText-1652 benchmark and the unseen ERA dataset demonstrate that GRASP achieves competitive performance in drone-view fine-grained image-text retrieval, validating its effectiveness for cross-modal understanding in aerial scenarios. Our code implementation is available at this https URL.
91. 【2608.09266】Did the Grid Erase the Event? EndoClock for Auditing Medical World-Model Pipelines
链接:https://arxiv.org/abs/2608.09266
作者:Yarin Udi,Tom Sharon-Shahak,Roee Masad,Dan Pri-Tal
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Medical world models, multimodal recordings synchronized, Medical world, world models commonly, models commonly learn
备注: Accepted for publication at the 1st MICCAI Workshop on Medical World Models (MWM 2026)
点击查看摘要
Abstract:Medical world models commonly learn from multimodal recordings synchronized onto a fixed-rate grid. This preprocessing resamples each native stream onto a shared time axis. Each stream has an observation clock that governs when observations are emitted or updated. When this clock depends on the latent or acquisition state, it is endogenous. In such settings, synchronization may not be neutral and can erase task-relevant evidence before the model sees the data. We introduce a four-regime taxonomy that characterizes where the evidence needed to distinguish a target event or state survives. The relevant witness may remain in the sampled values, in grid-cell update patterns, in native timing, or only in an external acquisition channel. EndoClock operationalizes this taxonomy as a conservative pretraining audit. It reports the lowest witness-bearing representation supported by the available evidence, or unresolved when no regime can be established. We illustrate this failure in echocardiography, where B-mode video write-outs cease during pulsed-wave Doppler acquisition while the corresponding measurement events remain recorded only in an external acquisition log. This work is a preliminary failure alert and executable audit. Its practical message is to preserve the native observation process long enough to determine whether synchronization has erased information required by the intended task.
92. 【2608.09264】ask-Adaptive 3D Cross-Field MRI Translation via Field-Conditioned Content-Style Pretraining
链接:https://arxiv.org/abs/2608.09264
作者:Haowen Pang,Yingqi Hao,Pengli Zhu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:magnetic resonance imaging, Magnetic field strength, magnetic resonance, cross-field MRI translation, spatial detail
备注: MICCAI 2026 Workshop on MRIxFields
点击查看摘要
Abstract:Magnetic field strength is a major source of domain shift in magnetic resonance imaging (MRI), affecting signal-to-noise ratio, tissue contrast, spatial detail, and the visibility of anatomical boundaries. The MRIxFields 2026 challenge investigates this problem through cross-field MRI translation across acquisitions at 0.1T, 1.5T, 3T, 5T, and 7T. Its three tasks, Any-to-7T, 0.1T-to-High, and Any-to-Any synthesis, require the generation of target-field image characteristics while preserving subject-specific anatomy. This problem is particularly challenging because paired acquisitions of the same subject across multiple field strengths are rarely available for training. We propose a 3D unpaired cross-field MRI translation framework based on field-conditioned content-style pretraining. The proposed framework first learns controllable field-to-field translation across all available field strengths by disentangling anatomical content from field-dependent contrast characteristics. The pretrained backbone is then adapted to task-specific target domains. Our model comprises a 3D content encoder, a 3D style encoder, a field-conditioned style generator, an AdaIN-modulated decoder, and a multi-field discriminator. Adversarial learning encourages realistic target-field appearance, while cycle-consistency, identity, content, style, and diversity constraints promote anatomical fidelity and controllable translation. We evaluate the proposed method on MRIxFields data spanning five field strengths and three MRI modalities. Experiments on paired test data demonstrate that the framework can adapt to the three challenge settings while preserving three-dimensional anatomical structure in the synthesized volumes. The implementation code is publicly available at this https URL.
93. 【2608.09244】In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation
链接:https://arxiv.org/abs/2608.09244
作者:Yushun Tang,Weiming Chen,Siyi Liu,Yi Zhang,Feng Wu,Zhihai He
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:achieved remarkable success, reference image, generating high-quality images, diffusion model, image
备注:
点击查看摘要
Abstract:Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized images to mimic the appearance of subjects in given reference images within different visual contexts specified by the text prompts. The central challenge here is that, when the reference image changes, the diffusion model cannot efficiently adapt to different visual contexts while consistently maintaining the subject identity. Existing methods either train the model with a large domain-specific dataset or fine-tune the model using the reference image for hundreds of iterations before actual image generation. In this work, we explore a new approach, called \textit{In-Loop Model Adaptation} (IMA), which adapts the core diffusion model at each generation step during the actual process of image generation, without being trained on the reference image before the generation process. To this end, we establish a DDIM inversion chain that maps the reference image to a sequence of latent, as well as a text-to-image generation chain which generates the image from the text prompt only. We then introduce a masked latent consistency loss and a noise regularization loss to characterize the latent-noise difference between the diffusion model and these two chains at each generation step. This coupled latent-noise loss is used to guide the in-loop model adaptation to preserve the subject identity specified by the reference image while maintaining accurate alignment with the text prompt, resulting in high-fidelity text-to-image generation. Our extensive experiments demonstrate that our proposed IMA method significantly improves the performance of subject-driven text-to-image generation.
94. 【2608.09238】RealDenseFace: Real-time Monocular 3D Face Reconstruction from Dense UV-space Priors
链接:https://arxiv.org/abs/2608.09238
作者:Linzhou Li,Tianjia Shao,Kun Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Morphable Model, Recent monocular, dense priors predicted, face reconstruction methods, face reconstruction method
备注:
点击查看摘要
Abstract:Recent monocular 3D face reconstruction methods achieve high fidelity by fitting a 3D Morphable Model (3DMM) to dense priors predicted by networks, but the optimization stage is computationally expensive, often taking tens of seconds per image. We present RealDenseFace, a real-time optimization-based 3D face reconstruction method with dense UV-space network predictions. Our key idea is to formulate 3DMM fitting as a nonlinear least-squares problem and solve it with a tailored Gauss-Newton solver that converges in only a few iterations. The reconstruction is conducted in two stages. In the first stage, the network predicts two dense UV-space maps from a single RGB image: a correspondence map for UV-to-image alignment, and a relative-depth map for geometric constraints along the viewing direction. In the second stage, the solver fits per-vertex targets sampled from these maps at the vertex UV coordinates. The solver supports all three reconstruction settings: single-image fitting, offline sequence reconstruction, and online tracking. Our method achieves state-of-the-art accuracy on the NeRSemble SVFR benchmark. The online tracker runs at 80+ FPS, and the offline sequence reconstruction is over 20 times faster than previous optimization-based baselines.
95. 【2608.09233】DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models
链接:https://arxiv.org/abs/2608.09233
作者:Mingfeng Lin,Chengfei Cai,Lin Xu,Yuxiang Wei,Liang Han
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:diverse downstream scenarios, downstream scenarios typically, scenarios typically relies, task-specific optimization objectives, image generation
备注:
点击查看摘要
Abstract:Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.
96. 【2608.09231】BAG: Budget-Aware Gating for Diffusion Caching
链接:https://arxiv.org/abs/2608.09231
作者:Tong Zhao,Mingkun Lei,Yucheng Han,Chi Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:accelerates Diffusion Transformers, Diffusion Transformers, lack instance adaptivity, existing paradigms face, heuristics lack global
备注: 22 pages, 12 figures, and 13 tables
点击查看摘要
Abstract:Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but existing paradigms face a fundamental trade-off: online heuristics lack global budget awareness, whereas static schedules lack instance adaptivity and fail to flexibly adapt to varying runtime budget constraints. To bridge this gap, we present BAG (Budget-Aware Gating), a novel caching policy that unifies global budget pacing with dynamic, instance-adaptive feature reuse. Rather than relying on hand-crafted rules, BAG employs a lightweight gating network that dynamically decides whether to execute a full computation or reuse cached features at each step by jointly conditioning on the budget state and local trajectory feedback. We train this policy via offline-to-online schedule distillation, transferring the decision-making of offline-searched schedules into a compact online gate. Extensive experiments on FLUX.1-dev and Wan2.1 demonstrate that BAG consistently outperforms state-of-the-art caching methods across various speedup tiers while remaining robust across different resolutions, seeds, and guidance scales. Code will be released.
97. 【2608.09226】RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
链接:https://arxiv.org/abs/2608.09226
作者:Yuhan Li,Fangao Zeng,Sicong Kang,Mengfei Xu,Hao Zhou,Wei Li,Pipei Huang,Bingbing Ni
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:typically performed sequentially, performed sequentially, gains during compression, reward gains, based reward alignment
备注:
点击查看摘要
Abstract:Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these procedures are typically performed sequentially, increasing training cost and risking the loss of reward gains during compression. We instead take an RL-native perspective: diffusion RL already generates reward-scored finite-step trajectories, whose intermediate states provide a natural source of distillation supervision rather than a disposable byproduct of sampling. Based on this insight, we propose REST (Reward-Enhanced Scored-Trajectory Distillation), a single-stage RL-distillation co-training framework that attaches a decoupled student to an arbitrary RL teacher. The student learns segment-wise from the teacher's evolving rollout trajectories while leaving the original teacher optimization unchanged. To prevent uniform imitation from preserving undesirable low-reward behaviors, we further introduce Advantage-Modulated Distillation (AMD), which transforms rollout advantages into signed weights over a base distillation loss. AMD strengthens supervision from preferred trajectories and mildly repels the student from low-reward ones. The resulting framework is lightweight and plug-and-play, requires no extra image rollouts, no separate distillation dataset, and no adversarial training. Experiments on compositional generation, visual text rendering, and human-preference alignment show that REST enables few-step CFG-free inference that matches or surpasses its 40-step RL teacher, with an overall additional training cost below 25% over pure RL. REST improves DrawBench PickScore over RTDMD by 0.82 while requiring only one-fifth of the training iterations.
98. 【2608.09223】PatchHead: Learning Spatial Patch Evidence for Generalizable AI-Generated Image Detection
链接:https://arxiv.org/abs/2608.09223
作者:Shengbo Qi,Hongyi Fang,Benjia Zhou,Rui Mao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:detectors generalize poorly, AI-generated image detectors, image detectors generalize, test images originate, generalize poorly
备注:
点击查看摘要
Abstract:AI-generated image detectors generalize poorly when their training and test images originate from different generators or datasets. Despite the rich spatial representations produced by vision foundation models like DINO, existing detectors typically classify images using only the globally aggregated CLS token. We hypothesize that globally aggregating DINO features into a single CLS token obscures spatially distributed generation traces. To test this hypothesis, we introduce PatchHead, a lightweight spatial aggregation head that preserves the two-dimensional organization of DINO patch tokens and integrates evidence across neighboring regions. During training, we freeze the pretrained DINO backbone and optimize only the inserted LoRA adapters, PatchHead, and auxiliary projection head. Across nine cross-dataset benchmarks spanning manually curated and in-the-wild settings, PatchHead ranks first on seven datasets and second on the remaining two. It improves the strongest prior method from 91.6% to 94.6% in average balanced accuracy (+3.0 points) and raises the worst-case accuracy from 82.4% to 89.4% (+6.9 points), while introducing only 8.6% more trainable parameters and 0.08% additional FLOPs. Further qualitative analysis suggests that PatchHead (i) reduces class-conditional domain discrepancy, and (ii) redirects the representation from content-dominated saliency toward spatially distributed authenticity evidence. Together, these observations provide a representation-level account of why spatial patch aggregation transfers more reliably across generators and datasets than a single CLS-based global representation. Our code and models will be made available upon acceptance.
99. 【2608.09221】FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning
链接:https://arxiv.org/abs/2608.09221
作者:Radwan Selo,Majid Kundroo,Taehong Kim
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)
关键词:Federated Learning, enables collaborative model, preserving data privacy, collaborative model training, distributed client devices
备注:
点击查看摘要
Abstract:Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, particularly in terms of label distribution skewness and variations in dataset sizes, which can lead to biased model updates and hinder convergence. To address this, we propose FedTVD, a novel FL algorithm that weights client contributions during aggregation by considering both data quality and quantity. Unlike traditional FL approaches such as FedAvg, which rely solely on dataset size for client weighting, FedTVD integrates Total Variation Distance (TVD) to measure the divergence between each client's local label distribution and a uniform global distribution. Clients with highly skewed distributions receive lower weights, preventing unbalanced datasets with imbalances from disproportionately influencing the global model. At the same time, dataset size is incorporated to ensure scalability and fairness. This dual-weighting mechanism effectively mitigates the impact of data imbalance, leading to more stable and generalized global models. Experimental results show that FedTVD consistently outperforms state-of-the-art methods across all datasets (FMNIST, CIFAR-10, and CIFAR-100) and all levels of data heterogeneity. Notably, it achieves up to 10.6% improvement over FedAvg on CIFAR-10 under highly skewed data, while maintaining top performance even under moderate and IID settings.
100. 【2608.09208】FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning
链接:https://arxiv.org/abs/2608.09208
作者:Van Truong Vo,Khoa Nguyen,Taehong Kim
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC)
关键词:Decentralized intelligence systems, decentralized federated learning, Decentralized intelligence, decentralized federated, limited coordination increasingly
备注:
点击查看摘要
Abstract:Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). However, DFL suffers from convergence inefficiency under data heterogeneity due to the use of a uniform learning rate (LR) that ignores layer-specific optimization needs. Foundational layers are responsible for maintaining network consensus, while specialized layers adapt to local data characteristics, leading to conflicting gradients and degraded performance under non-IID conditions. To address this fundamental tension, this work introduces FedA2L, a method that dynamically adjusts layer-wise LRs based on model divergence signals. By leveraging local update intensity and network consensus constraints, FedA2L seamlessly integrates into existing DFL protocols without additional communication or coordination. Extensive evaluations across DFL algorithms, various model architectures, and datasets demonstrate that FedA2L achieves up to 4.94 times faster convergence than vanilla DFL and reduces communication rounds by up to 59% compared to scheduler-based baselines. Furthermore, FedA2L exhibits resilience to severe data heterogeneity, larger network sizes, and sparse topologies, reducing communication overhead and establishing it as a versatile optimization tool for resource-constrained or large-scale distributed learning in edge and IoT deployments. The code is released at this https URL.
101. 【2608.09200】NBA_Streaming: A Large-Scale Benchmark for Fine-Grained Basketball Commentary Generation in Continuous Streams
链接:https://arxiv.org/abs/2608.09200
作者:Lifang Wu,Yuyang Wu,Yangdong Gao,Fengyu Liu,Ya Jing,Liang Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Live basketball commentary, subsequent events unfold, generation requires determining, Live basketball, requires determining
备注:
点击查看摘要
Abstract:Live basketball commentary generation requires determining when an event is sufficiently observable and describing it before subsequent events unfold. However, existing methods are primarily designed for pre-segmented clips or complete videos, making them unsuitable for continuous streams. Existing datasets also provide limited supervision for player identities, fine-grained actions, event attributes, and coherent event chains, restricting the factual richness of generated commentary. To address these limitations, we introduce NBA_Streaming, a large-scale benchmark for online fine-grained basketball commentary generation. It contains 307 hours of basketball broadcasts and approximately 35K temporally aligned events, with annotations of event boundaries, player identities, fine-grained actions, event chains, and natural-language commentary. By moving from isolated clips to continuous streams, NBA_Streaming enables unified evaluation of event localization, response reliability, factual grounding, and commentary quality under causal constraints. We further propose a causal two-stage framework that combines completion-first localization with ball-centric semantic grounding, enabling the system to identify complete events from observed streams and organize scene, event, identity, and action cues for commentary generation. Extensive experiments reveal the difficulty of NBA_Streaming, where existing baselines struggle with online timing, factual grounding, and fine-grained description. Our framework consistently improves over strong alternatives, while the remaining gap highlights NBA_Streaming as a valuable benchmark for streaming sports video understanding and generation.
102. 【2608.09186】RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing
链接:https://arxiv.org/abs/2608.09186
作者:Hao Li,Ju Dai,Feng Zhou,Mengting Shi,Haofei Wang,Zhen Song,Wei Zhou,Lei Li,Junjun Pan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remains challenging due, translating long-form descriptions, editing remains challenging, fine-grained facial geometry, remains challenging
备注:
点击查看摘要
Abstract:Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text--mesh annotations and paired text--blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at this https URL, and the source code and dataset will be released upon paper acceptance.
103. 【2608.09182】Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction
链接:https://arxiv.org/abs/2608.09182
作者:Jingxian Xu,Yuhao Huang,Rusi Chen,Yanfeng Zhou,Dong Ni
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:quantitative clinical measurement, Accurate landmark localization, Accurate landmark, downstream analysis, medical images
备注: 10 pages, 4 figures, 3 tables. Accepted by MICCAI MLMI 2026
点击查看摘要
Abstract:Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Existing localization methods have advanced, among which multi-stage refinement is a superior solution. Although this strategy mitigates the anatomical ambiguity inherent in single-stage global predictions, its high computational cost limits practical applicability. In this work, we propose a parameter-economic model, PPOC-LL, which leverages Prototype learning-based Progressive Offset Correction for Landmark Localization. Our contribution is three-fold. First, to drive coarse-to-fine landmark optimization, we introduce a multi-scale dynamic perception strategy for patch-level feature pyramid modeling. Second, to effectively handle anatomically similar patterns, we design a similarity-driven prototype learning mechanism that captures informative local semantics for robust offset prediction. Last, to stabilize the model learning and improve the overall performance, we incorporate a novel error-aware reliability regularization via tolerance-based balancing. We collected a large validation cohort, including two public and one private datasets spanning X-ray and ultrasound modalities, covering cephalometric, symphysis-fetal head, and fetal heart landmarks. Extensive experiments demonstrate that PPOC-LL achieves satisfactory performance with a favorable trade-off between accuracy and model complexity.
104. 【2608.09176】Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
链接:https://arxiv.org/abs/2608.09176
作者:Jingbo Wen,Liang He,Mingyu Cao,Haoyu Wang,Minxuan Hu,Kangning Cui,Xilu Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:maximize average accuracy, carry equal cost, fixed compute budget, language models, errors carry equal
备注:
点击查看摘要
Abstract:Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed compute budget, implicitly assuming that all errors carry equal cost. However, the consequence of an incorrect prediction on downstream tasks is rarely symmetric: misreading an invoice amount can be far more costly than misclassifying a background color. Motivated by this, we introduce consequence-sensitive visual token compression, which allocates visual computation across requests according to their potential error costs. Our method follows a calibrate-then-allocate procedure, estimating consequence-specific error-budget curves offline and applying the calibrated token budgets online using consequence signals available from question or task information. On a controlled within-task benchmark, high- and low-consequence questions are drawn from the same document images, so content alone cannot reveal which questions are costly to get wrong. In this setting, our method reduces high-stakes errors from 0.300 to 0.133 under the same total token budget, whereas a content-driven allocator performs no better than uniform allocation. Measuring how error rates change with token budget across different cost ratios, we derive an allocation frontier: uniform allocation is optimal when errors are equally costly, and token transfer toward high-consequence questions becomes increasingly beneficial as the cost gap grows. This allocation principle generalizes well across three dense vision-language benchmarks, two budget realization mechanisms (token deletion and resolution reallocation), two VLM architectures, and multiple token selection strategies. On a realistic mixed workload, consequence-sensitive allocation reduces cost-weighted error by 38% while achieving approximately 21% lower latency than full-resolution inference.
105. 【2608.09152】LightAIR: Lightweight Action Inversion and Riemannian Rectification for Text-based Person Anomaly Search
链接:https://arxiv.org/abs/2608.09152
作者:Yulun Zhang,Zixu Li,Zhiwei Chen,Zhiheng Fu,Wenbo Wang,Zihang Qiu,Zhilin Wang,Ruxin Wang,Yupeng Hu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Traditional Text-based Person, Text-based Person Search, Person Anomaly Search, Text-based Person Anomaly, severely neglecting dynamic
备注: Accepted by ACM MM 2026
点击查看摘要
Abstract:Traditional Text-based Person Search (TPS) is typically limited to matching static appearance attributes, severely neglecting dynamic action information. The Text-based Person Anomaly Search (TPAS) task bridges this gap, requiring models to locate micro-level specific abnormal behaviors while matching macro-level appearance of pedestrians. However, current TPAS methods face fundamental limitations: external explicit pose estimators are fragile in unconstrained surveillance scenarios, and implicit learning encounters visual decoupling failure under pixel-level entanglement, causing dominant appearance information to easily swallow and contaminate subtle action features. Furthermore, performing contrastive optimization on hard negative samples (``same appearance, different actions'') in conventional Euclidean spaces induces severe shortcut learning. To address these, we propose the Lightweight Action Inversion and Riemannian rectification network (LightAIR). First, it introduces textual semantic priors as anchors via a lightweight action inversion operator to extract pure action features, thereby overcoming visual-inherent coupling. Subsequently, it employs orthogonal null-space projection to constrain appearance features within the orthogonal complement space of action features, guaranteeing strict forward decoupling. Finally, we designed a gradient rectification module that computes the Riemannian gradient to constrain the backpropagation trajectory, forcing the gradient flow to update strictly along the tangent space that preserves decoupling properties, thereby cutting off harmful shortcuts. Extensive experiments on the widely used TPAS and TIPR datasets demonstrate that LightAIR significantly outperforms existing state-of-the-art methods. Codes are available at this https URL
106. 【2608.09150】OGG-FR: Orthogonal Gradient Gaming and Frequency Rectification for Unmanned Aerial Vehicle Infrared Image Super-Resolution
链接:https://arxiv.org/abs/2608.09150
作者:Yongsong Huang,Qingzhong Wang,Xiaofeng Liu,Tomo Miyazaki,Yaohou Fan,Shinichiro Omachi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Unmanned aerial vehicle, infrared image super-resolution, image super-resolution aims, Unmanned aerial, recover weak thermal
备注: This manuscript is currently under peer review. Copyright may subsequently be transferred to the publisher, after which the availability of this version may be subject to the publisher's policy
点击查看摘要
Abstract:Unmanned aerial vehicle (UAV) infrared image super-resolution aims to recover weak thermal structures for deployment on resource-constrained platforms; lightweight models are therefore preferred, but multi-loss training can be unstable. A common strategy combines pixel-domain and frequency-domain objectives; however, low contrast, limited high-frequency content, and sensor-specific noise often make their gradients weakly aligned or conflicting. To address this optimization ambiguity, we propose Orthogonal Gradient Gaming and Frequency Rectification (OGG-FR), a plug-and-play optimization framework that decomposes the frequency gradient into a redundant parallel component and an orthogonal innovation component relative to the pixel gradient. In the conflict regime, OGG-FR computes a safe base gradient using the Multiple Gradient Descent Algorithm (MGDA) and adds a variance-rectified orthogonal innovation; in the compatible regime, it discards redundant parallel information and injects the orthogonal innovation according to a confidence score estimated from the high-frequency residual. Experimental results on the UAV thermal benchmark show broad gains under BI and BD degradations at $\times 4$ and $\times 8$ scales, while gradient analyses support the effectiveness of the proposed conflict-aware update rule.
107. 【2608.09147】RefineAny3D: Depth Refinement as Semantic Alignment for Monocular 3D Detection
链接:https://arxiv.org/abs/2608.09147
作者:Zhihao Zhang,Gengwei Zhang,Tianlong Chen,Xiaoming Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:depth foundation models, leveraging depth foundation, fixed category vocabulary, depth foundation model, localize arbitrary categories
备注:
点击查看摘要
Abstract:Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector's predicted depth degrades accuracy, even falling below the detector's own prediction. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM's vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.
108. 【2608.09146】Multi-Submap Implicit Neural SLAM with Local-to-Global Loop Closure for Large-Scale Scene Reconstruction
链接:https://arxiv.org/abs/2608.09146
作者:Tianchen Deng,Chongdi Wang,Nailin Wang,Lei Zhao,Ziqi Ma,Tianjun Zhang,Zhe Liu,Danwei Wang,Hesheng Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Neural Radiance Fields, Radiance Fields, complex environments remains, accumulated trajectory drift, environments remains challenging
备注:
点击查看摘要
Abstract:Neural Radiance Fields (NeRF)-based SLAM has demonstrated impressive results in small-scale scene reconstruction, yet scaling these methods to extensive, complex environments remains challenging due to catastrophic forgetting and accumulated trajectory drift. This paper presents a robust, large-scale neural SLAM system featuring a multi-submap architecture and a dual-tier loop closure mechanism. Specifically, we propose a progressive mapping strategy that dynamically allocates neural submaps to maintain high-fidelity representations without memory explosion. For robust pose estimation, an optical-flow-based tracking module is integrated to handle aggressive motions. To address global consistency, we introduce a local-to-global loop closure framework leveraging the foundation model for high-performance global descriptor extraction, significantly enhancing relocalization accuracy under varying viewpoints. Furthermore, an inter-submap online distillation algorithm is designed during back-end optimization to enforce geometric and appearance consistency across overlapping submap boundaries. To validate the system, we developed a customized handheld mechatronic platform and conducted extensive evaluations on both public benchmarks and our large-scale indoor-outdoor datasets. Experimental results, including direct deployment on an onboard computing unit, demonstrate that our approach outperforms state-of-the-art neural SLAM methods in reconstruction quality and localization robustness, providing a scalable solution for real-world robotic perception and digital twinning. We will release the code publicly on \href{this https URL}{this https URL} .
109. 【2608.09145】Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images
链接:https://arxiv.org/abs/2608.09145
作者:Yongsong Huang,Xiaofeng Liu,Tomo Miyazaki,Yaohou Fan,Shinichiro Omachi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:General-purpose multimodal large, multimodal large language, General-purpose multimodal, large language models, multimodal large
备注: This manuscript is currently under peer review. Copyright may subsequently be transferred to the publisher, after which the availability of this version may be subject to the publisher's policy
点击查看摘要
Abstract:General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy alone. However, a correct answer does not ensure that the model's explanation is grounded in infrared thermal evidence. We introduce an explanation-aware evaluation framework that separates answer correctness, output-level explanation groundedness, and thermal grounding for infrared visual questions. Using a Dual-LLM Consensus Judge with a preliminary human-anchor calibration check, we find that correct answers can still rely on weak or visible-light evidence; withholding the original infrared image and showing only a visible-like rendering erodes thermal grounding with little accuracy change; and this erosion is observed most strongly for more capable models but disappears when infrared remains available. We further propose Thermal-Grounded Feedback (TGF), a training-free feedback loop that diagnoses explanation-side failures and revises the explanation while preserving the selected answer. On local paired-input validation, TGF improves explanation-side grounding without changing answers. These findings suggest that future trustworthy MLLMs for infrared scene understanding should be evaluated and developed to produce thermally grounded explanations rather than merely accurate answers.
110. 【2608.09143】UniMoFlow: Grounding Instruction-Driven 3D Human Motion Editing in Generation
链接:https://arxiv.org/abs/2608.09143
作者:Yilei Hua,Beibei Jing,Ce Zheng,Hanyu Zhou,Yawei Luo,Wei Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:precise spatiotemporal localization, requires precise spatiotemporal, rich semantic grounding, human motion requires, motion requires precise
备注: 18 pages, including supplementary material; 8 figures and 7 tables. Code: [this https URL](https://github.com/Yilei-Hua/UniMoFlow) . Submitted to AAAI 2027
点击查看摘要
Abstract:Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity. To overcome this bottleneck, we ground motion editing directly within text-to-motion generation across data, architecture, and inference. At the data level, we develop a closed-loop synthesis-and-verification pipeline that produces Omni-MoEdit, a large-scale dataset spanning body-part, amplitude, temporal, action, and style edits. At the architectural level, we introduce UniMoFlow, a unified latent flow-matching model that shares broad semantic and kinematic knowledge between generation and editing. At the inference level, SAFE (Source-Anchored Flow Editing) complements UniMoFlow with controllable, source-anchored refinement. Furthermore, we augment standard evaluations with semantics-aware metrics to account for valid edits that inherently deviate from a single ground-truth reference. Extensive experiments demonstrate improved target-text alignment, edit effectiveness, and cycle consistency, while maintaining competitive source fidelity and text-to-motion generation quality.
111. 【2608.09139】CodecArena: Codec Quality Assessment via Visual Reinforcement Learning
链接:https://arxiv.org/abs/2608.09139
作者:Jiaye Fu,Weiqi Li,Qiankun Gao,Yanchen Zhao,Xiandong Meng,Jian Zhang,Siwei Ma,Jiaqi Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:video generation models, jointly optimized neural, optimized neural networks, generation models, low and ultra-low
备注: The project page is: [this https URL](https://jyfu-vcl.github.io/codecarena)
点击查看摘要
Abstract:Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.
112. 【2608.09137】Bright-Channel Retinex Enhancement with a Conditional Overdispered-Noise Analysis
链接:https://arxiv.org/abs/2608.09137
作者:Jongpil Jeong
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:training-free low-light enhancement, combines local bright-channel, bright-channel illumination estimation, low-light enhancement method, local bright-channel illumination
备注: 5 pages, 3 figures, 2 tables
点击查看摘要
Abstract:I present a training-free low-light enhancement method that combines local bright-channel illumination estimation, Retinex division, and edge-preserving denoising. For a fixed illumination estimate, a conditional Negative -Binominal psueduo-count method characterises the heteroscedastic noise amplified by division. The unconstrained reflectance ratio is the pixelwise maximum-likelihood estimate, with a boundary solution for zero-valued observations; the implemented estimate additionally applies illumination filtering and range clipping. The NB model is a diagnostic noise analysis rather than a calibrated sensor model, and the final fixed-bandwidth bilateral filter is an empirical approximation rather than the exact Bayesian solution. On the LOL-v1 dataset, the methodobtains mean PSNR/SSIM of 17.74dB/0.739, the highest values among the evaluated with conventional methods. A 400X600 image is processed at approximately 43 FPS on an Apple M2 Pro CPU.
113. 【2608.09133】When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution
链接:https://arxiv.org/abs/2608.09133
作者:Yu Shi,Yuyao Zhang,Yu-wing Tai
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:remarkable perceptual quality, observation remains challenging, recently achieved remarkable, achieved remarkable perceptual, perceptual quality
备注:
点击查看摘要
Abstract:Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging. In particular, we observe that diffusion transformers (DiTs) built on latent representations suffer from a critical limitation: the compression bottleneck of the VAE weakens fine-grained spatial information, leading to hallucinated details that are weakly grounded in the input image. In this work, we revisit generative SR from a representation perspective and propose a pixel-grounded super-resolution (PGSR) framework that preserves LR-observed pixel evidence before VAE compression and reuses it throughout restoration. Instead of relying solely on the compressed latent condition, PGSR extracts pre-VAE pixel evidence from the upsampled LR image and reuses it at two stages. First, Condition-Side Trajectory Guidance fuses LR-derived pixel evidence with the latent LR condition to guide the latent restoration trajectory. Second, Decoder-Side Pixel Grounding injects multi-scale pixel features into the frozen VAE decoder to ground the final rendering with LR-observed cues. To efficiently adapt large pretrained DiT models, we keep the latent autoencoder and main flow-matching backbone frozen, and train only lightweight restoration modules. We further study an efficient local-window attention variant for improved high-resolution efficiency and scalability. Extensive experiments demonstrate that PGSR improves the realism--fidelity trade-off and produces more faithful, visually convincing results than existing latent generative SR approaches.
114. 【2608.09122】Visual Distortion Detection in UGC Images Using Large Multimodal Models
链接:https://arxiv.org/abs/2608.09122
作者:Ziheng Jia,Yingji Liang,Jiaying Qian,Xiongkuo Min
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Toggle, textit, Toggle Hugging Face, distortion detection, distortion
备注:
点击查看摘要
Abstract:The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal models (LMMs) predominantly rely on text-driven supervised fine-tuning (SFT). However, this training paradigm exhibits notable limitations in detection accuracy. Moreover, synthetically distorted images, which are often used as the primary training data source, show a significant generalization gap when deployed in real-world scenarios; thus, the \textbf{synthetic-to-authentic (\textit{S2A})} problem represents a critical challenge. Motivated by these issues, we propose \textbf{\textit{VIGIL}}, which leverages the LMM architecture for precise visual distortion detection. From a candidate pool of over 1000K samples, we construct the \textbf{\textit{VIGIL-140K}} training set, which consists of over 140K distorted images. These images are obtained through rigorous quality filtering and carefully crafted distortion injection, covering 8 major synthetic distortion categories. Our model leverages different layers of the large language model (LLM) decoder, treating them as \textit{multiple detectors} that perform synchronous distortion detection using multi-level features. Additionally, we retain distortion cues from predictions assigned to the non-distortion class, which helps mitigate the ambiguous foreground-background (\textit{FG-BG}) separation commonly encountered in the \textit{S2A} problem. After post-processing, our model consistently outperforms strong baselines on both in-domain synthetic distortion detection and \textit{S2A} tasks.
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as:
arXiv:2608.09122 [cs.CV]
(or
arXiv:2608.09122v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.09122
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Ziheng Jia [view email] [v1]
Mon, 10 Aug 2026 05:00:12 UTC (16,150 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Visual Distortion Detection in UGC Images Using Large Multimodal Models, by Ziheng Jia and 3 other authorsView PDFHTML (experimental)TeX Source
view license
Current browse context:
cs.CV
prev
|
next
new
|
recent
| 2026-08
Change to browse by:
cs
cs.AI
References Citations
NASA ADSGoogle Scholar
Semantic Scholar
export BibTeX citation
Loading…
BibTeX formatted citation
loading…
Data provided by:
Bookmark
checked="checked"class=“labs-tab-input”>
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv’s community? Learn more about arXivLabs.
Which authors of this paper are endorsers? |
Disable MathJax (What is MathJax?)
mathjaxToggle();
We gratefully acknowledge support from
our major funders,
member institutions, ,
and all contributors.
About
Help
Contact
Subscribe
Copyright
Privacy
Accessibility
Operational Status (opens in new tab)
Major funding support from
115. 【2608.09110】View-Adaptive Renderer for View-Consistent 2D-to-3D Generation
链接:https://arxiv.org/abs/2608.09110
作者:U-Chae Jun,Jaeeun Ko,Jiwoo Kang
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:single image remains, Neural Radiance Field, computer vision, remains a fundamental, fundamental yet challenging
备注:
点击查看摘要
Abstract:Reconstructing 3D shapes from a single image remains a fundamental yet challenging problem in computer vision. Traditional monocular 3D generation pipelines typically synthesize multiple views from a single input image before applying Neural Radiance Field (NeRF)-based reconstruction. However, inherent projective ambiguities often produce visual discontinuities across generated viewpoints, leading to inaccuracies in reconstructed 3D models. Current solutions either incur significant additional computational burdens or fail to adequately resolve practical inconsistencies between synthesized views. To address these limitations, we propose a novel viewpoint-adaptive neural rendering framework that enables robust 3D reconstruction even when given partially inconsistent multi-view inputs. Our approach introduces view-adaptive neural renderers that independently correct viewpoint-dependent errors while simultaneously sharing a global feature backbone to preserve structural coherence. Furthermore, we propose a self-attention fusion module that adaptively integrates multi-view information, ensuring geometric consistency without relying heavily on indirect regularizations or computationally intensive methods. Through extensive experiments, we demonstrate that our method consistently improves 3D reconstruction fidelity. Importantly, our approach achieves near state-of-the-art performance without diffusion-based SDS supervision, relying primarily on photometric rendering loss with lightweight attention regularizers. This balance between accuracy and efficiency makes the proposed framework highly practical for real-world applications.
116. 【2608.09101】Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation
链接:https://arxiv.org/abs/2608.09101
作者:Shuaishuai Cao,Shuwei Peng,Meng Tang,Min Huang,Youjin Wang,Jie Chen,Jing Ouyang,Zhiwei Zhai
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Semantic segmentation models, high overlap scores, Semantic segmentation, Contrastive Mask Fidelity, high overlap
备注:
点击查看摘要
Abstract:Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox. We introduce Contrastive Mask Fidelity (CMF), a training-free, reference-free metric that scores competing class masks directly against image evidence. CMF composites keep and erase counterfactual views of each mask and asks a frozen vision-language judge whether class evidence is concentrated inside the mask and absent outside. We validate CMF on controlled mask corruptions, then audit 10,731 image-class pairs across ten remote-sensing benchmarks using candidate masks from Seg-Probe, a training-free open-vocabulary probe built on SegEarth-OV3 that outperforms prior baselines on nine of ten datasets. The audit reveals systematic, class-dependent annotation distortion: man-made classes such as buildings, roads, and cars favor the candidate mask on 62-85% of pairs, whereas ambiguous land cover more often favors human annotations. On a blinded three-annotator consensus, CMF matches expert judgment on 81% of pairs, exceeding keep-only scoring, model confidence, and a trained label-quality baseline. Finally, conservative class-wise arbitration yields supervision that improves cross-domain transfer over raw annotations and matched replacement controls, positioning CMF as a scalable tool for auditing ground truth rather than presuming it infallible.
117. 【2608.09100】Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition
链接:https://arxiv.org/abs/2608.09100
作者:Yani Guan,Dengpan Dong,Zi Wei,Shuang Luo,Dan Hannah,Yumin Zhang,Kang Xu
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:scale requires reading, Millions of chemical, information at scale, scale requires, requires reading
备注:
点击查看摘要
Abstract:Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR appears nearly solved on synthetic images yet remains difficult on real documents: the starting recognizer, Qwen2.5-VL-7B, exceeds 91% accuracy on synthetic renders but falls below 16% on three real-world benchmarks (ACS, CLEF-IP, USPTO). To identify the main source of improvement, 21 recognizers were fine-tuned on mixtures of synthetically rendered structures and labeled real depictions from patents, journal figures, and hand-drawn collections, varying the vision language model (VLM) base, the fraction of real training data, and the vision-tower adaptation strategy. Labeled real training images make the largest difference. For Qwen2.5-VL, ACS exact match rises from 0.15 with no real data to 0.37 at 9.5% and 0.46 at 50.2%; a controlled experiment across three base models reproduces the trend. A vision-tower LoRA, in contrast, does nothing for Qwen (+0.00, paired p=1.00), substantially helps InternVL3-8B (+22.8 to +34.6 pt), and modestly helps GLM-4.1V-9B (+1.0 to +9.6 pt), so its value depends on the base model. The best configuration reaches 0.96 exact match on clean renders and 0.49, 0.65, 0.84, and 0.76 on ACS, CLEF-IP, UOB, and USPTO, respectively. Gaps between base models are largest without real data (0.21), shrink to 0.06 at 70% real data, and reorder the ranking; base model and real-data mixture must therefore be selected together. Small-scale experiments on handwritten image-to-LaTeX recognition and chart-to-table conversion show that base-model rankings also vary beyond chemistry. More generally, model and adaptation choices for visual structure recognition should be evaluated on the target task.
118. 【2608.09097】SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision
链接:https://arxiv.org/abs/2608.09097
作者:Weixin Ye,Wei Wang,Hongguang Zhu,Xuecheng Nie
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:fine-grained local deformations, achieving pixel-level precision, persistent challenge, Large Language Models, rapid advances
备注: accepted by ACM MM 2026
点击查看摘要
Abstract:Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions. To address this issue, we first introduce **SI-Data**, a high-quality dataset specifically designed for instruction-guided local sketch editing. We develop an automated pipeline leveraging Multimodal Large Language Models (MLLMs) to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images. By providing both reliable spatial anchors and explicit semantic intent, SI-Data uniquely enables collaborative spatial-semantic learning. Building upon this, we propose a collaborative framework called **SI-Edit** that integrates semantic instructions with precise geometric constraints. Furthermore, to address the lack of standardized evaluation, we establish a comprehensive set of metrics designed to measure both structural fidelity (e.g., sketch-to-edge alignment) and semantic adherence. Experimental results demonstrate that SI-Edit provides more reliable structural control than baselines for sketch-based image editing, and achieves precise, pixel-level local refinements aligned with user intent. The data and code are released on the [project page](this https URL).
119. 【2608.09091】LDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset
链接:https://arxiv.org/abs/2608.09091
作者:Jing Ning,James D. Braza
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:limited training data, transfer learning dataset, Transfer learning, fine tuning dataset, learning dataset
备注:
点击查看摘要
Abstract:Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO . Qualitatively, it seems a transfer learning dataset should have both more classes and more examples per class than the fine tuning dataset; however, a quantitative method to choose the best transfer learning dataset does not currently exist. In this paper, we design TLDChoiceNet, a model to choose the best transfer learning dataset given a fine tuning dataset by predicting the test-set accuracy after fine-tuning. A simple version 1 achieves 0.154 MSE on the test dataset, while a version 2 leveraging an ImageNet pre-trained ResNet50 v2 embedding with per-class information attains a 5X lower MSE of 0.031. We further design two metrics that enable an unsupervised method of choosing an optimal transfer learning dataset: distribution distance (DD), which linearly regresses against fine-tune accuracy with an R2 of 0.89, and average class correlation (ACC), which improves the R2 to 0.97. Our results underscore that a dataset's low-level statistics can explain the transfer learning effect, and that using a pre-trained ImageNet can embed different classes further apart in latent feature space.
120. 【2608.09083】Learning human joint torques from pixels
链接:https://arxiv.org/abs/2608.09083
作者:Chen Chen,Rui Cheng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:real-world movement scenarios, Estimating human joint, Estimating human, movement scenarios, key step
备注:
点击查看摘要
Abstract:Estimating human joint torques from visual observations is a key step toward bringing biomechanical analysis from controlled laboratories to real-world movement scenarios. Existing torque estimation methods typically depend on surface electromyography, motion-capture markers, force plates, or simulated imitation data, which limits their applicability to ordinary RGB images. In this work, we introduce VID, a vision-based inverse dynamics dataset and benchmark for predicting human joint torques directly from real monocular images. VID contains 63,369 synchronized frames with real human images, kinematic annotations, anthropometric attributes, and OpenSim-derived dynamic labels, providing paired visual and biomechanical supervision for real-image inverse dynamics. We further define a standardized evaluation protocol covering overall torque estimation, joint-specific analysis, and action-specific prediction. To establish a strong reference model, we propose VID-Network, which combines pose-pretrained spatial probabilistic features, marker regression, and temporal torque inference to recover joint torques from image sequences. Experiments on VID show that VID-Network achieves an overall mPJE of 1.7612 N$\cdot$m/kg, improving over the best compared baseline by 39.81\%, and obtains the lowest error across all evaluated joint types and most action categories. VID establishes a first practical benchmark for vision-driven human inverse dynamics and provides a foundation for studying biomechanical inference in less constrained environments.
121. 【2608.09057】Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective
链接:https://arxiv.org/abs/2608.09057
作者:Hongyi Fang,Chuwen Xie,Benjia Zhou,Yu-Xuan Qiu,Chenggong Hu,Zhibin Wang,Chao Chen,Jianbin Qin,Rui Mao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:powerful generative paradigm, producing high-quality images, Next-scale visual autoregressive, Next-scale visual, generative paradigm
备注:
点击查看摘要
Abstract:Next-scale visual autoregressive models (VARs) have emerged as a powerful generative paradigm, producing high-quality images through efficient coarse-to-fine prediction. However, their potential for text-guided image editing remains largely underexplored. Existing training-free VAR editing approaches often formulate editing as target-conditioned regeneration guided or constrained by the source image, and may rely on inversion, test-time optimization, attention control, or user-provided masks. This generation-centric formulation does not fully exploit the multiscale source representations provided by VARs and may introduce additional computation or intervention. We instead take a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes. Based on this perspective, we propose \textbf{EditMod}, which compares source- and target-conditioned predictions under a shared autoregressive context, treats their difference as a scale-wise editing direction, and applies it as a residual update to source tokens at selected scales. Experiments show that EditMod achieves leading source-image fidelity while maintaining strong text alignment, and completes end-to-end editing of a 1K image in only 1.57 seconds on a single A100 GPU without per-image preparation.
122. 【2608.09052】riple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation
链接:https://arxiv.org/abs/2608.09052
作者:Xuanyu Liu,Zheng Fang,Hongyang He,Yundi Hong,Daizong Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:vision foundation models, updates lightweight modules, foundation models, commonly freezes, vision foundation
备注: Accepted for publication at the British Machine Vision Conference (BMVC) 2026. Official list of accepted papers: [this https URL](https://bmvc2026.bmva.org/programme/accepted_papers/)
点击查看摘要
Abstract:Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such as LoRA. However, pseudo-labels have mixed reliability, and a single LoRA adapter must absorb reliable, ambiguous, and noisy gradients in the same low-rank space. This can make VFM adaptation sensitive to pseudo-label noise. We propose \textbf{TriNoL}, a \textbf{Tri}ple-expert learning framework from \textbf{No}isy \textbf{L}abels for semi-supervised VFM adaptation. TriNoL routes unlabeled samples into three confidence regions and assigns them to three LoRA experts: a Positive Expert for high-confidence pseudo-labels, an Alignment Expert for medium-confidence ambiguous samples, and a Negative Expert for low-confidence noisy samples. The VFM backbone remains frozen, and only the LoRA experts and classifier head are updated. By separating different pseudo-label reliability regions into specialized adaptation paths, TriNoL improves robustness to noisy supervision while keeping the training cost low.
123. 【2608.09048】GeoAI-based post-segmentation quality validation of building footprints via spatial feature engineering
链接:https://arxiv.org/abs/2608.09048
作者:Shah Imran Ahsan Chowdhury,Kazi Jihadur Rashid,Rajsree Das Tuli,Rahul Saha,Bulbul Ahammad
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Deep learning-based building, produces topologically inconsistent, inconsistent vectors unfit, topologically inconsistent vectors, Deep learning-based
备注: 19 pages, 10 figures, 8 tables
点击查看摘要
Abstract:Deep learning-based building footprint extraction from high-resolution imagery often produces topologically inconsistent vectors unfit for direct GIS database ingestion. To address this, we present a multidomain GeoAI quality control framework that automates error detection to systematically purify vector footprint databases. Candidate footprints were generated across five UAV survey sites in Bangladesh using U-Net (ResNet-34) and SAM-LoRA (ViT-B). The extracted raster masks were vectorized, geometrically regularized, and consolidated under a spatial-exclusivity constraint to eliminate duplicate representations. We used twenty-four predictors capturing geometric, spatial-contextual, and raster-derived spectral and texture properties. Machine Learning (ML) classifiers were trained on a development partition (Sites B-D) and rigorously validated on a spatially independent test set (Site E) excluded from hyperparameter tuning and class balancing. The experimental results demonstrate that geometric and spatial-contextual predictors using Decision Tree (DT) provide the most effective discriminatory evidence for identifying object-level boundary deformations. DT achieved an accuracy of 95.31%, an F1-score of 91.06%, and a Matthews correlation coefficient (MCC) of 0.880 on the unseen testing site. At the database level, this framework successfully identified 87.34% of erroneous footprints while maintaining 98.31% of acceptable structures, reducing the residual error proportion from 27.32% to 4.62% and improving final database purity to 95.38%. This translates into a relative error reduction of 83.09%. The findings indicate that post-segmentation object-level ML provides a highly transferable, robust mechanism for automated quality assurance in production-ready geographic information system (GIS) workflows.
124. 【2608.09047】Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks
链接:https://arxiv.org/abs/2608.09047
作者:Yi Yang,Xiaoke Chen,Jinyang Huang,Feng-Qi Cui,Yu-Tong Guo,Jia-Cheng Zhao,Haiming Jin,Xiaokang Zhou,Meng Li
类目:Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV)
关键词:compromise training data, model retains clean, triggered inputs, model retains, predicts an attacker-chosen
备注:
点击查看摘要
Abstract:Backdoor attacks compromise training data so that a model retains clean accuracy but predicts an attacker-chosen target on triggered inputs. At very low poisoning rates, only a few samples convey the trigger--target association, making poison-sample selection critical. Existing methods typically rank candidates using per-sample scores, which can select redundant samples from similar semantic regions, and many require task-specific surrogate training. We propose Distributional Feature Coverage Sample Selection (DFCS), a training-free, trigger-agnostic method that clusters fixed pretrained features into one region per poisoning slot and selects the centroid-nearest sample from each region. A local first-order analysis relates this allocation to feature-coverage and representative-mass terms. Across BadNets and Blended attacks on CIFAR-10, Tiny-ImageNet, and Imagenette, DFCS achieves the highest mean attack success rate among seven selectors in all six dataset--attack settings, averaging $96.30\%$ and exceeding the strongest comparator in each setting by 4.60 percentage points on average while preserving clean accuracy. These results support distributional feature coverage as an effective selection principle for low-budget dirty-label backdoor attacks.
125. 【2608.09045】Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production
链接:https://arxiv.org/abs/2608.09045
作者:Xiao Liu,Shiwei Gan,Yafeng Yin,Jiaxin Yin,Bowen Guo,Yaqi Sun,Zhiwei Jiang,Lei Xie
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:sign language recognition, sign language, language recognition, Recent advances, sign language understanding
备注:
点击查看摘要
Abstract:Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolated sign language recognition (ISLR), continuous sign language recognition (CSLR), and sign language translation (SLT), within a single framework, leading to substantial progress. Meanwhile, sign language production (SLP), which generates sign sequences from text, has also attracted growing attention. This naturally raises an important question: can sign language understanding and production be unified within a single framework? Compared with unifying SLU subtasks, this problem is substantially more challenging. Existing SLU tasks largely share the same direction of mapping, namely from sign inputs to linguistic outputs, whereas SLT and SLP lie in opposite directions of sign-text mapping. A unified framework must therefore address two key challenges: (1) bridging the modality gap between continuous sign motions and discrete text tokens through a shared sign tokenizer that supports both linguistic abstraction and motion reconstruction; and (2) learning a single conditional autoregressive model that can take either sign or text as input and generate the corresponding target sequence in the opposite modality. To this end, we propose Uni-SLTP, a unified framework for SLT and SLP with two key components: (1) a shared sign tokenizer that converts sign sequences into discrete tokens and latent representations, capturing both semantic and reconstructive information; and (2) a unified autoregressive generation model that formulates both tasks as conditional sequence generation. Experiments on widely used public datasets show that Uni-SLTP achieves superior motion accuracy for SLP while maintaining competitive SLT performance.
126. 【2608.09006】SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
链接:https://arxiv.org/abs/2608.09006
作者:Shiwei Gan,Xiao Liu,Yafeng Yin,Zhiwei Jiang,Bowen Guo,Lie Xie,Sanglu Lu,Hongkai Wen
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Large Language Models, achieved remarkable success, Large Language, Sign Language Translation, Language Translation
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Translation (GFSLT) remains a challenge. In this paper, we investigate how to effectively adapt LLMs to the GFSLT task. We show that there are two key issues that need to be solved: (1) the inherent distributional gap between visual feature inputs and text feature inputs makes it difficult for LLMs to interpret visual inputs; and (2) existing approaches typically concatenate visual and textual features in an autoregressive framework, which leads to the model overemphasizing textual inputs and deprioritizing visual cues, as LLMs are pretrained predominantly on text-centric data. To address the first challenge, we propose a simple yet effective method named Filtered Pseudo-Gloss CTC Pretraining, which leverages filtered pseudo-gloss sequences generated from text sequences to supervise the training of the visual backbone. To tackle the second issue, we introduce a Visual-Prioritized Distillation training strategy. Specifically, we define a visual-only prediction path in which text inputs are masked, and the model is required to generate the target sequence relying solely on visual inputs. To guide this path, the outputs from the standard visual-textual prediction are then distilled into the visual-only prediction path, encouraging the model to prioritize visual features. Comprehensive experiments and qualitative analyses demonstrate the effectiveness of the proposed model. The proposed SignLlama achieves very competitive performance on multiple datasets for GFSLT tasks, without using any extra modalities or external sign language datasets for pretraining.
127. 【2608.08999】DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models
链接:https://arxiv.org/abs/2608.08999
作者:Chen-Hsiu Huang,Mario Köppen,Ja-Ling Wu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:AI-generated images produced, raised critical concerns, Latent Diffusion Models, Diffusion Implicit Model, infringement and misinformation
备注: accepted by APSIPA ASC 2026
点击查看摘要
Abstract:The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infringement and misinformation. Although existing frequency-domain watermarking methods embed handcrafted geometric patterns into the initial latent noise prior to generation, they suffer from limited capacity and rigid pattern designs. We propose DeepFreqMark, an end-to-end learnable frequency-domain watermarking framework that replaces manual pattern engineering with a neural message encoder and decoder. To circumvent the computational bottleneck caused by Denoising Diffusion Implicit Model (DDIM) inversion during training, we introduce a Spherical Linear Interpolation (Slerp)-based attack simulation. This approach operates directly on the noise latent while strictly preserving the Gaussian variance. Extensive experiments demonstrate that DeepFreqMark achieves significantly lower Bit Error Rates (BER) than baseline methods under real-world attacks and scales to 256 bits message capacity. Our source code is available at this https URL.
128. 【2608.08977】Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach
链接:https://arxiv.org/abs/2608.08977
作者:Parisa Farmanifard,Arun Ross
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:lenses, contact lenses, clear, iris recognition, clear contact
备注:
点击查看摘要
Abstract:This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic or patterned contact lenses has been extensively studied under the presentation attack detection (PAD) paradigm, clear prescription contact lenses, that are typically transparent, have received comparatively less attention despite their widespread use. Unlike patterned lenses, clear lenses introduce no salient texture artifact, making them difficult to detect and are often assumed to have no impact on iris recognition. We first examine this assumption using the commercial VeriEye matcher on four benchmark datasets and show that clear lenses marginally degrade genuine match scores and increase verification error. We then propose a two-stage contact-lens detection framework. Stage~1 uses an existing PAD model to identify patterned lenses, while Stage~2 focuses on the more challenging clear-lens versus no-lens distinction using a ConvNeXt-Base model equipped with Mask-Guided Spatial Attention (MGSA). The proposed MGSA module incorporates a Hough-derived anatomical ROI mask together with learned spatial attention and Squeeze-and-Excitation channel recalibration, allowing the network to focus on subtle limbal cues associated with clear lens wear. Across four datasets, the full pipeline consisting of both patterned and clear contact lens detection achieves between 90.0\%--98.8\% accuracy. Finally, we introduce a z-score calibration method that adjusts VeriEye match scores when a clear lens is detected in the input images. This calibration reduces EER by 4.1\%--28.3\% across datasets, demonstrating that reliable clear contact lens detection can directly improve iris verification performance.
129. 【2608.08965】CoRe-UIE: Rethinking Coexisting and Region-wise Degradation for Underwater Image Enhancement
链接:https://arxiv.org/abs/2608.08965
作者:Weifeng Kong,Chenghao Xu,Lin Chen,Ziheng Cao,Guanying Huo
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:including color distortion, Underwater Image Enhancement, scattering haze, texture attenuation, Underwater images
备注: 9 pages, 5 figures
点击查看摘要
Abstract:Underwater images often suffer from diverse and coexisting degradations, including color distortion, scattering haze, texture attenuation, and uneven illumination. These degradations vary across regions and may coexist locally, making conventional uniform restoration difficult to adapt to different degradation patterns. To address this problem, we propose Coexisting and Region-wise Degradation for Underwater Image Enhancement (\textbf{CoRe-UIE}), a degradation-oriented expert collaboration framework. CoRe-UIE combines a content-preserving shared expert with four shared-backbone routed experts for color correction, scattering suppression, texture recovery, and illumination protection. The routed experts share the same architecture but have independent parameters, and are assigned to different regions through input-derived degradation cues and region-adaptive Top-\(k\) routing. We further introduce a Hilbert--Schmidt Independence Criterion (HSIC)-based representation constraint to reduce statistical dependence among expert features and alleviate redundant expert responses. Experiments on UIEB, LSUI, and U45 demonstrate that CoRe-UIE achieves competitive quantitative performance and visually balanced enhancement under diverse underwater degradation conditions.
130. 【2608.08963】Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
链接:https://arxiv.org/abs/2608.08963
作者:Sarah Rastegar,Mina Ghadimi Atigh,Pascal Mettes,Yuki M. Asano,Cees G. M. Snoek
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Generalized Category Discovery, Category Discovery aims, unlabeled data, aims to recognize, Category Discovery
备注: Accepted by ECCV 2026
点击查看摘要
Abstract:Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typically based on self-supervision and contrastive learning, often struggle to capture fine-grained distinctions, relying on superficial visual cues rather than the intrinsic attributes humans use for categorization. We introduce Fourier Self-Supervision, that leverages the Fourier transform of images to enhance the discrimination of subtle differences and support the discovery of new categories. Our method employs a dual frequency filtering strategy: a low-pass filter first extracts broad, abstract attributes that capture high-level category information, while a high-pass filter emphasizes fine details such as edges and textures that are essential for fine-grained recognition. Each operates on a dedicated latent space, and their overlapping representations together yield a richer, more complete feature space. This dual-frequency approach not only refines feature extraction to identify novel categories, but also strengthens the model's discriminative power in fine-grained category discovery. Experiments on multiple fine-grained datasets show that incorporating Fourier Self-Supervision outperforms state-of-the-art methods, even when the number of classes is unknown, demonstrating its effectiveness for Generalized Category Discovery. Our code is available at: this https URL.
131. 【2608.08957】RMR-Net: Degradation-Evidence-Guided Road-Image Restoration for Defect Detection
链接:https://arxiv.org/abs/2608.08957
作者:Amir Ghorbani,Amirali K. Gostar,WeiQin Chuah,Vahid Ghorbani,Aidan Blair,Alireza Bab-Hadiashar
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Vehicle-mounted road cameras, erase thin cracks, pothole boundaries needed, Vehicle-mounted road, poor illumination
备注: Submitted to ICCAIS 2026
点击查看摘要
Abstract:Vehicle-mounted road cameras are vulnerable to motion blur, defocus, poor illumination, and noise, which can erase thin cracks and pothole boundaries needed by road defect detectors. This paper presents RMR-Net, a compact task-aware restoration front end that estimates degradation evidence from the image, optionally fuses it with existing corruption context/parameters, conditions lightweight restoration blocks, and returns high-frequency pavement detail through a bounded residual path. The experimental scope is deliberately controlled: the conditioning information used on the Image and Vision Computing New Zealand (IVCNZ) pothole dataset and the Road Damage Dataset: Potholes, Cracks and Manholes (PCM) consists of saved synthetic-generator parameters, not measured vehicle telemetry. A clean-trained, frozen YOLO11s detector evaluates every image source. Across eight held-out degradation conditions, RMR-Net obtains the highest mAP50 in seven, including 0.140-0.427 for IVCNZ motion blur and 0.060-0.233 for PCM defocus. A compact ablation identifies the bounded detail path as the largest local contributor, while degradation conditioning and detector-aware stability terms provide complementary guidance.
132. 【2608.08955】Damage Classification for 3D Point Cloud Data via 3D Data Analysis and Vision Foundation Model-based 2D Projections
链接:https://arxiv.org/abs/2608.08955
作者:Evan Perez,Kalelo Dukuray,Erika Ardiles-Cruz,Jie Wei
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:point cloud data, limited labeled data, high computational demands, point cloud, remains a persistent
备注:
点击查看摘要
Abstract:Fine-grained damage classification of 3D point cloud data (PCD) remains a persistent challenge, constrained by high computational demands and limited labeled data. This study examines two methods: 3D PCD-based damage assessment (3PDA) algorithm and 2D projection damage assessment (2PDA) In our 3PDA analysis algorithm, TDA is used to derive compact representations of 3D PCD segmented by pointNet, which are then integrated with anomaly detection algorithms to quantify structural degradation. We show that TDA effectively compresses geometric structure from VFM-segmented components into discriminative feature vectors and that anomaly detection models can reliably distinguish components with varying damage severity using only 3D PCD inputs. In the 2D projection analysis algorithm, we leverage large VFMs for granular damage detection by projecting 3D PCD into 2D views. These projections allow VFM based models to achieve competitive classification performance while requiring only a fraction of the computational cost associated with full 3D data processing. Our results demonstrate that 2D VFM pipelines in 2PDA can perform strongly on fine-grained damage classification tasks, highlighting their viability as lightweight, resource-efficient alternatives to traditional 3PDA architectures. Comparative evaluation shows that the 3PDA attains higher accuracy but only for a narrow subset of object geometries and at substantially higher computational cost due to its reliance on TDA and the scarcity of high-fidelity 3D datasets. In contrast, the 2PDA algorithm yields slightly lower accuracy but offers an order of magnitude reduction in time complexity and generalizes across a far broader range of object categories.
133. 【2608.08951】opology-Aware Global-Local Mamba Networks for Palm Vein Biometrics
链接:https://arxiv.org/abs/2608.08951
作者:Zhengxi Wu,Felix Marattukalam,Waleed H. Abdulla
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:carry discriminative information, fine-grained biometric task, vessel tree carry, tree carry discriminative, Palm-vein recognition
备注:
点击查看摘要
Abstract:Palm-vein recognition is a fine-grained biometric task in which both local vascular texture and the global layout of the vessel tree carry discriminative information, while public datasets remain limited. We propose a topology-aware global-local backbone that combines multi-scale local features, a structureguided directional stream built on a fixed Sobel-magnitude edge prior, and a four-direction state-space scan global pathway within six Topology-Aware Blocks. A staged gated fusion integrates local, structural, and global representations in that order. On HKPUNIR, our method achieves 99.13% top-1 accuracy and 0.08% EER with 7.2 M parameters; on VERA Palm Vein, it achieves 92.42% accuracy and 0.61% EER. Across both datasets it attains the lowest EER among ResNet50, Vim-S, ViT-S, and GLVM at the smallest parameter count, while GLVM remains the strongest in top-1 accuracy and the cheapest in FLOPs. Code is available upon request.
134. 【2608.08949】EndoMD-SLAM: Endoscopic Gaussian Splatting SLAM under Optical Degradation with Memory and Static-Transient Decomposition
链接:https://arxiv.org/abs/2608.08949
作者:Nuo Chen,Kangqi Ni,Lulin Liu,Joga Ivatury,Ying Ding,Farshid Alambeigi,Tianlong Chen,Zhiwen Fan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:clinical endoscopic navigation, Gaussian Splatting SLAM, reconstruction is critical, navigation and documentation, Splatting SLAM systems
备注: Project page: [this https URL](https://endomd-slam.github.io/)
点击查看摘要
Abstract:Dense 3D reconstruction is critical for clinical endoscopic navigation and documentation. While Gaussian Splatting SLAM systems show promise in this domain, they fundamentally rely on strict multi-view photometric consistency. In routine procedures, this assumption is severely violated by intermittent optical degradations like moving debris and water flushing. Standard systems erroneously fuse these cameraattached artifacts into the persistent 3D geometry, causing severe tracking drift and irreversible map corruption. To address this limitation, we propose EndoMD-SLAM, a framework designed to maintain stability under optical degradation through specialized tracking and mapping mechanisms. On the tracking side, a memory-driven gating mechanism detects unreliable observations to suspend map updates and utilizes historical keyframes for drift-aware relocalization. On the mapping side, a self-supervised static-transient decomposition isolates visual contaminants into a dedicated transient field. This explicit separation prevents artifacts from structurally entangling with the persistent anatomical map. We curate a degradationfocused benchmark from colonoscopy videos to systematically evaluate these failure modes. Extensive experiments show that while standard baselines fail under severe optical degradation, EndoMD-SLAM preserves geometric integrity, reducing absolute trajectory error by 91% and improving rendering fidelity by 9.9 dB PSNR.
135. 【2608.08947】Can Webcam Gaze Constrain Mesa-Objectives in Driving Models? An Instrument Precision Analysis
链接:https://arxiv.org/abs/2608.08947
作者:Lennox Anderson,Ahmed Boutar,Jonah Mulcrone,Tal Erez
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG)
关键词:develop mesa objectives, learned internal goals, Current hazard detection, achieve high training, high training performance
备注: 6 pages, 3 figures, 4 tables
点击查看摘要
Abstract:Current hazard detection systems in autonomous driving may develop mesa objectives, learned internal goals that achieve high training performance through spurious correlations rather than genuine hazard recognition. We investigate whether human gaze patterns, captured via webcam-based eye tracking (this http URL), can serve as privileged information to constrain mesa-objective formation. We collected 137,663 frame-level gaze samples synchronized with hazard annotations across 388 real dashcam clips, then test this hypothesis across two calibration protocols (9-point/45-click and 11-point/440-click), two model architectures (Random Forest and causal Transformer), and five random seeds per experiment with paired t-tests. No experiment yields a statistically significant improvement from gaze (p = 0.919, 0.578, and 0.667 respectively). A geometric analysis reveals the root cause: WebGazer's reported error (~130-257 px depending on configuration) exceeds 93% of detected hazard object sizes (median 36 px), rendering object-level gaze attribution physically impossible at this instrument precision.
136. 【2608.08929】Zero-shot 2D Grounding with Novel Affordance Types
链接:https://arxiv.org/abs/2608.08929
作者:Haomeng Zhang,Raymond A. Yeh
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:aims to locate, human can interact, affordance grounding aims, affordance, grounding aims
备注:
点击查看摘要
Abstract:2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types seen during training and does not study models' ability to generalize to novel affordances, which is crucial for real-world applications. We propose the task of zero-shot 2D grounding with novel affordance types (NAT) and introduce the NAT benchmarks. We then propose AffordAnything, a training-free method that leverages segmentation cues, motivated by the strong correlation between affordance regions and object subparts. To further improve performance, we develop AffordAnything+, a trainable variant that learns to combine these cues. On the proposed AGD20K-NAT benchmark, our best model AffordAnything+ achieves a substantial improvement of 12.3% (absolute) in IoU@0.4 over the SOTA affordance grounding method, OOAL.
137. 【2608.08924】From Noise to Meaning: Meaningful Secret Sharing with Tamper Detection for Facial Recognition
链接:https://arxiv.org/abs/2608.08924
作者:Ajnas Muhammed,Iurii Medvedev,Nuno Gonçalves
类目:Computer Vision and Pattern Recognition (cs.CV); Cryptography and Security (cs.CR)
关键词:recognition system directly, system directly demands, directly demands protection, Popularity of AI-based, AI-based face recognition
备注: Accepted at the IEEE International Joint Conference on Biometrics (IJCB) 2026
点击查看摘要
Abstract:Popularity of AI-based face recognition system directly demands protection of sensitive biometric data used for training. Visual secret sharing is an interesting idea, as it splits facial images into secret shares that look random and spread across many institutions. However, these shares look like noise and can easily spark suspicion and recognized as encrypted content. This makes them open to targeted collection and harvest-now-decrypt-later attacks. Additionally, visual secret sharing does not detect tampering, allowing attackers to modify shares and threaten the integrity of reconstruction. In this paper, we introduce a new method that turns distracting noise-like secret shares into visually appealing cover images with additional cryptographic tamper detection. The proposed technique works with visual secret sharing and introduces cover images to embed the shares using adaptive least significant bit steganography. Here, cover images with perceptual transparency are used to store secret shares while guaranteeing complete privacy. A two layer authentication using strong digital watermarking and cryptographic hashing is used to protect the integrity of shares. The proposed technique shows high resilience in stopping bit-flipping, cropping, and substitution attacks. Extensive experiments on multiple public face datasets show that the technique shows better FR accuracy, while eliminating share conspicuousness and guaranteeing integrity. The proposed framework sets a new standard for protecting facial data in such a way that privacy, security, and integrity are protected.
138. 【2608.08907】oolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
链接:https://arxiv.org/abs/2608.08907
作者:Delin Mao,Chenghao Sun,Jingwei Song,Chishui Chen,Linfeng Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Thinking with images, invoking visual tools, compensate for limited, limited perception, perception by invoking
备注: 18 pages, 13 figures
点击查看摘要
Abstract:Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevailing SFT-then-RL recipe creates a different supervision misalignment at each stage. SFT is expected to teach how to use tools, but trajectories from stronger teachers may succeed through perceptual capabilities that a smaller student cannot reliably reproduce or exploit, causing the student to imitate tool-call patterns without learning how to make them useful. RL is expected to teach when to use tools, but outcome-only rewards make fallible tool execution a liability and suppress tool use, whereas a blanket bonus for every correct tool-using trajectory encourages valid but ineffective operations. To address these two misalignments, we introduce ToolVision. During SFT, a multi-agent pipeline explores candidate trajectories, and a committee including student-scale models scores stepwise evidence gain to rank and prune the search branches. Only successfully executed trajectories with correct final answers are retained for SFT. Before RL, ToolVision compares the learner's performance with and without tools, then rewards successful tool use only on questions where tools provide a clear benefit. Both signals are constructed automatically from public task data without additional human annotations of tool use or necessity. ToolVision-8B improves over its base on all seven main benchmarks, surpasses Thyme-7B, CodeVision-8B, and CodeDance-7B on all three high-resolution benchmarks, and outperforms Qwen3-VL-32B-Thinking on V* and HRBench 8K. We will publicly release the datasets and source code.
139. 【2608.08904】From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
链接:https://arxiv.org/abs/2608.08904
作者:Alexander Hackett,Arnaud Denis-Remillard,Axel Cassou
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:spatial understanding remains, base VLM, vision-language model, base VLM depth, open-source base VLM
备注: Accepted to the archival proceedings track of the Embodied Multimodal Reasoning (EMR) Workshop at ECCV 2026
点击查看摘要
Abstract:How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-language-action model (VLA)? We probe depth perception, a primitive of spatiogeometric understanding, from every decoder layer of a weight-matched open-source base VLM/VLA pair: Molmo2-ER and MolmoAct2-LIBERO. First, the VLA decodes depth worse at every layer, a persistent gap we call the floor. Second, the degradation is not uniform: while the base VLM's depth decodability improves through its final layers, the VLA's collapses, an additional late-layer drop we call the cliff. We causally localize the cliff to late-layer MLP interference: ablating the late-layer MLP writes recovers the majority of the terminal decodability cliff, while matched attention ablations and the same intervention in the weight-matched base VLM produce no comparable recovery. A module-level decomposition explains this dissociation: the base VLM carries depth most accessibly in accumulated MLP writes, whereas action post-training collapses depth decodability in the late accumulated writes.
140. 【2608.08887】City Sentinel: A Unified AI-Based Smart Surveillance Framework for Real-Time Multi-Threat Detection Using Deep Learning
链接:https://arxiv.org/abs/2608.08887
作者:Hanan Syed Shabir,Noor Fatima,Safia Baloch,Masroor Hussain
类目:Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY)
关键词:public safety risks, monitor multiple public, multiple public safety, Rapid urbanization, urbanization has increased
备注: 6 pages, FIT 2026
点击查看摘要
Abstract:Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems often use separate solutions for facial recognition, vehicle identification, fire detection, and behavioral analysis, resulting in fragmented infrastructure and multiple interfaces for operators to manage. This paper presents City Sentinel, a unified AI-based surveillance framework that integrates six detection capabilities into one scalable platform: facial recognition, automatic number plate recognition (ANPR), fire and smoke detection, weapon and knife detection, violence detection, and road accident detection. The system combines a this http URL operator dashboard, FastAPI backend, cloud-based PostgreSQL event storage, InsightFace and YOLOv8 vision models, and EasyOCR for plate recognition. Camera streams are processed through dedicated inference workers using RTSP. On a workstation equipped with an NVIDIA RTX 3060 GPU, the system achieves a median end-to-end latency of 743 ms and supports four concurrent RTSP streams within a two-second latency limit. It achieves a 91.2% face-match rate, 85.7% plate-reading accuracy, and mAP@0.5 scores of 0.846 to 0.889 across the fire, knife, and weapon detection modules. In user-acceptance testing, operators could enroll a new identity in under one minute and identify a flagged person from live footage in an average of 12 seconds. The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.
141. 【2608.08874】AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images
链接:https://arxiv.org/abs/2608.08874
作者:Rui Li,Chenxi Duan,Haoyang Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Spoken language offers, Spoken language, offers a natural, hands-free interface, written expressions
备注:
点击查看摘要
Abstract:Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote-sensing image segmentation benchmarks accept only written expressions. To bridge this gap, we introduce \dataset, a spoken-query benchmark derived from RISBench that adds accent- and voice-diverse speech while preserving the original image, mask, and data splits. Its hard evaluation sets combine rotor, wind, and mixed interference with three signal-to-noise levels. We also propose \model, an efficient bilateral network that combines a boundary-preserving visual path with token-preserving speech encoding, kernel linear cross-modal attention, and a resolution refinement head. The design conditions visual features at two scales without materializing a dense speech--visual affinity matrix, then restores fine boundaries using high-resolution visual features. On the clean test split, \model with Swin-Base achieves 62.09\% mean intersection over union (mIoU) and 68.22\% overall intersection over union (oIoU), outperforming the strongest audio-adapted remote-sensing baseline by 5.38 and 2.08 percentage points, respectively. It retains the best hard-set mIoU at 54.09\%. To the best of our knowledge, this is the first benchmark and model study of full-sentence spoken-query referring segmentation for remote-sensing imagery. The code will be made publicly available.
142. 【2608.08873】Sparse Attention to Emotion: Efficient Facial Emotion Recognition via Token Reduction
链接:https://arxiv.org/abs/2608.08873
作者:Aya Manel Zitouni,Aicha Zenakhri,Karim Haroun,Larbi Boubchir
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Facial Emotion Recognition, Current Vision Transformer-based, Vision Transformer-based approaches, Emotion Recognition, Facial Emotion
备注: 6 pages, 2 figures, published at ICIP 2026
点击查看摘要
Abstract:Facial Emotion Recognition (FER) is an important task that has significant implications across various fields such as biometrics, health, and human-computer interaction. Current Vision Transformer-based approaches display quadratic complexity $\mathcal{O}(N^2)$, with N being the input sequence length, making them cumbersome to deploy at the edge. In this paper, we hypothesize that the FER task does not necessarily require all facial information to correctly interpret emotional states, as specific regions such as the eyes, the mouth, and parts of the cheeks carry discriminative information that can be sufficient to recognize emotions. Based on this, we propose Sparse Attention to Emotion (SAE), a model that discards image tokens that have no added value to the emotional context, while preserving good accuracy and achieving a significant gain in computational cost. Surprisingly, even after suppressing 90\% of the image tokens, our model achieves competitive accuracy to state of the art methods at much lower cost, providing a lightweight Facial Emotion Recognition approach. Experimental results demonstrate that SAE achieves new state of the art results on the RAF-DB dataset while reducing the computational complexity by up to 90\%.
143. 【2608.08867】Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline
链接:https://arxiv.org/abs/2608.08867
作者:Dipit Saha,Shah Mohammad Abdul Mannan,Mohammad Raihan Rashid,Ruwad Naswan,Ahnaf Tahmid
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Traffic surveillance cameras, raw CCTV footage, converting raw CCTV, capture accidents continuously, surveillance cameras capture
备注: Accepted at the AUTOPILOT Workshop, CVPR 2026, Denver, CO
点击查看摘要
Abstract:Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, and what type of collision occurred remains unsolved at scale. The ACCIDENT @ CVPR benchmark evaluates exactly this joint prediction under a strict constraint: no labeled real-world training data is available. We introduce a training-free, two-pass coarse-to-fine pipeline that pairs a frozen Qwen3-VL-32B-Instruct vision-language model with YOLO11x object detection and BoT-SORT tracking. A first pass sparsely samples the full clip to anchor the collision moment in time; a second pass re-examines a tight window around that estimate using frames annotated with stable vehicle identities and normalized bounding-box coordinates, which gives the model both a visual overlay and an explicit numeric description of the same scene. On the official 2,027-clip real-CCTV test set, our system achieves a three-way harmonic mean score of 0.504, surpassing all organizer-published baselines including the best multi-model ensemble (0.412) by a 22% relative margin.
144. 【2608.08844】oward Mask Annotation-Free Surgical Instrument Segmentation from Endoscopic Images Using Text-Prompted Segment Anything Model 3 (SAM3)
链接:https://arxiv.org/abs/2608.08844
作者:Nakul Poudel,Richard Simon,Cristian A. Linte
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:existing methods rely, computer-assisted interventions, scalability and automation, manual spatial prompts, fundamental task
备注: Accepted at the Medical Image Understanding and Analysis (MIUA)
点击查看摘要
Abstract:Surgical instrument segmentation is a fundamental task for computer-assisted interventions, yet most existing methods rely on pixel-level annotations or manual spatial prompts, which limit scalability and automation. The recently introduced Segment Anything Model 3 (SAM3) offers a pathway to annotation-free, automatic segmentation via text-based prompting; however, the instrument name as a text prompt could not be directly used due to a large domain gap. To overcome these limitations, we propose a two-stage framework that achieves instance-level segmentation without requiring ground truth masks or manual interaction. In the first stage, we leverage a natural-language-aligned generic prompt - "tool" - to produce binary masks using SAM3's zero-shot capability. In the second stage, these masks are extended to instance-level by integrating a vision-language model (Qwen) that is fine-tuned on SAM3-generated masked regions for instrument classification. We evaluate our approach on the EndoVis 2017 and 2018 datasets. Results show that, while our two-stage approach does not reach the performance of current fully supervised methods, it significantly outperforms the direct use of SAM3 for instance-level instrument segmentation with text prompts. Overall, our findings highlight both the limitations and potential of SAM3, suggesting a promising direction toward annotation-free surgical instrument segmentation.
145. 【2608.08840】SLAP: Selective Local Vision-Language Alignment for Fish Re-Identification via Partial Optimal Transport
链接:https://arxiv.org/abs/2608.08840
作者:Cigdem Beyan,Tonje Knutsen Sordalen,Kim Tallaksen Halvorsen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Individual fish re-identification, fine-grained recognition problem, Individual fish, specific body regions, Partial Optimal Transport
备注: This is the author version prior to incorporating the camera-ready comments. The final version will be included in the Proceedings of the European Conference on Computer Vision (ECCV) 2026
点击查看摘要
Abstract:Individual fish re-identification (ReID) is a fine-grained recognition problem in which identity-discriminative cues are often localized to specific body regions rather than distributed uniformly across the animal. Nevertheless, recent CLIP-based ReID methods rely predominantly on global image-text alignment, allowing background and weakly discriminative regions to contribute to cross-modal supervision. We propose a selective local vision-language alignment framework that establishes localized correspondences between visual patch embeddings and multiple identity-aware prompt embeddings through Partial Optimal Transport (POT). Rather than enforcing exhaustive correspondence, POT enables selective matching between visual patches and prompt embeddings, allowing the model to emphasize the strongest cross-modal correspondences while avoiding forced alignment of weakly matching regions, thereby yielding more discriminative visual representations for retrieval. The framework is trained end-to-end, while only the adapted visual encoder is retained during inference. Experiments on the longitudinal Symphodus melops dataset demonstrate consistent improvements over recent CLIP-based ReID methods under both closed-set and open-set evaluation protocols. Additional evaluations on other datasets further demonstrate the generalization capability of the proposed method across diverse marine ReID benchmarks.
146. 【2608.08839】SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models
链接:https://arxiv.org/abs/2608.08839
作者:Junjie He,Junfeng Li,Zhide Zhong,Haodong Yan,Ruixin Li,Yangyang Zheng,Jiaguan Zhu,Tianran Zhang,Yuqiao Du,Wen Chen,Shunbo Zhou,Haoang Li
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:promising paradigm, paradigm for robotic, World-Action Models, semantic, semantic guidance
备注:
点击查看摘要
Abstract:World-Action Models (WAMs) have emerged as a promising paradigm for robotic manipulation. However, most existing WAMs generate future videos and actions by relying mainly on visual cues rather than language instructions, since off-the-shelf text encoders embed instructions independently of visual observations. As a result, the videos predicted by these WAMs are often semantically misaligned with their corresponding language instructions, which degrades the accuracy of the predicted actions. To overcome this limitation, we propose SG-WAM, a semantic guidance method for world-action models that leverages a vision-language model (VLM) as a semantic planner to enhance the instruction-grounding capacity of world-action models. Specifically, we train a VLM-based planner to predict text-grounded and spatial-aware semantic foresight. The text-grounded semantic foresight grounds the instruction by identifying the correct target objects, and the spatial-aware semantic foresight provides the scene geometry for precise manipulation. We then inject this foresight into the world-action model as high-level semantic guidance, ensuring that both future-video generation and action prediction faithfully follow the language instruction. Extensive experiments in simulation and the real world demonstrate the superiority of our semantic guidance method, showcasing precise manipulation and strong instruction-following capabilities.
147. 【2608.08832】Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
链接:https://arxiv.org/abs/2608.08832
作者:Donghui Feng,Fengxi Zhang,Changsheng Gao,Wenhan Yang,Qi Wang,Qunshan Gu,Hongwei Hu,Zhengxue Cheng,Li Song
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:large vision foundation, making efficient feature, exchanges intermediate token, vision foundation models, intermediate token features
备注: 13 pages, 9 figures
点击查看摘要
Abstract:Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and patch tokens into an L x C pseudo image, causing entropy models to mainly capture sequence-axis dependencies while overlooking the native two-dimensional patch-grid structure. In this paper, we show that ViT patch tokens retain strong local spatial correlations on the original grid. To exploit this structural prior, we propose the Visual Token Codec (VTC), a dual-path learned codec that separates global and patch tokens into dedicated coding paths. Global tokens are compressed with a lightweight factorized prior, whereas patch tokens are encoded on the patch-token grid using a spatial-channel context entropy model. To support intermediate-layer compression and practical rate adaptation, VTC further incorporates feature-matching supervision after subsequent ViT blocks and variable-rate modules within a single codec. Experiments on DINOv2 and SAM3 show that VTC consistently outperforms representative ViT feature coding baselines on classification, segmentation, and detection tasks. At 90% of uncompressed-feature performance, VTC reduces bitrate by 15.7x-37.4x across these tasks. We further provide intermediate-layer rate-utility analyses for practical transmission- and storage-oriented deployment scenarios.
148. 【2608.08820】LogiShot: Logically Coherent Cross-Shot Video Generation
链接:https://arxiv.org/abs/2608.08820
作者:Shuai Guo,Yuhang Yang,Zeyu Zhang,Pengfei Yu,Wei Zhai,Yang Cao,Zheng-Jun Zha
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generating cross-shot videos, Generating cross-shot, logically connected, connected is essential, Generating
备注:
点击查看摘要
Abstract:Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production, still rely on isolated textual scripts or explicit reference images to specify the generated content. Consequently, when user instructions are underspecified or ambiguous, a generated clip may appear visually plausible on its own but fail to align with the overall narrative, leading to disjointed content. We argue that achieving cross-shot logical coherence in video generation requires establishing logical connections across shots and maintaining visual consistency. To this end, we propose LogiShot, which incorporates information through two complementary paths: 1) LogiShot jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation; 2) the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots. Additionally, we construct a dataset with 110K samples and a dedicated benchmark for evaluating cross-shot logical coherence. Experiments demonstrate that LogiShot consistently outperforms existing baselines in terms of logical coherence across multiple shots. Model and data will be made publicly available.
149. 【2608.08819】MRI super-resolution in ten sampling steps using a diffusion bridge model
链接:https://arxiv.org/abs/2608.08819
作者:Mojtaba Safari,Hang Yu,Zach Eidex,Mingzhe Hu,Ryan J. Sanford,Alexandru Florea,Shansong Wang,Chih-Wei Chang,Erik H Middlebrooks,Aditya Juloori,Stanley L. Liauw,Ralph Weichselbaum,Xiaofeng Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Medical Physics (physics.med-ph)
关键词:Objective, SR-DBM, super-resolution diffusion bridge, diffusion bridge model, long acquisition times
备注:
点击查看摘要
Abstract:Objective. MRI provides excellent soft-tissue contrast, but long acquisition times can cause patient discomfort and lead to motion artifacts, forcing a trade-off between spatial resolution and scan time. Diffusion-based super-resolution (SR) reconstructs high-resolution (HR) images from low-resolution (LR) inputs, but typically needs many sampling steps and initializes from a Gaussian prior ill-suited to image restoration. We developed an efficient diffusion framework that reconstructs HR MRI directly from LR data. Approach. We propose super-resolution diffusion bridge model (SR-DBM), a super-resolution diffusion bridge model that casts SR as a stochastic transport between the LR and HR image distributions. Through a Doob's h-transform of a mean-reverting stochastic differential equation, SR-DBM pins the process to the paired HR and LR images at its endpoints, initializing reconstruction from the measured anatomy rather than from Gaussian noise. The HR image is recovered by a deterministic reverse trajectory in which a network predicts the clean image at each of only ten sampling steps. We evaluated SR-DBM on ultra-high-field 7T brain T1 MP2RAGE maps and pelvic T2-weighted prostate images against nine comparison methods using PSNR, SSIM, GMSD, and LPIPS. Main results. SR-DBM attained the highest PSNR and SSIM and the lowest GMSD on both datasets (brain: 27.66+-1.52 dB, 0.96+-0.02, 7.96+-1.86$; prostate: 27.87+-2.29 dB, 0.80+-0.05, 8.38+- 1.44), with statistically significant gains over every comparison method (two-sided Wilcoxon signed-rank test with Holm correction, p0.05). The strongest baseline, SR-EMamba, ranked second. Qualitatively, SR-DBM produced the smallest residual errors and best preserved fine structures and lesions.
150. 【2608.08814】360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
链接:https://arxiv.org/abs/2608.08814
作者:Kenta Watanabe,Atsuyuki Miyai,Mizuki Takenawa,Kiyoharu Aizawa,Toshihiko Yamasaki
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Spatial Reasoning, urban exploration capabilities, photorealistic environment constructed, Reasoning, environment constructed
备注: ECCV2026. Project Page: [this https URL](https://360mm-team.github.io/360CityArena/)
点击查看摘要
Abstract:We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.
151. 【2608.08808】AdapterMoE: A Two-Stage Hard-Routing Mixture-of-Experts Architecture for Multi-Crop Disease Recognition with Calibrated Rejection and Incremental Learning
链接:https://arxiv.org/abs/2608.08808
作者:Pin-Hsun Huang,Shaou-Gang Miaou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Timely crop-disease identification, Timely crop-disease, food security, crop-disease identification, identification is critical
备注: 17 pages, 7 figures, 14 tables
点击查看摘要
Abstract:Timely crop-disease identification is critical to food security. Multi-crop recognition suits Mixture-of-Experts (MoE), but conventional soft-routing MoE learns crop assignment freely end-to-end, letting a few experts dominate (expert collapse) with no semantic correspondence to crops, and facing high retraining costs, unstable rejection of non-target inputs, and a saturated accuracy ceiling. We shift the objective from accuracy toward a trade-off among deployment cost, scaling flexibility, and rejection stability, using deterministic hard routing. We propose AdapterMoE: a RouterHead classifies the crop and rejects non-target crops via a Maximum Softmax Probability threshold, with a dual-gate Energy+KNN out-of-distribution module catching distribution-shifted inputs; five per-crop Adapters atop a frozen EfficientNet-B0 backbone discriminate diseases, each calibrated via Temperature Scaling. Because experts are hard-isolated at the data level, the design avoids expert collapse and exposes an add_crop interface for local, per-crop updates instead of full retraining. On PlantVillage (5 crops, 26 classes), across a fair five-system comparison, AdapterMoE attains accuracy statistically indistinguishable from the best baselines (Macro-F1 within a 0.24-point band) while cutting training cost to about 9% of full-network baselines, expanding to a new crop in
152. 【2608.08805】LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation
链接:https://arxiv.org/abs/2608.08805
作者:Jinhong Zhu,Weiqi Yan,Shengchuan Zhang,Liujuan Cao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generalization Semantic Segmentation, Domain Generalization Semantic, unseen target domains, Semantic Segmentation, Generalization Semantic
备注: 10 pages, 4 figures
点击查看摘要
Abstract:Domain Generalization Semantic Segmentation (DGSS) focuses on generalizing knowledge from labeled source domains to unseen target domains where data is unavailable during the training phase. While conventional methods utilize style randomization or feature normalization to mitigate domain shifts, they often impair feature integrity. Specifically, style randomization distorts the underlying feature manifold due to its coarse-grained nature, while feature normalization suppresses discriminative, domain-sensitive semantic details owing to its rigid design. To address these limitations, we propose the Language-and-Source-Anchored Alignment (LASA) framework, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO). Concretely, the TSGST module addresses manifold distortion by utilizing source features as structural anchors and vision-language model (VLM) priors as fine-grained guidance. To restore suppressed discriminative and domain-sensitive details, the DAQA module recalibrates object queries via categorical guidance and domain-aware signatures, while the DADO module aligns the resulting query distributions with a shared classifier to ensure consistent categorical responses across domains. Extensive experiments on challenging benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches.
153. 【2608.08771】UPolarSQ: Polar Representation Learning for Optic Disc and Peripapillary Atrophy Segmentation and Quantification in Fundus Photographs
链接:https://arxiv.org/abs/2608.08771
作者:Mengxian He,Yunyun sun,Ziyue Gao,Wengkei Lam,Shunyi Zhang,Wu Yuan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Myopia-induced posterior-pole remodeling, Peripapillary Atrophy, provide clinically relevant, clinically relevant structural, deformation and Peripapillary
备注:
点击查看摘要
Abstract:Myopia-induced posterior-pole remodeling is frequently accompanied by Optic Disc (OD) deformation and Peripapillary Atrophy (PPA), both of which provide clinically relevant structural biomarkers. In Cartesian fundus images, however, PPA often appears as an irregular and partially visible crescent adjacent to the OD, leading to fragmented segmentation and post-processing-dependent quantification. We propose UPolarSQ, a unified polar-domain framework for OD/PPA segmentation and biomarker quantification in myopic fundus images. UPolarSQ first maps an OD-centered region of interest into polar coordinates, where OD and PPA boundaries can be represented as radial profiles. It then employs UPolarSeg, a U-Net-based segmentation network enhanced with a Radial-Angular-Decoupled Module and boundary-aware auxiliary supervision to model anisotropic polar features and radial boundary transitions. Clinical biomarkers, including disc shape and PPA-width-related measurements, are deterministically extracted from the predicted polar masks, aligning segmentation and quantification within a shared geometric representation. Experiments on internal and external cohorts demonstrate that UPolarSQ improves OD/PPA segmentation and supports reliable polar-native biomarker estimation for myopic analysis.
154. 【2608.08753】Parcel2Progression: An Anatomy-aware Longitudinal Framework for Alzheimer's Disease Diagnosis
链接:https://arxiv.org/abs/2608.08753
作者:Madhumitha Venkatesh,Shanawaj S Madarkar,Konda Reddy Mopuri
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:subtle pathological cues, early stages, Longitudinal Transformer Framework, process with subtle, subtle pathological
备注:
点击查看摘要
Abstract:Alzheimer's disease (AD) progression is a longitudinal process with subtle pathological cues in the early stages. Yet, computational constraints have limited most neuroimaging models to either compromise spatial information or limit the number of longitudinal scans. We aim to overcome this bottleneck and fully leverage high-resolution, variable-length T1w structural MRI (4D sMRI) scan sequences. We introduce Parcel2Progression (P2P), a Longitudinal Transformer Framework which tackles this challenge using an Atlas-guided Parcel Encoder that tokenizes 3D scans into a set of richer anatomically grounded representations. A Longitudinal Transformer then integrates irregular, arbitrary-length longitudinal visits with patient age. This synergy delivers two key advantages: (1) parcel-specific interpretability, and (2) computational tractability for long-term analysis, which scales linearly with the number of scans compared to a naive quadratic 4D ViT cost. P2P outperforms prior works and baselines in both MCI (Mild Cognitive Impairment) to AD conversion prediction and AD vs. CN (Cognitively Normal) classification tasks across ADNI, AIBL, and MIRIAD datasets. Leveraging longitudinal scans boosts performance over single-scan baselines by up to 5% and 7% in balanced accuracy for AD classification and MCI conversion prediction tasks, respectively. Interpretability analysis using parcel saliencies and attention rollouts reveals clinically consistent atrophy patterns in AD and MCI subjects. We also demonstrate the frameworks' reliability in anomaly detection using a synthetic dataset, and test the model's generalizability for other neurodegenerative diseases like Frontotemporal Dementia.
155. 【2608.08734】IDATA: Scalable Invertible Diffusion for Unrestricted Adversarial Transfer Attack
链接:https://arxiv.org/abs/2608.08734
作者:Yi Pan,Jun-Jie Huang,Tianrui Liu,Zihan Chen,Lin Liu,Zhao Wentao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Unrestricted adversarial transfer, important for evaluating, Unrestricted adversarial, IDATA, visual
备注: Adversarial attack,Invertible diffusion model,Memory-efficient,Low-frequency
点击查看摘要
Abstract:Unrestricted adversarial transfer attacks are important for evaluating the black-box robustness of deep visual models. Diffusion-based attacks have shown promising transferability and visual imperceptibility by optimizing adversarial perturbations along denoising trajectories in latent space. However, existing methods are limited by two challenges: memory-intensive multistep backpropagation and frequency-agnostic perturbation over intermediate latents. To address these issues, we propose IDATA, a memory-efficient diffusion framework for unrestricted adversarial transfer attack. IDATA consists of two key components: an Invertible Diffusion Module (IDM) and a Low-Frequency Constraint Module (LFCM). Specifically, IDM reformulates adversarial optimization over diffusion trajectories as an invertible process, enabling constant-memory backpropagation through on-demand reconstruction of intermediate states instead of storing the full denoising chain. Moreover, LFCM leverages Discrete Wavelet Transform (DWT) to decompose latent variables into low- and high-frequency components, restricting perturbations to semantically stable low-frequency subspaces, thereby improving transferability while preserving visual imperceptibility. Extensive experiments on multiple benchmarks and diverse model architectures demonstrate that IDATA consistently outperforms state-of-the-art baselines in attack success rate, memory efficiency, and visual imperceptibility. These results suggest that IDATA is a promising tool for black-box robustness evaluation of deep visual models. Code is available at this https URL.
156. 【2608.08732】AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval
链接:https://arxiv.org/abs/2608.08732
作者:Haoyu Zuo,Yibo Yan,Xin Zou,Shuliang Liu,Yi Cao,Mingdong Ou,Xuming Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Multi-vector vision-language retrievers, fine-grained Visual Document, incurs substantial overhead, vision-language retrievers enable, retrievers enable fine-grained
备注: 24 pages, 7 figures
点击查看摘要
Abstract:Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive compression, whereas merging does not explicitly prioritize important regions when forming representatives. We introduce AnchorFold, a training-free focus-then-fold framework for document-side index compression. AnchorFold applies Recursive Attention Propagation over visual self-attention graphs, performing multi-step propagation within each attention head and integrating scores across heads and layers. The focus stage selects the highest-centrality tokens as anchors. The fold stage assigns remaining tokens to their most similar anchors in the normalized retrieval space and summarizes each anchor-centered group through centrality-weighted aggregation. This preserves non-anchor contributions while concentrating capacity on structurally important tokens. Across ViDoRe v1/v2 and REAL-MM-RAG with three diverse retrieval backbones, AnchorFold consistently outperforms all evaluated training-free baselines at $\gamma \leq 0.20$. On ViDoRe v1/v2, it retains 98.3% of full-index NDCG@5 on average at $5\times$ compression, achieving near-lossless compression, and 92.4% at $20\times$ compression.
157. 【2608.08727】omaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases
链接:https://arxiv.org/abs/2608.08727
作者:Gia-Han Truong,Khang Nguyen Quoc,Luyl-Da Quach
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Tomato leaf disease, tomato disease understanding, large-scale Tomato leaf, leaf disease MultiModal, disease MultiModal Understanding
备注: Accepted at ECCV CVPPA Workshop
点击查看摘要
Abstract:To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a benchmark for evaluating VLMs on tomato disease understanding. TomaMMU comprises 28,808 high-quality images spanning 15 categories and 213,119 human-annotated visual question-answer pairs, generated through a three-stage pipeline comprising Data Collection, Human Annotation, and Question-Answer Generation. Building on this foundation, TomaBench organizes seven agricultural tasks into a hierarchical three-level taxonomy spanning Basic Perception, Pathology Understanding, and Expert Diagnosis, which together enable systematic evaluation from low-level visual recognition to high-level diagnostic reasoning. The tasks assess visual symptom recognition, taxonomic relationships, and diagnostic reasoning, offering a comprehensive view of how well models grasp plant pathology. Our results pronounced gaps in fine-grained recognition and factually grounded reasoning with 14 state-of-the-art VLMs, consistently underperforming on both challenging MCQs and open-ended questions. These results suggest that current VLMs struggle to translate visual perception into reliable diagnostic knowledge, motivating the need for targeted domain adaptation. Simple fine-tuning on TomaMMU substantially narrows this gap, boosting accuracy on challenging MCQs to 96.09%, outperforming recent VLMs, and pointing toward promising directions for future work. All data and code is available in this https URL.
158. 【2608.08720】High-Quality Exposure Correction with Diffusion-Based Image Generation Priors
链接:https://arxiv.org/abs/2608.08720
作者:Ziwen Li,Meng Cao,Jinpu Zhang,Chunyang Li,Long Bao,Heng Sun,Yuehuan Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:place excessive focus, extreme exposure regions, exposure correction, pixel-wise accuracy, place excessive
备注: Accepted by IEEE Transactions on Multimedia (TMM)
点击查看摘要
Abstract:Although most existing exposure correction methods achieve high fidelity, they often place excessive focus on overall pixel-wise accuracy, making it challenging to effectively model extreme exposure regions, which results in suboptimal perceptual quality. Recently, diffusion models have received significant attention due to their remarkable performance in the realm of image generation. However, their successful application to exposure correction remains a challenging and open question. The key challenge lies in generating accurate image structures and maintaining high image fidelity during stochastic diffusion processes. In this paper, we propose DPEC (Diffusion Prior-based Exposure Correction), a novel framework for image exposure correction that utilizes diffusion-based image generation priors encapsulated in pre-trained large-scale diffusion models. Specifically, we first propose an efficient fine-tuning strategy to derive an exposure corrector from pre-trained models, enabling the generation of enhanced images in a single-step denoising process. Moreover, we seamlessly combine the strengths of diffusion models and regression models, and design a joint cross-attention module to integrate multi-scale diffusion prior features, thereby effectively preserving high-frequency details and minimizing random artifacts. The diffusion model focuses on dealing with low-frequency content rather than all the intricate texture details. The experimental results demonstrate that the proposed DPEC method consistently outperforms existing state-of-the-art methods on multiple exposure correction datasets, whether in terms of fidelity, perceptual quality, or visual effects.
159. 【2608.08713】Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation
链接:https://arxiv.org/abs/2608.08713
作者:Jonathan Suprijadi,Raphael Stock,Moritz Langenberg,David Zimmerer,Kim-Celine Kahl,Stefan Denner,Yannick Kirchhoff,Karol Gotkowski,Maximilian Rokuss,Jeremias Traub,Tassilo Wald,Constantin Ulrich,Klaus Maier-Hein
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Vision-language models offer, volumes poses substantial, substantial computational challenges, poses substantial computational, radiology report generation
备注:
点击查看摘要
Abstract:Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses substantial computational challenges. Modern foundation vision encoders (VEs) can produce tens of thousands of vision tokens per scan, making the visual sequence passed to the large language model (LLM) a primary computational bottleneck. Vision-to-language projectors can compress this sequence to reduce computation, but may discard clinically relevant detail; conversely, effective compression can accommodate higher-resolution inputs while keeping the downstream token count fixed. How this vision-token budget should be allocated across input field of view, spatial resolution, and vision-to-language projection therefore remains an open design question. We systematically evaluate four heterogeneous VEs (CNN- and ViT-based), five token-reducing projectors at up to 64x compression alongside a non-reducing MLP projector baseline, and five instruction-tuned LLMs (1.7B--4B) on two large-scale CT report datasets (CT-RATE and Merlin). At matched LLM token budgets, anatomy-guided region of interest cropping is the most consistent strategy, improving clinical macro F1 in 19 of 20 settings by +3.7 points on average for the 3D ViT Primus encoder and +1.1 for the slice-based 2D ViT Curia encoder. Increasing input resolution further is strongly projector-dependent: the PerceiverResampler, paired with higher-resolution Curia features, yields the strongest configuration in the resolution study on both datasets. Our best configurations achieve state-of-the-art clinical macro F1 on the test sets, reaching 49.5 on CT-RATE and 49.0 on Merlin. Code and models will be published upon publication.
160. 【2608.08702】SRE-FER: Regional residual evidence learning for mitigating local evidence dilution in fine-grained facial expression recognition
链接:https://arxiv.org/abs/2608.08702
作者:Jiaye Song,Ruochen Zhang,Yuliang Wang,Jiaqi Wu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Fine-grained facial expression, distinguish adjacent emotions, facial expression recognition, capturing subtle muscular, subtle muscular cues
备注: 10 pages, 5 [this http URL](http://figures.Accepted) at the 9th International Conference on Artificial Intelligence and Pattern Recognition (AIPR 2026)
点击查看摘要
Abstract:Fine-grained facial expression recognition (FER) hinges on capturing subtle muscular cues that distinguish adjacent emotions. Yet capturing these cues presents a dilemma. Detector-based methods depend on fragile landmark pipelines, whereas we find that directly transferring foundation models such as DINOv3 under conventional global readouts can cause local evidence dilution: early global aggregation washes out sparse muscular signals and leaves persistent confusion between categories such as fear/surprise and sad/neutral. To recover this evidence, we propose SRE-FER, a readout-level regional residual evidence learning framework. Its core module, RERA, adds zero-initialized residual logits that refine class boundaries while preserving the backbone's global prediction. Training-time action unit (AU) guidance steers regional features toward expression-relevant areas using Facial Action Coding System (FACS)-based anatomical priors, without requiring an external facial pipeline at inference. An optional Full setting further routes sample-specific non-redundant tokens. On three benchmarks, SRE-FER attains 92.76% on RAF-DB, 91.32% on FERPlus, and 67.78% on AffectNet-7, demonstrating highly competitive performance compared to existing FER methods.
161. 【2608.08696】OccAnyScene: Towards Unified Indoor-Outdoor 3D Occupancy Predictio
链接:https://arxiv.org/abs/2608.08696
作者:Junjie Liu,Wanshui Gan,Zitong Dai,Guiping Cao,Yan Li,Ke Chen,Dongmei Jiang,Xiangyuan Lan,Jianguo Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Semantic Occupancy Prediction, semantic occupancy methods, fixed scene types, occupancy prediction, semantic occupancy
备注:
点击查看摘要
Abstract:3D occupancy prediction is fundamental to scene understanding, yet existing 3D semantic occupancy methods are typically specialized to fixed scene types and occupancy protocols. We introduce Cross-Scene 3D Semantic Occupancy Prediction, a new task setting which requires a single model to handle heterogeneous indoor and outdoor scenes with varying cameras, spatial ranges, voxel specifications, and semantic taxonomies. This setting poses a fundamental challenge: achieving metric-consistent yet scene-adaptive image-to-3D lifting across varying camera configurations and scene scales. To address this challenge, we propose OccAnyScene, a pixel-frustum-centered Gaussian framework built upon a pretrained depth foundation model. Specifically, the framework employs Pixel-Aligned Frustum Feature Aggregation to construct a camera-aware frustum query for each feature pixel, and Frustum-Parameterized Gaussian Construction to decode each query into multiple Gaussians whose positions and sizes are constrained by the predicted pixel depth and corresponding frustum geometry. OccAnyScene sets new state-of-the-art results, achieving 59.92% mIoU on the indoor Occ-ScanNet and 23.06% mIoU on the outdoor SurroundOcc-nuScenes.
162. 【2608.08693】CUPA-T2*: Covariance-Aware Uncertainty Propagation and Alignment for T2* Mapping in Accelerated MRI
链接:https://arxiv.org/abs/2608.08693
作者:Gideon N. L. Rouwendaal,Natascha Niessen,Hannah Eichhorn,Dirk H. J. Poot,Christine Preibisch,Julia A. Schnabel
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:long scan times, scan times, rendering them impractical, clinical settings, strong potential
备注: Accepted at the 2nd workshop on: Reconstruction and Imaging Motion Estimation (RIME) MICCAI. This is the submitted manuscript with added link to GitHub repo, funding acknowledgements, disclosure of interests, and authors' names and affiliations. No further post submission improvements or corrections were integrated
点击查看摘要
Abstract:Quantitative T2* maps have strong potential for biomarker discovery but are limited by long scan times, rendering them impractical in clinical settings. Significant acceleration can be achieved through undersampling in k-space combined with learning-based reconstruction. However, reconstruction artifacts and noise can propagate into downstream T2* fitting, degrading its accuracy. We introduce CUPA-T2*, a framework that explicitly propagates voxel-wise inter-echo uncertainty from stochastic Monte Carlo dropout reconstructions to downstream T2* fitting via covariance-aware sampling. T2* fitting is performed with a heteroscedastic MLP and a correlation-based regularizer that encourages alignment between predicted variance and reconstruction uncertainty. Experiments on accelerated brain MRI data show tissue-dependent behavior: CUPA-T2* achieves competitive overall T2* fitting performance and improves white-matter performance at higher accelerations. Compared with a heteroscedastic baseline, the proposed framework substantially increases alignment between reconstruction uncertainty and predicted T2* variance, while also revealing a trade-off with calibration (ECE) and selective prediction performance (AURC). CUPA-T2* enables reconstruction uncertainty-aware T2* fitting and delivers voxel-wise uncertainty maps to support the interpretation of quantitative T2* estimates.
163. 【2608.08685】Semi-Dense Matching Uncertainty Is Not Just Local Confidence
链接:https://arxiv.org/abs/2608.08685
作者:Khoa Hoang,Hoang-Tuan Nguyen,Huong Ninh,Hai Tran,Long Q. Tran
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Reliable semi-dense matching, Reliable semi-dense, geometric vision systems, modern geometric vision, vision systems
备注: 14 pages, 12 figures, including supplementary material
点击查看摘要
Abstract:Reliable semi-dense matching is essential for modern geometric vision systems. Designed under a coarse-to-fine paradigm, it achieves an optimal balance between performance and computational cost. However, existing methods often struggle to provide well-quantified uncertainties, where catastrophic coarse-assignment failures are ignored, leading to truncated error distributions and severely misjudged geometric estimations. In this paper, we propose a lightweight, post-hoc overall uncertainty estimation framework that introduces a two-component calibrated Laplace mixture model with only 9 learnable parameters. The objective is to explicitly capture both the sharp local refinement noise and the broader tail of coarse-assignment failures. We introduce the Coarse-success posterior Refit (CoRe) method, a geometric refitting module that utilizes the posterior probability of coarse-assignment success as soft correspondence weights. Extensive experiments show that our method consistently improves downstream geometric accuracy across various pretrained-only matchers and robust estimators with minimal computational overhead. Our code is available at this https URL.
164. 【2608.08676】UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
链接:https://arxiv.org/abs/2608.08676
作者:Jinbo Yan,Limeng Qiao,Jie Qin,Junyan He,Feize Wu,Guanglu Wan
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Semantic vision encoders, vision encoders, visual, generation, Semantic
备注:
点击查看摘要
Abstract:Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine-grained visual details, leading to poor pixel reconstruction and limiting their use in reconstruction-sensitive tasks such as image generation and editing. In this work, we ask whether understanding, generation, and editing can be modeled in a single visual representation space built from a pretrained semantic ViT. We show that the frozen Transformer blocks of a semantic ViT are not intrinsically unable to preserve visual details. Instead, the original patch parameterization drives the representation toward semantic abstraction, making fine-grained information difficult to recover from the final tokens. Based on this observation, we introduce \emph{Patch Reparameterization}, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks. The resulting unified representation preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off. We further scale this representation into \emph{UniSpace}, an 8B Mixture-of-Transformer-Experts model that performs understanding, generation, and editing in the same visual space without a separate VAE pathway. System-level evaluations demonstrate practical text-to-image generation and instruction-based image editing, showing that a reparameterized pretrained ViT can serve as a unified visual interface for scalable multimodal modeling.
165. 【2608.08664】FiRe: Fixed-Noise Refinement for Visual Counterfactual Explanations
链接:https://arxiv.org/abs/2608.08664
作者:Yan Zeng,Changlu Guo,Oskar Kristoffersen,Anders Nymark Christensen,Morten Rieger Hannemose,Anders Bjorholm Dahl
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:preserving decision-irrelevant content, decision-irrelevant content, decisions through realistic, preserving decision-irrelevant, counterfactual explanations aim
备注: Accepted at the British Machine Vision Conference (BMVC) 2026
点击查看摘要
Abstract:Visual counterfactual explanations aim to change classifier decisions through realistic and localized edits while preserving decision-irrelevant content. Existing DDPM-based methods typically perform classifier-guided editing along a long reverse denoising trajectory. The changing noise levels make semantic editability and spatial control difficult to balance, and the editable state is noisy, whereas the target classifier is trained on clean images. As a result, these methods require either costly recursive denoising or low-quality one-step estimates to obtain classifier-facing clean images. We propose FiRe, a Fixed-noise Refinement framework for visual counterfactual explanations. Rather than following a reverse denoising trajectory, FiRe maps the input to a fixed noise level and iteratively refines the noisy state at that level. To provide clean images for classifier guidance, FiRe first adapts Pixel Mean Flow to visual counterfactual explanation, enabling direct clean-image prediction from noisy states. To make fixed-noise refinement produce minimal and localized counterfactual edits, FiRe introduces three FiRe-specific controls: a dynamic dual-mask strategy, adaptive guidance, and early stopping, which determine where edits accumulate, which changes become visible, and when refinement stops. Experiments on five tasks across three datasets show that, compared with the strongest recent baseline, FiRe achieves about 3$\times$ faster online inference and 8$\times$ fewer FLOPs while obtaining comparable or state-of-the-art counterfactual quality.
166. 【2608.08663】A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual Data
链接:https://arxiv.org/abs/2608.08663
作者:Joseph Bingham
类目:Human-Computer Interaction (cs.HC); Computer Vision and Pattern Recognition (cs.CV)
关键词:process psycholinguists call, psycholinguists call lexical, converge on shared, process psycholinguists, psycholinguists call
备注: 23 pages, 5 figures
点击查看摘要
Abstract:Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. Leading vision-language models fail at this: recent empirical work documents that they do not shorten references, reuse successful expressions, or maintain stable pact state across turns. We present a framework that addresses the gap by externalizing pact state into three explicit, inspectable sets of referent-object bindings ($\Gamma, \Xi, \Omega$), updated by a dynamic-semantics context-change rule. The symbolic layer sits on top of a lightweight perceptual-alignment pipeline that grounds noisy human referring expressions in crowd-sourced imagery via SIFT homographies and the Universal Quality Index. Evaluated on the Stanford Repeated Reference Game corpus (over 15{,}000 director-matcher utterances on abstract tangram stimuli), the framework places the correct target in its top-5 hypothesis set 83.56% of the time from a single director utterance. Human matcher top-1 accuracy on the same corpus is approximately 77-80%. We also report results on a held-out condition in which obvious tangram-adjacent images are excluded from the retrieved set, which provides a more conservative measurement of the grounding signal. Ablations isolate the contribution of each component: SIFT alignment, UQI, query preprocessing, and image augmentation. The central contribution is the combination: a transparent, auditable symbolic layer that recovers the structure of lexical entrainment turn by turn, paired with a perceptual channel whose behavior can be examined ablation by ablation. We also discuss in detail what the framework does not do. It is not interactive, it does not close the loop with the director, and its retrieval-driven perceptual channel is vulnerable to a class of leakage effects that we quantify and bound rather than wave away.
167. 【2608.08661】Degradation-Guided Underwater Image Restoration with Task-Oriented Latent Control
链接:https://arxiv.org/abs/2608.08661
作者:Xu Zhang,Xuhui Cao,Kangzhe Yuan,Laibin Chang,Yichu Xu,Shi Chen,Huan Zhang,Yong Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:underwater images plays, guide adaptive restoration, dual role, regulation during decoding, images plays
备注:
点击查看摘要
Abstract:Degradation information in underwater images plays a dual role: its spatial and spectral cues can guide adaptive restoration, while degradation-entangled features may be propagated without explicit regulation during decoding. Existing methods largely overlook this dual role, either underexploiting degradation cues or directly forwarding encoder features through skip connections. To address this issue, we propose PROTEUS, which couples degradation-guided feature adaptation with task?oriented latent control. PROTEUS tackles this problem from two complementary perspectives. At the feature level, the Guided Dynamic Feature Modulation Block exploits spatially varying degradation cues to adapt feature processing across network stages. At the representation level, the task-oriented latent controller learns a structured control code under discriminative regularisation and uses it for channel-wise modulation of skip features, without requiring the code to form a metrically cleaner embedding. Extensive experiments on five paired and four non-reference underwater benchmarks demonstrate that PROTEUS achieves highly competitive restoration performance, with a favourable balance between restoration quality and computational cost.
168. 【2608.08659】JSGS: JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views
链接:https://arxiv.org/abs/2608.08659
作者:Jinhua Cui,Anhong Wang,Kai Hu,Donghan Bu,Peihao Li,Tammam Tillo,Hao Jing,Shiao Xu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:image faithfully samples, Gaussian Splatting, samples scene radiance, faithfully samples scene, input image faithfully
备注:
点击查看摘要
Abstract:Standard 3D Gaussian Splatting (3DGS) assumes that every input image faithfully samples scene radiance. However, mixed-quality JPEG images violate this assumption because compression-induced blocking and ringing artifacts can corrupt updates to Gaussians shared across views. To address this problem, we propose JPEG State-Guided Supervision for 3D Gaussian Splatting from Mixed-Quality Views (JSGS). JSGS uses luminance and chrominance quantization tables stored in each JPEG file to construct a view-specific JPEG observation operator. This operator encodes and decodes each rendered view for domain-matched comparison with the corresponding decoded input image. The luminance quantization table supplies continuous weights within a fixed middle frequency band. A loss in the low frequency band anchors coarse structure, while the weighted middle frequency loss redistributes supervision among the selected DCT coordinates. The resulting block disagreement also guides the Gaussian Controller to regularize small primitives with high opacity in disagreement regions. Across seven scenes and three mixed-quality schedules, JSGS achieves the lowest mean LPIPS and the highest mean SSIM under every schedule while rendering at approximately 150 FPS. Code: this https URL.
169. 【2608.08648】Agentic Visual Reasoning in Whole-Slide Pathology Images via Active Perception
链接:https://arxiv.org/abs/2608.08648
作者:Jingyun Chen,Fengchun Liu,Linghan Cai,Songhan Jiang,Shenjin Huang,Hongpeng Wang,Lequan Yu,Yongbing Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:requires identifying sparse, identifying sparse diagnostic, reasoning requires identifying, Whole-slide visual reasoning, Whole-slide visual
备注: 14 pages, 5 figures
点击查看摘要
Abstract:Whole-slide visual reasoning requires identifying sparse diagnostic evidence in gigapixel pathology slides and integrating observations across spatial scales. Existing WSI methods either compress densely sampled patches into global representations or use pretrained vision-language models with heuristic region selection, weakening links between predictions and morphology or lacking pathology-trained observation policies. We present AdaptivePath, an active-perception framework that formulates WSI evidence acquisition as sequential decision making. The Navigator learns question-agnostic abnormality-driven navigation from pathologist-reviewed labels to select observation locations and spatial extents, avoiding costly question-specific trajectory annotations. We train this policy through alternating representation learning and proximal policy optimization, followed by fine-tuning with geometric and appearance consistency objectives to stabilize focus trajectories. During inference, the Navigator hierarchically acquires sparse observations from low to high magnification under a limited ROI budget. A Morphology Interpreter converts observations into question-conditioned evidence, while the Deliberator evaluates evidence and revises intermediate answers across magnifications. The Arbiter integrates deliberation history to produce final answers. AdaptivePath achieves state-of-the-art zero-shot performance on WSI and region pathology VQA benchmarks and reaches 80.14% accuracy for cancer subtype classification across six TCGA cohorts. In a blinded diagnostic-utility study, pathologists using AdaptivePath-selected observation sequences achieve 82.9% accuracy. These results demonstrate that learned active perception enables effective and traceable visual reasoning over gigapixel pathology slides.
170. 【2608.08646】Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking
链接:https://arxiv.org/abs/2608.08646
作者:Zhifeng Chu,Bin Wang
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:personalized location-aware services, user mobility analysis, Trajectory-User Linking, candidate users, aims to identify
备注:
点击查看摘要
Abstract:Trajectory-User Linking (TUL) aims to identify the owner of an anonymous trajectory from a set of candidate users, providing a basis for user mobility analysis and personalized location-aware services. Existing methods often learn Point of Interest (POI), temporal, and semantic features independently, make limited use of structural knowledge shared across trajectories, and compress structural and sequential information before classification. To address these issues, we propose Multi-Relational Knowledge Graph Enhanced Embedding for Trajectory-User Linking (MakeTUL), which, to the best of our knowledge, is the first attempt to introduce knowledge graph representation learning into TUL. MakeTUL organizes visit-time, POI-category, and transfer-speed information as typed relations in a multi-relational mobility knowledge graph, allowing heterogeneous mobility semantics to jointly constrain the learned embeddings. The resulting POI representations are further enriched with high-order co-occurrence patterns extracted from the trajectory collection, providing structural prior knowledge for sparse and overlapping trajectories. By integrating these prior-enhanced representations with temporal, category, and transfer information, the trajectory sequence learning module captures ordered mobility patterns, while a dual-branch classification layer preserves and combines global structural evidence and sequential evidence at the decision level.
171. 【2608.08630】VLZip: Unified Visual and Textual Compression for Interleaved Long-Context Modeling
链接:https://arxiv.org/abs/2608.08630
作者:Yuqi Zhang,Cheng Chen,Yuyu Guo,Wenjie Yang,Lingchen Meng,Peng Di,Hang Yu,Zuxuan Wu,Yu-Gang Jiang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision Language Models, Vision Language, Language Models, face significant challenges, interleaved image-text sequences
备注:
点击查看摘要
Abstract:Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer-specific "soft prefixes" and injects them into each decoder layer's hidden states, drastically shortening the attention sequence while preserving fine-grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative-level reasoning. Extensive experiments show VLZip achieves leading performance on long-context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long-context multimodal AI. Code is available at this https URL.
172. 【2608.08622】VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models
链接:https://arxiv.org/abs/2608.08622
作者:Dong Xing,Jiaxin Chen,Hang Yang,Peixun Liu,Qiushi Yang,Yuqing Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large vision-language models, open-ended video understanding, demonstrated strong performance, fluent responses unsupported, Large vision-language
备注:
点击查看摘要
Abstract:Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence. Existing training-free methods typically apply a globally fixed visual intervention or construct a contrastive branch through input perturbation. The former cannot accommodate video-dependent fusion paths, while the latter can be compensated by cross-frame redundancy. We therefore propose Video-Adaptive Debiasing via Evidence Reweighting (VADER), a training-free framework with two complementary modules. Visual Focus Reallocation (VFR) automatically instantiates an intervention policy for each video-question input: it diagnoses layer-wise visual-to-text evidence flow, determines where to intervene, and derives how strongly to reallocate pre-softmax attention from system-token to video-token blocks. Selective Evidence Erasure (SEE) independently masks high-importance visual tokens in every frame, constructing a prior-biased branch that is difficult to compensate through neighboring frames. Contrastive decoding then down-weights predictions that remain confident after selective evidence erasure. Across multiple VideoLLMs, VADER yields substantial improvements on event-level grounding and temporal consistency; on LLaVA-Video-7B, it reaches 72.60% accuracy on EventHallusion.
173. 【2608.08612】REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering
链接:https://arxiv.org/abs/2608.08612
作者:Caijun Yan,Yang Zhou,Meixing Shi,Haoran Sun,Yichen Li,Yuxiang Cai,Yankai Jiang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:long-video question answering, retrieval-augmented and memory-augmented, question answering, promising paradigms, Recently
备注:
点击查看摘要
Abstract:Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on rigid, fixed-length temporal chunking (e.g., 10s) and static offline memory banks, which not only fragment coherent continuous events but also fail to adapt during real-time reasoning. Moreover, whether using multi-scale summaries or multimodal knowledge graphs, current approaches prioritize retrieval relevance while overlooking evidence sufficiency, often stopping to answer once only semantically relevant clues are retrieved, even when key temporal, causal, or fine-grained action evidence is still missing. To tackle these challenges, we propose REVEAL, a rubric-guided agent framework. As a foundation, we introduce an adaptive visual-similarity-based preprocessing pipeline that groups visually coherent adjacent frames into natural event units to construct an offline-online video memory---capturing global video context offline while dynamically maintaining question-conditioned memory online. Built upon this structured memory, REVEAL uses an automatically constructed rubric library to explicitly verify whether retrieved evidence satisfies sufficiency criteria, pinpoints missing clues upon verification failure, and directs targeted re-retrieval for complementary information. Without any extra training, REVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments. These results show that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasoning.
174. 【2608.08600】Population-Scalable Multi-Agent World Modeling
链接:https://arxiv.org/abs/2608.08600
作者:Renjie Zhao,Yuxiang Wu,Mingyu Zhang,Jiaxin Li,Sisi Li,Yimin Sheng,Tianxi Tan,Zhenkai Zhang,Jianyi Zhu,Yong-Lu Li
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:recently achieved impressive, achieved impressive progress, fundamental scalability challenge, recently achieved, achieved impressive
备注: Technical report. Project page: [this https URL](https://rhos.ai/research/khora) . Online demo: [this https URL](https://ophilus.ai/khora)
点击查看摘要
Abstract:World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent environments introduces a fundamental scalability challenge. Existing methods generally assume a fixed number of agents during training and inference, which ties the model to a pre-determined agent population and limits inference-time scalability. Our key insight is that cross-view consistency should arise from a shared world state whose evolution does not assume a predefined number of agents, while agent-specific observations should be generated by querying this state through a unified rendering interface. Based on this insight, we propose Khora, a scalable multi-agent world model that supports inference-time expansion to arbitrary numbers of agents without retraining. Our framework decouples world-state evolution from visual rendering and introduces a population-agnostic rendering mechanism for incorporating other agent information. This design maintains cross-view consistency through the shared world state rather than through dense interactions among observation streams inside the expensive video generator, enabling approximately linear practical scaling with the number of queried views. Qualitative experiments demonstrate that our approach generalizes to unseen numbers of agents while maintaining visual quality and multi-agent consistency. We further implement a real-time interactive system to demonstrate scalable open-world simulation.
175. 【2608.08596】Goal-oriented Navigation Instruction Generation with Tour Video Priors
链接:https://arxiv.org/abs/2608.08596
作者:Fangdi Li,Juncheng Liao,Changxu Cheng,Jiazhi Wang,Senda Chen,Tao Wang,Wuyue Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Navigation Instruction Generation, natural language instructions, Instruction Generation, aims to produce, natural language
备注:
点击查看摘要
Abstract:Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily treat NIG as an auxiliary task for vision-andlanguage navigation (VLN), focusing on data augmentation or multi-task learning. However, generating navigation instructions from compact environmental priors requires meticulous spatial reasoning, especially when the target route does not simply follow the demonstrated tour, and remains challenging for current multimodal models. In this work, we introduce VideoNIG, a goal-oriented video-grounded NIG task that generates navigation instructions from ego-centric tour videos, an initial observation, and a textual or visual goal, without relying on intermediate representations such as graphs and maps. We instantiate VideoNIG in a controlled simulator benchmark with 60K tour videos across continuous indoor environments and 37K multimodal prompts with progressive difficulty levels. We further introduce a diagnostic evaluation protocol that combines text similarity, choice-based spatial consistency tests, and downstream navigation execution. To address this task, we propose a two-stage Curriculum Learning framework that decomposes the learning into foundational motion perception and long-horizon navigation reasoning. Specifically, we first employ Action Warmup for spatial action-view alignment, followed by Complexity Progression using trajectories with increasing exploratory difficulty. Extensive experiments show that existing MLLMs struggle with VideoNIG, while our approach significantly improves instruction quality across complementary diagnostic metrics. Finally, integrating VideoNIG-generated instructions with a VLN agent demonstrates the executability of this task formulation for end-to-end navigation.
176. 【2608.08589】RobustDefect-LLM: Explainable and Robustness-Aware Industrial Surface Defect Classification with Decision Support and AI-Assisted Reporting
链接:https://arxiv.org/abs/2608.08589
作者:Nazlıcan Düşünmez,Halûk Gümüşkaya
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:paper presents RobustDefect-LLM, surface-defect inspection framework, inspection framework integrating, framework integrating deep-learning, controlled AI-assisted reporting
备注: 27 pages, 11 figures, 16 tables, preprint manuscript, the source code is publicly available
点击查看摘要
Abstract:This paper presents RobustDefect-LLM, an industrial surface-defect inspection framework integrating deep-learning classification, operator-facing visual evidence, confidence-aware decision support, controlled AI-assisted reporting, traceable storage, and mobile interaction in a unified quality-control workflow. Here, robustness-aware denotes explicit evaluation under controlled image degradation and confidence-aware review routing, not an intrinsic robustness guarantee. Four transfer-learning-based convolutional neural networks, ResNet50, EfficientNet-B0, DenseNet121, and MobileNetV3-Large, were evaluated on 1,799 images from the six-class NEU-DET dataset using fixed training, validation, and held-out in-domain test partitions. MobileNetV3-Large achieved the highest numerical test accuracy (99.26%) and macro F1-score (0.9926), with a bootstrap 95% accuracy CI of 0.9815-1.0000. An exact paired McNemar test found no significant difference from DenseNet121 (p = 1.000). The selected model averaged 0.060 s per CPU forward pass (16.66 FPS). Under combined synthetic degradation, accuracy fell to 87.78% at mild intensity and below 40% at stronger intensities, revealing sensitivity to severe image-quality deterioration. Grad-CAM supplied visual evidence, while predictions with confidence below 0.90 or a top-2 margin below 0.10 were routed to HUMAN REVIEW. This conservative policy provided 12.22% automatic coverage and 100% observed selective accuracy among 33 eligible cases (95% CI: 89.43%-100.00%), while routing both observed classification errors to review. Under nominal controlled conditions, all 100 generated reports passed deterministic consistency checks, with a mean latency of 1.66 s. Results support the feasibility of the integrated workflow while emphasizing the need for calibration, repeated evaluation, and real-world industrial validation.
177. 【2608.08585】EvTrajGS: Accurate and Efficient 3D Gaussian Splatting from Unposed Event Streams
链接:https://arxiv.org/abs/2608.08585
作者:Zixuan Chen,Jiakai Zhang,Junhao Dong,Guangcong Wang,Jianhuang Lai,Yew-Soon Ong,Xiaohua Xie
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:asynchronous sensing characteristics, shown great potential, high temporal resolution, high dynamic range, temporal resolution
备注:
点击查看摘要
Abstract:Event cameras, with high temporal resolution, high dynamic range, and asynchronous sensing characteristics, have shown great potential for dense 3D reconstruction. Traditional reconstruction methods based on off-the-shelf pose estimates achieve high efficiency but produce low-fidelity results, as inaccurate pose initialization introduces cumulative reconstruction errors. In contrast, recent SLAM-style methods stabilize joint pose-scene optimization through incremental tracking and mapping, yielding higher reconstruction fidelity at the expense of considerable computational overhead. To address this trade-off, this paper presents EvTrajGS, an accurate and efficient 3D Gaussian Splatting framework for unposed event streams. Our method enables reliable joint pose-scene optimization initialized from coarse pose priors, eliminating the need for computationally expensive SLAM-style pipelines. EvTrajGS parameterizes camera motion as a continuous-time trajectory initialized from discrete camera poses, providing a unified representation for pose refinement. We then aggregate adjacent trajectory states into a temporally coupled pose, promoting temporally consistent pose updates during joint optimization. Additionally, we introduce a loss-reweighted event sampling strategy to adaptively emphasize temporally under-reconstructed intervals. Extensive experiments on both synthetic and real-world datasets demonstrate that EvTrajGS outperforms state-of-the-art methods in terms of both geometric reconstruction quality and pose estimation accuracy, achieving 3.8 dB higher PSNR, 0.1 higher SSIM, and over 40\% lower ATE RMSE while retaining high computational efficiency.
178. 【2608.08580】Where Is the Bee? Detecting Tiny Pollinators with a Single Collaborative-Head Transformer
链接:https://arxiv.org/abs/2608.08580
作者:Junsu Kim,Seungryul Baek
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:BuzzSpot Challenge, field keyframes, detect bees, ECCV, median box occupies
备注: The 1st rank method's paper at CVPPA@ECCV 2026 BuzzSpot Challenge
点击查看摘要
Abstract:The CVPPA@ECCV 2026 BuzzSpot Challenge asks us to detect bees, bumblebees, hoverflies, and moths in 1920x1080 field keyframes. Its annotations carry 2 difficulties: the median box occupies 0.16% of a frame, and bees account for 80% of the labels. To cope with the small boxes, we compare 10 recorded detector configurations on held-out keyframes; plain Co-DINO with a Swin-L backbone has the highest mAP in this comparison, so we select it. Training then addresses the bee dominance in 2 ways: fine-tuning on a crop-mosaic pool in which the combined annotation share of the 3 rare classes rises from 19.9% to 55.1%, and a class-weighted simplex equiangular tight frame (ETF) loss that pulls the projected states of matched decoder queries toward fixed class directions. The full schedule spans 12+3+2 epochs. Without inference-time ensembling or test-time augmentation, we rank first on FinalTest at 0.5062 mAP@[.5:.95].
179. 【2608.08575】CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling
链接:https://arxiv.org/abs/2608.08575
作者:Zheyang Jing,Qin Lu,Jianwang Li,Yujie Yang,Chen Yi,Shaofeng Jiang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:grouped hierarchical channel, medical image segmentation, image segmentation requires, requires the integration, long-range anatomical context
备注:
点击查看摘要
Abstract:Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing methods often model global and local features in separate modules or feature levels and perform channel recalibration independently. This may cause semantic mismatch between global context and local boundaries, insufficient channel relationship modeling, weak spatial-channel interaction, and redundant representations. We propose CDGC-Net, a 3D medical image segmentation network that combines cooperative dual-scale spatial attention with grouped hierarchical channel modeling. With-in each CDGC block, Cooperative Dual-Scale Self-Attention (CDSA) assigns attention heads to parallel local-window and global-sparse branches. The two branches capture fine spatial details and long-range anatomical context at the same feature level. Their outputs are concatenated into an $N\times C$ spatial representation and directly passed to Grouped Hierarchical Channel Attention (GHCA). GHCA organizes the channels into $r$ groups and models both within-group and cross-group dependencies. CDSA and GHCA reuse a shared key projection to maintain a consistent feature reference. Residual feature alignment subsequently integrates the refined features with the original representation. On the Synapse, ACDC, BraTS, and LA datasets, CDGC-Net achieved mean DSC values of 86.96\%, 92.91\%, 82.56\%, and 93.52\%, respectively, exceeding the next-highest reported values by 0.39, 0.47, 0.17, and 0.32 percentage points. CDGC-Net contains 25.83M parameters and 28.62G FLOPs for an input size of $64\times128\times128$, reducing these quantities by 39.87\% and 40.30\%, respectively, relative to UNETR++. These results indicate a favorable trade-off between segmentation accuracy and computational complexity.
180. 【2608.08566】On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation
链接:https://arxiv.org/abs/2608.08566
作者:Idaya Seidu,Ahmed Tahiru Issah,Charles B. Delahunt,Carine Mukamakuza
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:resource-limited settings, microscopists are scarce, remains a leading, mortality in resource-limited, expert microscopists
备注: Accepted at The Fifth Workshop on Applications of Medical AI (AMAI) 2026, a satellite event at MICCAI 2026. To appear in Springer Lecture Notes in Computer Science (LNCS)
点击查看摘要
Abstract:Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based on microscopy images thus has strong potential to improve care delivery. But for an algorithm to deploy, a necessary requirement is that it meet a suite of non-obvious (from a machine learning (ML) perspective) clinical constraints. Therefore, in close consultation with a national health center we developed a malaria diagnosis pipeline which addresses key requirements listed by the health care center but typically ignored in the ML malaria literature. In particular, it includes: (i) stopping criteria (to reduce image acquisition and time-to-result); (ii) human-in-the-loop functionality (for review and accountability); (iii) multi-species discrimination (since treatment varies by species); (iv) thick film detection (standard for microscopy); (v) computationally-efficient uncertainty calculations (to aid clinician review); and (vi) an edge device platform (since internet can be spotty in this catchment area). The mobile system performs all inference on-device using YOLOv13n deployed via TensorFlow Lite. It detects four species and white blood cells from Giemsa-stained thick blood smear images, aggregating per-image detections into slide-level parasitemia with World Health Organization (WHO)-standard quantification. This paper highlights these various clinical constraints and offers methods to address them. Evaluated on 2,739 annotated images across all four species, the system achieves mAP@0.5 of 0.863, per-image parasite count correlation of r = 0.812, slide-level r = 0.951 (soft counting, 10 images/slide), and runs entirely offline with a pipeline time of 10.27 +- 1.65 s per image.
181. 【2608.08557】OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories
链接:https://arxiv.org/abs/2608.08557
作者:Changhao Xiang,Shilin Zhang,Zheng Ma,Kanzhi Cheng,Ruize Ma,Yi Feng,Jianbing Zhang,Zhi Wang,Zhen Wu,Xinyu Dai,Lewei Lu
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:fixed image encoding, actively acquire evidence, image encoding, multimodal agents, agents to actively
备注:
点击查看摘要
Abstract:Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides effective supervision. We argue this assumption is flawed: a strong teacher often reaches the correct answer without needing its tool calls, and imitating such trajectories teaches a student that tool calls accompany correct answers, not that tool observations ground them. We present OpenVisTool, an open framework for constructing instructive visual tool-use trajectories that provide effective supervision for tool learning. The key insight is that a trajectory should be retained only if its answer is correct (outcome validity) and its tool observations causally contribute to that answer (causal utility). The framework operates in three stages: difficulty screening to select queries that are not reliably answerable without tools, domain-specific trajectory synthesis to elicit coherent tool-use trajectories, and supervision verification to jointly test both conditions. Rather than encouraging models to imitate tool calls, the resulting supervision teaches when and how visual evidence should be acquired. Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains. Across four backbones (4B-27B), fine-tuning on OpenVisTool-42K consistently improves visual tool-use performance and yields gains on two out-of-distribution benchmarks; the larger models approach leading closed-source systems. The evidence suggests that effective visual tool use is learned from causally grounded supervision rather than tool-calling patterns.
182. 【2608.08555】SC-Diff: Semantically Calibrated Diffusion for Visible-to-Infrared Image Translation
链接:https://arxiv.org/abs/2608.08555
作者:Junyin Zhang,Siyu Huang,Jianxiong Ye,Haowei Gong,Ruicheng Zhang,Deyu Meng,Chenqiang Gao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:abundant visible images, semantic, calibration, image translation, abundant visible
备注:
点击查看摘要
Abstract:Visible-to-infrared image translation provides a practical way to expand infrared training data using abundant visible images. Diffusion models are promising for this task because of their strong generative performance. However, existing diffusion-based methods typically use semantic priors only as external conditions, without explicitly regulating token interactions within the denoising network. Consequently, they struggle to preserve object locations, shapes, and semantic layouts required for reliable annotation reuse. We propose SC-Diff, a semantically calibrated latent diffusion framework that uses semantic priors for both conditional guidance and internal self-attention calibration. A pretrained SAM3 model with predefined text prompts first extracts category-specific semantic masks from visible images. These masks are merged into a semantic map and fused with the visible image as the input condition. The same map is converted into token-level semantic labels to calibrate self-attention in the denoising network. Based on these labels, we introduce Semantic-Guided Self-Attention Calibration (SGSC), which adaptively applies positive biases to query-key pairs of the same category. The query-wise calibration strength depends on the dispersion of attention across semantic categories and the attention assigned to the query's own category. The original attention scores further modulate the bias, giving greater calibration to same-category keys with stronger responses. This soft calibration reduces cross-category interference while retaining global contextual interactions, thereby improving semantic consistency in generated infrared images. Extensive experiments show that SC-Diff improves perceptual quality and produces more effective synthetic training data for downstream infrared object detection.
183. 【2608.08553】MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling
链接:https://arxiv.org/abs/2608.08553
作者:Rong Fu,Chunlei Meng,Yangchen Zeng,Xiaowen Ma,Yongtai Liu,Wangyu Wu,Shuo Yin,Zijian Zhang,Sicheng Li,Yingrui Ji,Chenhao Wang,Simon Fong
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM)
关键词:high-fidelity high-resolution videos, recover high-fidelity high-resolution, Video super-resolution, high-resolution videos, aims to recover
备注: 14 pages, 6 figures
点击查看摘要
Abstract:Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
184. 【2608.08541】Rethinking Attention Locality in Spiking Transformers
链接:https://arxiv.org/abs/2608.08541
作者:Zeqi Zheng,Zizheng Zhu,Yuping Yan,Wenxuan Pan,Zhaofei Yu,Yaochu Jin
类目:Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE)
关键词:Softmax-free Spiking Self-Attention, Spiking Transformers provide, Spiking Transformer architectures, efficient visual processing, Spiking Transformer
备注: 17 pages, 7 figures
点击查看摘要
Abstract:Spiking Transformers provide a promising paradigm for efficient visual processing with spike-driven computation, yet their Softmax-free Spiking Self-Attention (SSA) struggles to establish spatially localized token interactions. Although existing locality-enhanced SSA methods improve accuracy, it remains unclear whether they consistently induce spatial locality across layers and different Spiking Transformer architectures. Through Mean Attention Distance (MAD) analysis, we reveal that computational locality does not necessarily translate into spatial locality and show that uniformly applying the same locality enhancement overlooks architecture-dependent deployment requirements. Motivated by these observations, we propose Spatially Contiguous Local Attention with Boundary Continuity Pathway (SCLA-BCP). SCLA computes attention within non-overlapping regions of spatially adjacent tokens, while BCP facilitates cross-boundary information exchange through a lightweight convolutional pathway. Furthermore, we develop a hierarchical locality deployment strategy to effectively apply SCLA-BCP across the two major Spiking Transformer architectures. Extensive experiments on seven static and neuromorphic datasets covering classification, detection, and segmentation demonstrate consistent improvements with limited parameter and energy overhead. Notably, our approach improves mAP@50 by up to 9.50% on COCO 2017 and mIoU by up to 3.42% on ADE20K. Visualizations, MAD analysis, and ablation studies further validate its effectiveness.
185. 【2608.08531】ERF-GS: Reconstructing Fast Motion from Disjoint Event-RGB Viewpoints
链接:https://arxiv.org/abs/2608.08531
作者:Xiaoyang Bai,Zhenyang Li,Weiwei Xu,Edmund Y. Lam,Yifan Peng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Deep learning-driven representations, neural radiance fields, improved visual precision, Deep learning-driven, Gaussian splatting
备注: 18 pages, 12 figures
点击查看摘要
Abstract:Deep learning-driven representations such as neural radiance fields (NeRFs) and 3D Gaussian splatting (3DGS) have revolutionized the field of dynamic 3D scene reconstruction with improved visual precision and scalability. However, the reconstruction of fast-moving objects remains a challenge; existing methods based on conventional frame-based videos often struggle in scenarios such as sports events and animal videography. We propose an event-RGB fusion Gaussian splatting (ERF-GS) framework that integrates event information into both optimization and densification stages of the Gaussian splatting pipeline, taking advantage of novel event sensors with high frame-rate. Unlike many other event-assisted scene reconstruction methods, ERF-GS was developed using realistic simulation settings and realizes event-based learning detached from RGB inputs. This design enables its application beyond straightforward synthetic data into the realm of natural video with complex layout, low frame rates and severe motion blur. Our experiments show that ERF-GS outperforms both the 4DGS baseline and the concurrent E-D3DGS on different variants of the Neu3D and Nvidia datasets which include blurry RGB frames and disjoint RGB-event viewpoints. Our code is available at this https URL.
186. 【2608.08521】A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems
链接:https://arxiv.org/abs/2608.08521
作者:Sangiya Pararajasingham
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
关键词:commonly-separated failure modes, enrolled template due, Face recognition systems, recognition systems face, Minimum Euclidean Distance
备注: 5 pages, 3 figures
点击查看摘要
Abstract:Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of an authorized user, and disguise, where a legitimate user is rejected because their appearance differs from their enrolled template due to accessories, facial hair, illumination, or pose. This paper proposes and compares five combined feature-extraction and classification pipelines that address both problems within a single framework: PM (PCA and Minimum Euclidean Distance, MED), LPM (Local Binary Patterns with PCA and MED), HPM (Histogram of Oriented Gradients with PCA and MED), SM (Speeded-Up Robust Features with MED), and HM (Harris corner features with MED). Each pipeline follows a common two-phase process comprising pre-processing, feature extraction, feature filtering, and classification. The methods were trained on 115 subjects drawn from the FEI, Disguised Faces Database, and NUAA databases and evaluated on six test conditions covering mixed appearances, frontal faces, dark illumination, left- and right-turned poses, and photo-spoofing attempts. The HOG-based pipeline (HPM) achieved the most consistent performance across conditions, with 94.59% accuracy on mixed-appearance disguise, 81.5-93.2% across pose and illumination variants, and 91.67% on spoofing, while the LBP-based pipeline (LPM) achieved the second-highest spoofing-detection accuracy (93.2%), behind PM (96.67%), but weaker robustness to pose change. These results reveal a measurable trade-off between spoof sensitivity and disguise robustness among classical feature representations, motivating the deep-learning and cross-database extensions discussed in the concluding sections.
187. 【2608.08519】BIRD: Event-based Intensity Image Reconstruction Using Controllable Diffusion Models
链接:https://arxiv.org/abs/2608.08519
作者:Ignacio Bugueno-Cordova,Fabian Valderrama,Rodrigo Verschae
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:challenging problem due, event streams remains, streams remains, remains a challenging, challenging problem
备注:
点击查看摘要
Abstract:Intensity-image reconstruction from event streams remains a challenging problem due to the binary, sparse, and asynchronous nature of event data. This work proposes eBIRD, an event-guided reconstruction framework that combines a DDPM with ControlNet-based conditioning. We analyze generic and specialized diffusion learning strategies for handwritten digit (N-MNIST) and face (RGBE-Gaze) reconstruction using 33ms event windows. On N-MNIST, the general model achieves the best reconstruction quality (MSE 0.0052, SSIM 0.8982, PSNR 23.34dB), whereas the specialized model performs best on RGBE-Gaze (MSE 0.0161, SSIM 0.7605, PSNR 19.08dB). These preliminary results suggest that controllable diffusion models are a promising approach for event-guided intensity-image reconstruction, while highlighting that the preferred learning strategy depends on the reconstruction domain.
188. 【2608.08508】owards Adaptive Super-Resolution and Quality Assessment via Test-Time Adaptation
链接:https://arxiv.org/abs/2608.08508
作者:Ajeet Kumar Verma
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:paper presents doctoral, presents doctoral research, real-world conditions, paper presents, presents doctoral
备注: 5 pages, 5 figures, 4 tables, 34th ACM International Conference on Multimedia, Accepted as Doctoral Symposium Track
点击查看摘要
Abstract:This paper presents doctoral research on adaptive video super-resolution and perceptual quality modeling under real-world conditions. Existing video super-resolution (VSR) methods struggle to generalize under unknown degradations arising from heterogeneous devices, codecs, and network environments. We address this challenge through test-time adaptation (TTA), a unified paradigm that improves robustness and perceptual quality without retraining or high-quality supervision. Specifically, we: 1) propose a TTA-based framework for no-reference video quality assessment (VQA), where adapted quality predictions provide perceptual guidance for VSR under unseen distortions; 2) develop a transformer-based architecture for screen-content super-resolution that preserves text clarity and structural fidelity; and 3) introduce a region-aware TTA strategy that selectively refines text and non-text regions without requiring high-resolution ground truth. Experimental results across diverse benchmarks demonstrate consistent improvements in perceptual quality and readability. We also outline ongoing work toward fully adaptive video enhancement systems capable of generalizing across unseen domains.
189. 【2608.08494】Linguistically-Aligned and Visually-Grounded Preference Optimization for Clinically-Augmented Medical Report Generation
链接:https://arxiv.org/abs/2608.08494
作者:Qiang Hu,Yuxuan Luo,Yingjie Guo,Hao Wang,Qimei Wang,Qiang Li,Zhiwei Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:reliability remains constrained, existing DPO-based MRG, DPO-based MRG methods, MRG, DPO-based MRG
备注: Accepted by ECCV 2026
点击查看摘要
Abstract:Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct Preference Optimization (DPO) has emerged as a promising post-training paradigm to enhance the performance of Supervised Fine-Tuned (SFT) MRG models, existing DPO-based MRG methods typically adopt a naive preference construction that directly pairs model-generated reports with ground truth reports. This strategy inadvertently entangles critical clinical findings with clinically irrelevant linguistic characteristics, and fundamentally lacks explicit vision-language alignment. To address these challenges, we propose DPO-Clin, a novel post-training framework that focuses preference optimization on clinical findings and cross-modal alignment. First, we introduce the Entity-level Clinical Diagnostic (ECD) module to perform a precise entity-level factual diagnosis. ECD guides the generation of linguistically-aligned report preference pairs, isolating clinical discrepancies from linguistic variations. Second, to achieve fine-grained cross-modal alignment, we develop M2DPO, a retrieval-augmented multi-modal DPO variant that enforces textual preference inversion triggered by visual context switches. Third, we locate correct yet highly uncertain predicted entities and apply counterfactual modifications to construct targeted preference data for latent risk mitigation, thereby further enhancing the model reliability. Extensive experiments on two public chest X-ray datasets (MIMIC-CXR and IU X-Ray) and an in-house endoscopy dataset demonstrate that DPO-Clin significantly improves the SFT baselines on clinical-aware metrics. Furthermore, it achieves superior performance over existing DPO-based MRG methods, exhibiting robust generalizability across distinct baseline architectures and diverse medical imaging modalities.
190. 【2608.08487】RenderMatte: Exact-Alpha Rendering and Group-Relative Alignment for Image Matting
链接:https://arxiv.org/abs/2608.08487
作者:Zecheng Ren,Yafei Hu,Jianing Zhao,Ruichen Cong,Qun Jin,Yiren Song
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:visual content production, downstream creation workflows, essential enabling technology, modern visual content, foreground extraction determines
备注:
点击查看摘要
Abstract:Image matting is an essential enabling technology for modern visual content production, where foreground extraction determines the realism and editability of downstream creation workflows. However, precise alpha estimation in open-world scenes remains challenging because real foregrounds exhibit highly diverse appearances and opacity patterns. This makes existing methods struggle with semantic ambiguity and fine-grained opacity variation, especially in sparse boundary regions that are fragile and difficult to supervise. To address this gap, we present RenderMatte, a trimap-guided matting framework that adapts FLUX.1 Kontext through full-parameter fine-tuning, leveraging image editing priors for structure-preserving alpha prediction. During supervised adaptation, an alpha-edge objective preserves the latent flow-matching signal while strengthening pixel-space boundary supervision. We further introduce group-relative alpha alignment for post-training. It compares multiple mattes sampled under the same trimap condition using matting-specific rewards for alpha accuracy, boundary fidelity, trimap compliance, and compositional consistency. To overcome the lack of precise edge annotations, we construct the RenderMatte dataset, a large-scale synthetic dataset combining 3D-rendered RGBA foregrounds with diverse multi-source assets. It features exact strand-level alpha annotations and diverse background composites. Experiments show state-of-the-art performance across all benchmarks, demonstrating a scalable path toward high-fidelity matting in open-world scenes.
191. 【2608.08476】RayLift: Lifting Complementary Ray-Wise Evidence with 3D Geometry Priors for Semantic Scene Completion
链接:https://arxiv.org/abs/2608.08476
作者:Meng Wang,Hongxia Yu,Wenzhe He,Xingdong Song,Huilong Pi,Jiapeng Zhang,Ruihui Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:semantic scene completion, comprehensive scene understanding, driving and robotics, understanding for autonomous, autonomous driving
备注:
点击查看摘要
Abstract:Camera-based 3D semantic scene completion (SSC) provides comprehensive scene understanding for autonomous driving and robotics. However, existing methods often treat stereo depth estimates as deterministic geometric constraints, causing depth uncertainty and local correspondence errors to propagate directly into voxel representations. To address this issue, we propose RayLift, a framework that uses stereo geometry as a metric reference while incorporating complementary ray evidence to recover reliable 3D structures adaptively. RayLift first employs a Complementary Context Encoder that extracts geometry-aware priors from a frozen 3D vision foundation model, thereby enriching the scene context. It then introduces a Depth Ray Evidence Lifter module that jointly models geometric dissimilarity, depth confidence, and spatial uncertainty to adaptively sample and weight candidate surface locations along each camera ray. Finally, a Semantic-Aware Voxel Integrator injects the resulting ray evidence into voxel features by explicitly modeling their spatial support. Extensive experiments on SemanticKITTI and SSCBench-KITTI-360 demonstrate that RayLift achieves competitive performance and consistently outperforms existing methods.
192. 【2608.08460】InstructionCrafter: Generating Consistent and High-Fidelity Visual Instructions
链接:https://arxiv.org/abs/2608.08460
作者:Shun Okamoto,Satoshi Iizuka,Kazuhiro Fukui
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:per-frame visual quality, image sequence requires, per-frame visual, visual quality, requires the simultaneous
备注:
点击查看摘要
Abstract:Given textual task instructions, generating step-by-step visual instructions as an image sequence requires the simultaneous satisfaction of multiple properties, specifically step faithfulness, cross-image consistency, and per-frame visual quality. Existing text-to-image generation approaches rarely meet all three properties, owing to independent sampling that breaks consistency, finetuning on low-quality video that degrades per-frame quality, and frozen backbones that lack multi-step understanding. In this work, we propose InstructionCrafter, a diffusion-based framework with the key idea of separating the optimization of temporal and instructional alignment from per-frame visual quality via (1) spatial-freeze training and (2) instruction-aware adapters. Built on a pretrained video diffusion backbone, InstructionCrafter freezes the spatial layers that control per-frame detail and updates only temporal and text-conditioning pathways to learn instruction semantics and inter-step relations, which preserves the generative prior for per-frame quality and reduces trainable parameters by about 50 percent compared with full finetuning. We also introduce two lightweight adapters that enhance the model's understanding of instructional context. The Consistent Adapter aggregates textual cues from the entire instruction sequence and from neighboring steps to keep object identity and attributes consistent across frames, and the Context-Aware Temporal Adapter converts cross-attention outputs into biases for temporal self-attention, explicitly propagating inter-frame relations. Extensive experiments on two benchmark datasets demonstrate state-of-the-art overall performance on step faithfulness, cross-image consistency, and per-frame visual quality while significantly reducing noise, blur, and spurious subtitles. Our code and trained models will be publicly available.
193. 【2608.08436】FreCast: Refining Radar Echo Intensity via Phase-Preserving Amplitude Residual Diffusion for Precipitation Nowcasting
链接:https://arxiv.org/abs/2608.08436
作者:Heping Fang,Zihuai Yin,Kaicheng Mao,Peiguang Zhang,Peng Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Precipitation nowcasting predicts, estimating the occurrence, radar echo, historical radar echo, predicts the spatiotemporal
备注:
点击查看摘要
Abstract:Precipitation nowcasting predicts the spatiotemporal evolution of future radar echoes from historical radar echo sequences, thereby estimating the occurrence, development, and movement of precipitation over the near term. In recent years, deep learning has become an important approach to precipitation nowcasting. Although state-of-the-art models can generally capture the overall spatial distribution of future precipitation, their predictions still exhibit substantial biases in radar echo intensity at individual locations. This observation motivates a more targeted strategy for reducing forecast errors. Instead of regenerating an entire radar echo sequence without spatial constraints, the predicted precipitation structure can be used to guide the refinement of echo intensities at individual locations. This structure-guided refinement directly targets echo intensity biases. Accordingly, we propose FreCast, a two-stage framework for radar echo prediction. The first stage generates an initial forecast of future radar echoes. The second stage uses the spatial structure of the initial forecast as a constraint to further correct intensity biases at individual locations in the first-stage prediction. Experiments on three datasets demonstrate that FreCast achieves consistent improvements across forecast skill metrics. Qualitative results further show that FreCast better preserves rainband continuity and intense precipitation structures at longer lead times.
194. 【2608.08418】Learning Deep Modality-Shared Self-Expressiveness for Image Clustering with Textual Information
链接:https://arxiv.org/abs/2608.08418
作者:Xianghan Meng,Wei He,Zhiyuan Huang,Chun-Guang Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Leveraging textual information, Leveraging textual, promising direction, largely owing, modality-shared self-expressive model
备注:
点击查看摘要
Abstract:Leveraging textual information for image clustering has emerged as a promising direction, largely owing to the powerful representations learned by Vision-Language Models (VLMs). Existing approaches typically retrieve a textual counterpart for each image and then refine multimodal representations by directly enforcing cross-modal agreement, e.g., maximizing image-text similarity inherited from pretrained VLMs. However, such a strategy aligns heterogeneous representations across modalities without explicitly modeling the intrinsic structure within each modality and thus might yield unreliable alignment or distort modality-specific structures that are crucial for clustering. In this paper, we propose a simple but principled approach, termed deep modality-shared self-expressive model (DeepMORSE), which discovers cross-modal structures via a modality-shared self-expressive model and simultaneously learns structured representations that conform to a union of modality-specific subspaces. Moreover, we theoretically justify that the modality-shared self-expressive coefficients suppress inter-class noise towards a subspace-preserving solution, and show that mini-batch optimization procedure introduces an implicit regularization onto the self-expressive model. We evaluate our DeepMORSE on six widely used image clustering benchmarks and observe performance improvements exceeding 3% on the UCF-101, DTD-47, and ImageNet-Dogs datasets. In addition, we demonstrate the strong transferability of the learned representations by achieving state-of-the-art performance on downstream tasks such as image retrieval and zero-shot classification---without requiring any task-specific losses or post-processing. The code is available at: this https URL.
195. 【2608.08402】Agentic AI-powered flexible fiber-bundle endoscopy for high-resolution NIR-II fluorescence imaging in vivo
链接:https://arxiv.org/abs/2608.08402
作者:Yanzhao Shi,Yuanhua Liu,Sixin Xu,Wayne Jason Li,Yuyuan Chen,Danyang Xu,Zhisheng Wu,Hanze Yu,Ian Yu-Hong Wong,Simon Ying-Kit Law,Hongjie Dai,Liangqiong Qu,Feifei Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Fiber-bundle endoscopy offers, low spatial resolution, natural human orifices, honeycomb artifacts, Fiber-bundle endoscopy
备注:
点击查看摘要
Abstract:Fiber-bundle endoscopy offers a compact and flexible route for clinical fluorescence imaging through natural human orifices, but since its first report in the 1950s, it has remained limited by low spatial resolution, honeycomb artifacts, and inter-core crosstalk. The crosstalk becomes more pronounced at near-infrared-II wavelengths (NIR-II, 1000-3000 nm), a spectral window that offers superior contrast, resolution, and tissue penetration depth for biomedical imaging. Here, we present an AI-powered flexible endoscopy platform that overcomes these constraints through optical-computational co-design: optimizing ultrathin fiber bundles to mitigate crosstalk-induced image blur and enable high-fidelity image transmission across the visible-to-NIR-II spectral range, and developing an Agent-Guided Mixture-of-Experts (GAME) pipeline for honeycomb-artifact removal and image restoration. GAME provides a single restoration entry point for diverse biomedical images acquired with our endoscope, spanning cell, mouse and human samples. It dynamically routes each input to suitable restoration experts via a vision-language model, facilitating image reconstruction with a fourfold resolution improvement beyond the NyquistShannon sampling limit. The utility of our endoscope is demonstrated through in vivo NIR-II imaging of anatomical structures in mice, as well as imaging of the digital micromirror device (DMD)-projected human gastric tube and lymphatic system, paving the way for future clinical translation.
196. 【2608.08401】Anatomically Consistent Cross-Contrast Super-Resolution of Anisotropic Brain T2w MRI
链接:https://arxiv.org/abs/2608.08401
作者:Mengqi Shen,Haicheng Wang,Meghna Trivedi,Tony J. Wang,Yuanguang Xu,Yingyan Zeng,Yading Yuan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:fluid-sensitive soft-tissue contrast, brain MRI, MRI provides fluid-sensitive, radiotherapy planning, fluid-sensitive soft-tissue
备注:
点击查看摘要
Abstract:T2-weighted (T2w) brain MRI provides fluid-sensitive soft-tissue contrast that is important for neuro-oncology and radiotherapy planning. However, T2w scans are acquired with anisotropic voxels and appear blurred or stair-stepped on coronal and sagittal views, which obscures small structures and weakens any downstream 3D analysis. We propose VIPP-SR (View-Independent Patched Projection Super-Resolution), a cross-contrast guided super-resolution framework that restores the inter-plane resolution of an existing anisotropic T2w volume without an isotropic ground-truth T2w. VIPP-SR first trains a view-independent patched generator (VIP-GAN) to learn local T1c-to-T2w anatomical correspondence from high-resolution axial slices. The trained generator is then applied to axial, coronal, and sagittal views of the T1c volume to generate three orthogonal T2w estimates. Shape-preserving patching and deepest-skip removal reduce view-specific shortcuts, thereby constraining the generator to learn patch-local representations and enabling the zero-shot inter-plane transfer. Central to VIPP-SR, a projection-based optimization then enforces anatomical consistency across the three view-specific volumes, fusing them by balancing inter-plane self-consistency against per-view data fidelity. The generator is trained on BraTS-MET and evaluated on both the held-out BraTS-MET testing set and the BraTS-GLI cohort without retraining, assessing the cross-cohort generalizability. The results validate that VIPP-SR improves downstream segmentation over the real anisotropic T2w baseline, raising mean-label Dice from 0.330 to 0.465 on BraTS-MET and, zero-shot, from 0.473 to 0.563 on BraTS-GLI and ablation studies identify inter-plane self-consistency as the main source of the gain.
197. 【2608.08381】DoRF++: Spherical Representation Learning over Doppler Radiance Fields for Robust Wi-Fi Sensing
链接:https://arxiv.org/abs/2608.08381
作者:Navid Hasanzadeh,Shahrokh Valaee
类目:Computer Vision and Pattern Recognition (cs.CV); Signal Processing (eess.SP)
关键词:Channel State Information, Wi-Fi Channel State, standardize advanced WLAN, State Information, Channel State
备注:
点击查看摘要
Abstract:Motivated by the IEEE 802.11bf effort to standardize advanced WLAN sensing, interest in Wi-Fi Channel State Information (CSI) for passive, device-free, and privacy-preserving activity and gesture recognition has grown rapidly. Recent studies have shown that Doppler velocity projections extracted from CSI, which directly reflect human-motion velocity, enable more robust human activity recognition (HAR) and stronger generalization across users and unseen conditions. Nevertheless, reliable generalization under real-world variability remains a major challenge, hindering the adoption of Wi-Fi sensing in real-world applications. To address this challenge, we introduce Doppler Radiance Fields (DoRF), bringing the concept of neural radiance fields (NeRF) from computer vision into Wi-Fi sensing. DoRF models Doppler velocity projections extracted from Wi-Fi CSI as sparse and diverse virtual-camera views of human motion. It then infers a latent 3D motion sequence whose projections along learned effective Doppler directions explain the CSI-derived Doppler observations. The recovered motion is subsequently projected onto an equiangular grid of directions on the unit sphere, producing a spherical representation of the underlying motion. Since DoRF naturally defines the Doppler representation on spheres, we further introduce DoRF++, a spherical-learning design that applies spherical Transformers for activity classification. Experiments on our collected hand-gesture dataset show that DoRF++ significantly outperforms state-of-the-art Wi-Fi-based HAR methods in cross-user generalization accuracy, especially for difficult gestures in settings with a single multi-antenna receiver access point (AP).
198. 【2608.08374】Gated Spatial Redundancy Projection for Pathology Transformer Attentions
链接:https://arxiv.org/abs/2608.08374
作者:Zhiyuan Yang,Jiahao Cheng,Vincent Quoc-Huy Trinh,Mahdi S. Hosseini
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:whole-slide image analysis, Transformer models, Gated SRP, computational pathology, models are increasingly
备注: Accepted at BMVC 2026 Conference
点击查看摘要
Abstract:Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images: neighbouring patches often contain highly similar tissue type, stain, texture, and cellular composition. We identify this local spatial redundancy as a pathology-specific failure mode of self-attention, where dominant neighbourhood features can be repeatedly mixed into patch-tokens and weaken subtle diagnostic or prognostic deviations. We propose Gated Spatial Redundancy Projection (Gated SRP), a lightweight drop-in correction module for self-attention layers. For each patch token and attention head, Gated SRP estimates a local redundancy axis from neighbouring value vectors, projects the attention output onto this axis, and applies a learned signed gate to correct the redundancy-aligned component geometrically. Across five TCGA survival cohorts, Gated SRP obtains the highest mean C-index among the compared attention variants in all cohorts, with an average improvement over the base attention, while adding only +0.02% parameters. Across five slide-level classification datasets, it improves the base attention on 12 of 16 reported metrics and achieves the best AUC on three datasets. Code is publicly available at this https URL.
199. 【2608.08368】PARAGraph: Pathology-Anatomy-Aware Hierarchical Graph for Diabetic Retinopathy Grading
链接:https://arxiv.org/abs/2608.08368
作者:Ziyang Zhang,Yuankai Huo,Yalin Zheng,He Zhao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:working-age adults worldwide, making reliable severity, Diabetic retinopathy, reliable severity grading, grading clinically important
备注:
点击查看摘要
Abstract:Diabetic retinopathy (DR) remains a leading cause of vision loss among working-age adults worldwide, making reliable severity grading clinically important. Despite strong performance, most deep models formulate DR grading as image-level classification and do not explicitly model clinically grounded evidence, such as lesion types and spatial relations. In this paper, we propose PARAGraph, a Pathology-Anatomy-Aware Hierarchical Graph framework for DR grading. PARAGraph represents each image as a three-level hierarchical graph with lesion-level nodes, intermediate category and region nodes, and global anatomical and semantic nodes. To incorporate medical priors into nodes, we construct an optic disc-fovea-anchored coordinate frame that provides a scale- and rotation-normalized retinal reference system. Within this frame, lesion nodes are encoded with category, normalized area, and anatomical coordinates. To mitigate noisy lesion segmentation, PARAGraph uses a dual-fusion strategy that introduces global visual context into a graph semantic node and a decision-level prediction branch, improving robustness when lesion evidence is unreliable. Extensive experiments on Messidor-2, APTOS, and DDR show that PARAGraph achieves consistent DR grading performance over state-of-the-art methods. Interpretability and robustness analyses further demonstrate that its predictions are clinically grounded, closely associated with lesion evidence and robust to lesion segmentation noise.
200. 【2608.08366】VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression
链接:https://arxiv.org/abs/2608.08366
作者:Xin Luo,Yicheng Tao,Haoxuan Zeng,Suyuan Wang,Chenzi Ouyang,Meiqi Zhu,Kai Liu,Shuibing Chen,Jie Liu
类目:Computer Vision and Pattern Recognition (cs.CV); Genomics (q-bio.GN)
关键词:Spatial transcriptomics, limited to targeted, number of samples, transcriptomics can resolve, small number
备注:
点击查看摘要
Abstract:Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. HE imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from HE images using paired Xenium data. VOICE first aligns cell centered HE morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.
201. 【2608.08354】ropical Cyclone Forecasting via Latent Rectified Flow using Satellite Imagery and Atmospheric Fields
链接:https://arxiv.org/abs/2608.08354
作者:Meheru Zannat,Sk. Md. Masudul Ahsan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Tropical cyclones, changing climate, cyclones are growing, growing more destructive, efficient forecasting
备注: 9 pages, 2 figures, 6 tables
点击查看摘要
Abstract:Tropical cyclones are growing more destructive in a changing climate, and efficient forecasting of their structure and track has become a necessity. Deep generative models promise an alternative to computationally expensive numerical weather prediction (NWP), yet current systems produce either satellite imagery or atmospheric fields, never both; they need many sampling steps, putting them out of reach of modest hardware; and their storm tracks come from regression heads with no physical link to the generated atmosphere. This work presents a single-pass model that jointly forecasts GRIDSAT-B1 infrared imagery and four ERA5 atmospheric fields (U-wind, V-wind, air temperature, and surface pressure) out to nine hours. A five-channel variational autoencoder compresses each 5 x 256 x 256 frame to a 4 x 64 x 64 latent, and a conditional rectified-flow UNet with a factorized temporal-attention module predicts the next three frames from three past frames, their best-track coordinates, and timestamps. The model is then reward-fine-tuned (DRaFT) against a differentiable track error derived from the predicted winds through a steering-flow calculation. On held-out 2022 storms the model reaches 16.35 dB PSNR and 0.759 SSIM, ahead of a reproduced cascaded-diffusion baseline at every lead time (+0.84 dB at +9 h) while sampling ~30x faster (56 ms vs. 1673 ms). Track error at +9 h is 62.4 km, 15% below the baseline, and a reward fine-tuning study demonstrates a further 8-11% track-error reduction across sampler budgets.
202. 【2608.08336】Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation
链接:https://arxiv.org/abs/2608.08336
作者:Uri Z. Kialy,Gil Ben-Artzi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:adapting Vision Transformers, Vision Transformers, adapting Vision, Parameter-Efficient Fine-Tuning, PEFT
备注:
点击查看摘要
Abstract:Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter count has been the dominant efficiency metric in PEFT, it does not imply \textit{compute efficiency}: parameter-sparse methods can still incur full-model training cost per step, and typically need long schedules to reach peak accuracy. We introduce Circuit Fine-Tuning (CFT), a compute-efficient framework that uses circuit discovery---conventionally used to explain trained models---to select modules for fine-tuning before training. Whereas attribution is conventionally formulated against a trained task head, we formulate it against a near-zero-initialized probe head, which isolates the response of the backbone to the target distribution rather than the preferences of a particular classifier. CFT then fine-tunes only the recovered subgraph. CFT needs no learning-rate warmup and reaches peak accuracy in ${\sim}20$ epochs on average---versus $44$--$96$ for strong PEFT baselines---yielding $2.3$--$6.6\times$ fewer training FLOPs and up to $16\times$ less wall-clock time, while adding zero parameters and no inference operations. Experiments across a standard visual transfer benchmark (VTAB-1k), hierarchical backbones (Swin), domain-shifted medical imaging (CBIS-DDSM), and a vision-language model (Gemma-3 on CUB-200) demonstrate the effectiveness of CFT. Code is available at this https URL
203. 【2608.08319】A continually expandable foundation model for brain MRI
链接:https://arxiv.org/abs/2608.08319
作者:Michail Mamalakis,Carmen Jimenez-Mesa,Yonghao Li,Hao Chen,Chao Li,Antonios Mamalakis,John Suckling,Richard Bethlehem,Stephen J. Price,Richard J. Gilbertson,Pietro Lio
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:brain MRI foundation, Brain magnetic resonance, magnetic resonance imaging, MRI foundation, brain MRI
备注:
点击查看摘要
Abstract:Brain magnetic resonance imaging (MRI) is central to neuroscience and clinical assessment, but models are commonly developed for individual diseases, populations or imaging protocols. Foundation models promise more general representations, yet they are usually pretrained once and can lose earlier capabilities when updated with new data. Here we show that Alcmaeon, a three-dimensional brain MRI foundation model pretrained without manual labels on more than 425,000 volumes and derived imaging maps, can be expanded sequentially across clinical domains. Alcmaeon combines volumetric encoding and latent diffusion generation with Graph-Blueprint Pruning (GBP), which protects network modules important to earlier domains while leaving the remaining capacity trainable. Across expansion from healthy ageing and neurodegeneration to developmental, psychiatric and tumour imaging, GBP showed less forgetting than sequential adaptation and elastic weight consolidation across voxel-level reconstruction measures, with its largest advantage after adaptation to tumour imaging. The blueprints provided an inspectable record of how model capacity was protected and reused. Representations from different model levels supported image synthesis, disease classification, survival modelling and postoperative prediction, although no single representation was optimal for every task. These findings provide a route towards brain MRI foundation models that can grow with emerging data while retaining earlier capabilities.
204. 【2608.08315】Your VLM Already Knows When: Training-Free Temporal Grounding by Asking Yes or No
链接:https://arxiv.org/abs/2608.08315
作者:Ji Huang,Barry Devereux,Hui Wang
类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:Multimodal LLMs, recognise events reliably, LLMs that recognise, reliably still fail, recognise events
备注:
点击查看摘要
Abstract:Multimodal LLMs that recognise events reliably still fail to say when they happen. Prompted for timestamps, strong VLMs reach as little as $3.8\%$ R@0.5 on Charades-STA, and $77$ to $80\%$ of their wrong predictions carry low output entropy: the models are confidently wrong, and entropy-based error detection stays below a random classifier. We show that this failure lives in the task interface, not in perception. Holding the weights fixed, replacing timestamp regression with a coarse-to-fine scan of binary questions, whose first-token probabilities are consumed only as a ranking, raises R@0.5 by $28$ to $50$ points across four frozen backbones. The residual failures decompose into two measurable axes: a perception axis that moves with the backbone, and a geometry axis that is analytically predictable from the ratio of the output-window and event widths. FV-Action, the training-free method built on this analysis, reaches $56.8\%$ R@0.5 on Charades-STA, above the same backbone's native grounding pipeline and the strongest training-free result on this benchmark; it surpasses every TVG-trained model evaluated zero-shot on TACoS, and improves over direct prediction on ActivityNet Captions and QVHighlights, with no temporal supervision at any stage.
205. 【2608.08309】hree Necessary Principles for Self-Supervised Visual Representation Learning
链接:https://arxiv.org/abs/2608.08309
作者:Nikos Giakoumoglou,Paschalis Giakoumoglou,Tania Stathaki
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:learning visual representations, signal jointly complete, training signal jointly, patch-level spatial prediction, augmented views
备注: Accepted at the 19th European Conference on Computer Vision (ECCV 2026) Workshops
点击查看摘要
Abstract:We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objectives: semantic invariance across augmented views, patch-level spatial prediction, and representational non-degeneracy. We formalize these as the observation, prediction, and regularization principles and prove (i) that combining observation and prediction without regularization admits the constant encoder as a global minimizer under negative-free alignment; (ii) that the two objectives are gradient-complementary and structurally non-conflicting at the encoder output; and (iii) that the momentum encoder converges to the same fixed point as the online encoder and provides no collapse guarantee at convergence. Contrastive alignment provides only self-limiting collapse resistance, formalized via an explicit gradient-decay argument. Dropping prediction withholds the spatial training signal by construction; dropping observation forfeits cross-view semantic invariance by construction; at the scale we study, no pair substitutes for the third. Every major self-supervised method is a special case of a single unified energy decomposition. We pair every theoretical claim with a controlled experiment, including a patch-retrieval evaluation for the spatial consequence of prediction.
206. 【2608.08308】Open-World Semantic Segmentation with Sensitivity Modeling
链接:https://arxiv.org/abs/2608.08308
作者:Anastasios Romanos Varvarigos,Nikos Giakoumoglou,Tania Stathaki
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Modern vision systems, Modern vision, vision systems, anomalous content, segmentation models operate
备注: Accepted at the 2026 IEEE International Conference on Image Processing (ICIP) Satellite Workshops
点击查看摘要
Abstract:Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous content. Conventional semantic segmentation models operate under a "closed-world" assumption, often producing overconfident misclassifications on novel content. We address open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design. The first decoder performs closed-set segmentation using Gaussian prototypes for known categories. The second uses contrastive feature learning to isolate unknown regions in embedding space. The third, our key contribution, is a sensitivity decoder that captures fine-grained texture irregularities and activation instabilities indicative of semantic uncertainty, which neither semantic prototypes nor contrastive norms can reliably detect. The three decoders provide genuinely complementary signals: class-level OOD distance in logit space, global feature energy in embedding space, and local activation instability across encoder scales. Experiments on Cityscapes and BDD-Anomaly show that our method improves anomaly segmentation and novel-class discovery while maintaining competitive closed-set accuracy, with gains of +2.4% AUROC and a 2.5 pp. reduction in FPR@95TPR on BDD-Anomaly over the baseline.
207. 【2608.08307】Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering
链接:https://arxiv.org/abs/2608.08307
作者:Yusra Tariq,Rakesh Chandra Joshi
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:including lesion texture, requires aligning subtle, Visual Question Answering, subtle visual evidence, aligning subtle visual
备注: 7 Pages, 4 figures, under review at AAAI 27
点击查看摘要
Abstract:Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse density changes, with clinical language. Existing multimodal fusion approaches operating in the spatial domain may not fully exploit complementary frequency information present in visual and textual representations. We introduce a dual-branch frequency-domain fusion module that conditions spectral filtering on the input question, enabling adaptive selection of global low-frequency structure and fine-grained high-frequency detail before reconstructing the spatial representation for answer generation. To provide a richer spectrum for filtering, we extract complementary features from early texture-sensitive and final semantic layers of a frozen BiomedCLIP encoder and align both with the question representation using a symmetric InfoNCE objective prior to staged joint training with a BioBART decoder. We pretrain the proposed model on PMC-VQA and fine-tune it on the VQA-RAD and SLAKE benchmarks, demonstrating that frequency-aware multimodal fusion improves medical VQA performance while maintaining a lightweight and efficient architecture.
208. 【2608.08294】A Controlled Study of Feature-Based Knowledge Distillation Across Student Designs
链接:https://arxiv.org/abs/2608.08294
作者:Abhinand Balachandran,Praveen Prashant
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:Knowledge distillation trains, Knowledge distillation, distillation trains, trains a smaller, match the outputs
备注: 7 pages, 3 figures, 2 tables
点击查看摘要
Abstract:Knowledge distillation trains a smaller student to match the outputs of a larger teacher. Feature-based methods also align intermediate representations, but this extra constraint may affect students differently. We study this question on CIFAR-100 using a ResNet-50 teacher, a width-controlled CustomResNet family and MobileNetV2 as a cross-design comparison. For each student, we evaluate each feature method against a matched logit-KD run using the same teacher, optimizer settings, training schedule and seed. We repeat the main comparisons across multiple seeds. Logit KD improved every tested student over its scratch baseline. Attention Transfer showed no clear relationship with size inside the CustomResNet family, but its average effect was negative for that family and positive for MobileNetV2. FitNets was below logit KD in all 15 paired runs. Within the constant-depth width sweep, its gap increased for wider students, although the different-depth w=48 student did not follow this trend. Finally, the same auxiliary coefficient produced different gradient scales across students, showing that a fixed coefficient does not create a uniform training condition.
Comments:
7 pages, 3 figures, 2 tables
Subjects:
Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2608.08294 [cs.LG]
(or
arXiv:2608.08294v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2608.08294
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
209. 【2608.08290】st-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation
链接:https://arxiv.org/abs/2608.08290
作者:Haozhe Wang,Jintao Cheng,Weibin Li,Xiaoyu Tang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:additional labeled supervision, Open-vocabulary semantic segmentation, pretrained CLIP encoder, Open-vocabulary semantic, repurposes a pretrained
备注: 17 pages, 12 figures, preprint
点击查看摘要
Abstract:Open-vocabulary semantic segmentation (OVSS) repurposes a pretrained CLIP encoder for dense prediction without additional labeled supervision. Existing methods improve CLIP's spatial behavior either by redesigning its internal attention or by injecting features from auxiliary vision foundation models; both require access to the host's internal computation and are tailored to its specific forward pass. In this work, we propose Test-time Prototype Adaptation (TPA), a training-free plug-in that operates at the output level, leaving the host's forward pass and weights unmodified. By leveraging a lightweight transductive adaptation phase, TPA identifies confident anchor patches from the host's own output predictions on a small pool of unlabeled deployment-domain images, and aggregates their frozen DINO features into per-class prototypes; at inference, a single cosine similarity lookup against this frozen bank provides an auxiliary score fused linearly with the host's logits. TPA composes with five representative OVSS hosts spanning attention-redesign and VFM-injection designs, across three CLIP backbones, eight benchmarks, and multiple internal VFM choices. Under a single set of hyper-parameters and without per-host tuning or parameter updates, TPA consistently improves segmentation accuracy, with as few as approximately 10% of unlabeled deployment-domain images sufficing for effective bank construction on most benchmarks.
210. 【2608.08287】What Irregularity Costs: CUDA C++, Rust, and Triton on a Hash-Blocked GPU Workload
链接:https://arxiv.org/abs/2608.08287
作者:Petr Korolev(Spacial Intelligence Labs)
类目:Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF); Programming Languages (cs.PL)
关键词:GPU language comparisons, dense linear algebra, tiled dense linear, GPU language, linear algebra
备注: 23 pages, 5 figures, 4 tables. Includes a correctness result for TSDF fusion implementations: at hash load factors reached by ordinary depth trajectories, the Triton implementation silently discards blocks. Code, raw measurement CSVs and an interactive viewer: [this https URL](https://github.com/realitymatrix/what-irregularity-costs)
点击查看摘要
Abstract:GPU language comparisons are almost always run on tiled dense linear algebra, where every toolchain is good and the differences are small. We implement the same hash-blocked TSDF fusion kernel in CUDA C++, in Rust through NVIDIA's cuda-oxide, and in Triton, and measure it on a workload with the opposite character: an open-addressed hash table with compare-exchange insertion, data-dependent per-lane probe depth, and contended scatter. The result is a split. On the regular stage, which walks a truncation band and accumulates, all three languages land within a small factor of each other. On the irregular stage, which probes and inserts, Rust stays close to hand-written CUDA C++ while Triton is more than an order of magnitude slower. Language choice is nearly free on the work that is usually benchmarked and expensive on the work that is not. We attribute both gaps to specific things the languages cannot express, not to ratios. Triton's cost follows from a probe loop that must run to a compile-time bound and from tl.atomic_cas taking no mask, which forces a scratch structure with no counterpart in CUDA. Rust's cost was invisible in every instruction count: its kernel issues fewer instructions, fewer compare-exchanges and fewer registers at identical occupancy, yet was slower. Hardware counters located it in L1 residency. A GPU-scope atomic load must be coherent across SMs, no NVIDIA L1 is, so the type-correct way to read a shared location bypasses the cache on every access. Triton's bounded probe is also a correctness problem for fusion: at load factors an ordinary depth trajectory reaches, it silently discards blocks and the reconstruction loses patches of surface with nothing reported. We also report a defect found and fixed in cuda-oxide itself, now merged upstream: its scoped atomic load and store could not be called at all in the build mode that produces real kernels.
Comments:
23 pages, 5 figures, 4 tables. Includes a correctness result for TSDF fusion implementations: at hash load factors reached by ordinary depth trajectories, the Triton implementation silently discards blocks. Code, raw measurement CSVs and an interactive viewer: this https URL
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Distributed, Parallel, and Cluster Computing (cs.DC); Performance (cs.PF); Programming Languages (cs.PL)
ACMclasses:
D.3.4; I.4.8; C.1.2
Cite as:
arXiv:2608.08287 [cs.CV]
(or
arXiv:2608.08287v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.08287
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Petr Korolev [view email] [v1]
Sat, 8 Aug 2026 18:39:16 UTC (2,526 KB)
211. 【2608.08285】Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
链接:https://arxiv.org/abs/2608.08285
作者:Gunjan Paul,Senthil Palanisamy,Satpal Singh Rathore,Pratyush Kumar Patnaik,Shubhanshu Khatana,Abhishek Anand
类目:Computer Vision and Pattern Recognition (cs.CV); Hardware Architecture (cs.AR); Robotics (cs.RO)
关键词:head-mounted stereo-inertial capture, embedded Linux SBC, head-mounted stereo-inertial, egocentric data collection, stereo-inertial capture device
备注:
点击查看摘要
Abstract:We present Ego-OSCAR, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild. EgoOSCAR pairs a hardware-synchronized global-shutter stereo camera with a 6- axis IMU, an embedded Linux SBC for on-device video encoding, and a realtime microcontroller for user feedback and watchdog functions. The complete bill of materials is under USD 200 per unit, using only commercially available components and 3D-printed parts. Alongside the device, we release a complete software stack (hardware-accelerated recording pipeline, IMU sampling daemon, time-synchronization tooling, and watchdog firmware) and roughly 550 hours of egocentric stereo video per camera with synchronized IMU, collected by a distributed contributor network across everyday indoor environments. The release is annotated rather than raw: free-form action captions cover essentially the entire recorded timeline with an open vocabulary, and per-frame 3D hand reconstructions ship alongside per-session stereo calibration. Ego-OSCAR does not aim to match the per-unit fidelity of research-grade systems such as Project Aria; it aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale. All hardware designs, software, and the dataset are open-sourced
212. 【2608.08273】Action- and Language-Conditioned Video Assessment for Embodied Control
链接:https://arxiv.org/abs/2608.08273
作者:Hwanhee Kim,Jaehyun Jang,Seungmin Cha,Hyeonseo Yun,Donghoon Lee,Chang D. Yoo
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision-based embodied agents, embodied agents executing, agents executing multi-step, Vision-based embodied, executing multi-step natural
备注: 21 pages, 7 figures, Published in Sensors
点击查看摘要
Abstract:Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete trajectories. Conventional approaches based on final-frame matching or continuous embedding similarity may overlook intermediate transitions that are necessary for determining whether an instruction has been completed. We propose ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction. The method uses a pre-trained vision-language model (VLM) in two stages: it first summarizes frame-to-frame visual transitions conditioned on the executed actions and then assesses the generated summary with respect to the instruction to produce a discrete trajectory-level progress score. In simulated 3D household environments, ALVA exhibits a conservative assessment pattern with near-zero false-positive rates. When used as terminal feedback for closed-loop policy optimization, it provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces the performance gap to a ground-truth oracle. These results support action- and language-conditioned video assessment as an interpretable feedback mechanism for the evaluated simulated embodied-control tasks.
213. 【2608.08219】VTO: Visual Tool Orchestration for Video Anomaly Detection
链接:https://arxiv.org/abs/2608.08219
作者:Rui Wang,Yeteng Wu,Xianling Zhang,Mengshi Qi
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Video anomaly detection, Video anomaly, challenging task due, critical yet challenging, challenging task
备注: Accepted by ACM MM 2026
点击查看摘要
Abstract:Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional deep learning approaches are fundamentally limited by poor generalization across diverse scenarios. While multimodal agents offer a promising tool-learning paradigm for VAD, current systems relying on supervised fine-tuning struggle with complex orchestration, and standard reinforcement learning often causes premature termination due to coarse-grained outcome rewards. To address these challenges, we propose VTO, a process-supervised reinforcement learning framework. Moving beyond static tool usage, VTO enables the agent to dynamically explore and interact with the environment. Specifically, we introduce a foundation model-driven cognitive evaluator to provide context-aware semantic feedback, which is seamlessly integrated into a Process-Supervised Cognitive Alignment that delivers fine-grained, step-wise supervision. By explicitly penalizing logical truncation and rewarding complete causal chains, the agent optimizes its multi-step reasoning policy for interrelated tool orchestration. To support our proposed framework, we meticulously crafted VAD-Tool, a hierarchical visual tool set comprising 12 specialized vision tools spanning from entity tracking to high-stakes hazard detection, and established the corresponding benchmark for rigorous multi-step reasoning evaluation. Extensive experiments on VAD-Tool demonstrate that VTO significantly outperforms baselines, achieving up to a 10.2\% absolute accuracy improvement in tool scheduling. Code and data are available at this https URL.
214. 【2608.08191】BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation
链接:https://arxiv.org/abs/2608.08191
作者:Satvik Praveen,Shengji Jin,Ahmed Lamidi,Xin Qian,Yi Sheng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:anatomically adjacent structures, localized boundary errors, segmentation remains challenging, Multi-organ ultrasound segmentation, ultrasound segmentation remains
备注: 9 pages, 5 figures. Source code available at [this https URL](https://github.com/SatvikPraveen/BAP-MOS)
点击查看摘要
Abstract:Multi-organ ultrasound segmentation remains challenging when anatomically adjacent structures must be delineated jointly, as localized boundary errors can persist even when Dice scores are high. To address these challenges, we propose Boundary-Adaptive Prompting for Multi-Organ Segmentation (BAP-MOS), a closed-loop adaptive prompting framework. BAP-MOS formulates prompt selection as an organ-specific multi-armed bandit problem over box, point, and combined prompts. An outer Tree-structured Parzen Estimator (TPE) loop selects the prompt-selection parameter vector, while an inner UCB-Tuned loop adapts per-organ prompt preferences during fine-tuning using a bounded Dice--MSD--HD95 validation-probe reward. The framework further introduces an organ-scaled negative prompt ring to adapt sparse prompt geometry across anatomical scales, while keeping the image and prompt encoders frozen and updating only the mask decoder. We evaluate BAP-MOS on pooled prostate-region TRUS cohorts against U-Net, nnU-Net, MedSAM, fixed-prompt SAM/MedSAM, and adaptive policy variants. On this benchmark, BAP-MOS achieves Dice 0.982, HD95 0.482, and MSD 0.204, reducing HD95 by approximately 48% and MSD by 45% relative to the strongest conventional baseline. To verify the generalization ability of the framework, we tested it on the external PFUS1 pelvic-floor ultrasound corpus using MedSAM and its adaptive strategy variants, and the results were good. These results support adaptive prompt allocation as an effective mechanism for improving boundary-sensitive multi-organ ultrasound segmentation without modifying the foundation-model backbone. Source Code is available at: this https URL
215. 【2608.08167】Wiener Representation Filtering for VLM Hallucination Suppression
链接:https://arxiv.org/abs/2608.08167
作者:Ameen Ali,Tamim Zoabi,Lidor Brami,Lior Wolf
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Vision-language models, excel at open-ended, open-ended captioning, captioning and visual, relations absent
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) excel at open-ended captioning and visual QA but often describe objects, attributes, or relations absent from the image, a phenomenon known as object hallucination. We propose a {training-free, post-hoc representation editing technique} that operates in the representation space of the language backbone. The method performs a lightweight, one-time offline calibration on a modest paired dataset to estimate the required covariance structures, using only forward passes and empirical second-order statistics with no gradient updates or fine-tuning, after which the correction is absorbed directly into the model's existing weights. By modeling hidden states as a superposition of truthful and hallucination-associated components, we derive a Wiener-type estimator whose optimal gains are given in closed form from the covariances of paired truthful and hallucinated representations. An eigendecomposition yields mode-wise attenuation that respects a stability criterion, i.e., the filter responds continuously to estimation noise. The correction is applied once to the feed-forward output projections of selected deeper layers, at inference time, the model runs unchanged and at the same speed. Experiments on LLaVA-1.5, MiniGPT-4, Gemma3, and mPLUG-Owl2 demonstrate consistent reductions in object hallucination on CHAIR, POPE, and MME while maintaining caption fluency and overall response quality. We further demonstrate the generality of our approach on the TempCompass video understanding benchmark and on discrete diffusion language models for grounded dialogue, showing that representation filtering reduces hallucinations even in temporal video reasoning and multi-step, sequence-wide denoising settings.
216. 【2608.08153】Learning Structural Illumination for Unsupervised Low-light Enhancement
链接:https://arxiv.org/abs/2608.08153
作者:Tianle Du,Peiyuan He,Hainuo Wang,Tianxiu Yu,Xiaojie Guo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Existing unsupervised low-light, preventing unreliable low, relative illumination structure, varying illumination pattern, estimate illumination directly
备注:
点击查看摘要
Abstract:Existing unsupervised low-light image enhancement (LLIE) methods often estimate illumination directly from the entire low-light input, without separating its spatially varying illumination pattern, termed relative illumination structure, from the absolute exposure level or preventing unreliable low signal-to-noise ratio regions from biasing the estimate. Moreover, fixed exposure targets impose a scene-agnostic enhancement criterion, limiting adaptation across diverse lighting conditions. Inspired by the spatial propagation of light, we propose a Relative Illumination Structure Estimation (RISE) framework that decouples relative illumination structure from absolute exposure and infers it from reliable bright regions, enabling interpretable and robust enhancement. For scene-adaptive exposure adjustment, we further propose a Dual-Metering Exposure Reference derived from each input, allowing RISE to adapt the enhancement strength to individual scenes and generalize across diverse lighting conditions. Extensive benchmark and real-world generalization experiments show that RISE achieves state-of-the-art performance among unsupervised LLIE methods while producing visually natural results.
217. 【2608.08138】EFFEKT: Efficient Federated Knowledge Transfer to Foundation Models
链接:https://arxiv.org/abs/2608.08138
作者:Matteo Caligiuri,Francesco Barbato,Pietro Zanuttigh,Francesco Restuccia
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Recent data protection, Recent data, data protection laws, privacy-preserving decentralized training, Federated Learning
备注: 12 main content pages, 8 appendix pages; 3 main figures, 9 appendix figures; 8 main tables, 9 appendix tables; 1 main algorithm, 4 appendix algorithms; accepted at TMLR
点击查看摘要
Abstract:Recent data protection laws have accelerated the adoption of Federated Learning (FL) for privacy-preserving decentralized training. Nevertheless, increasing model sizes impose substantial computational demands on client devices, limiting FL applicability in resource-constrained settings. We introduce a novel multi-domain federated learning framework in which lightweight client-side proxy models collaborate with a server-side Foundation Model (FM) to learn new concepts without sharing private data. Our approach, EFFEKT, enables efficient server-side training of domain-specific LoRA adapters while preserving feature-space alignment between the FM and proxy extractors via novel bi-directional cross-distillation strategies. Experiments on multiple real-world datasets and deployments on low-power edge devices demonstrate improvements over state-of-the-art baselines in most considered domains while maintaining lightweight computation at the client side.
218. 【2608.08135】Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching
链接:https://arxiv.org/abs/2608.08135
作者:Daniele Molino,Alessio Zoboli,Camillo Maria Caruso,Valerio Guarrasi,Paolo Soda
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Cross-modality medical image, field remains constrained, medical image translation, multi-modal acquisitions, coupled limitations
备注: Accepted ad Sashimi 2026
点击查看摘要
Abstract:Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two coupled limitations: methods operate on 2D slices or 3D patches rather than whole volumes, and train a separate model for each translation task. Both stem from a single cause, the absence of a sufficiently strong volumetric prior, which forces generative models to learn anatomical appearance and cross-modality mapping simultaneously, an ill-posed problem at the scale of available paired datasets. We propose to decouple these objectives. A large-scale pretrained 3D variational autoencoder provides a compact latent representation of volumetric appearance, reducing translation to a conditional flow-matching problem. This compression makes whole-volume processing tractable, while a resolution-aware sampling strategy preserves native anatomical scale. We train a single model jointly across inter-modality (MRI$\to$CT, CBCT$\to$CT) and intra-modality (MRI$\to$MRI) tasks over three multi-center datasets. Across all tasks, whole-volume processing outperforms its patch-based counterpart, and the multi-task model matches task-specific baselines while replacing $N$ networks with one. Crucially, joint training unlocks capabilities inaccessible to task-specific approaches: zero-shot generalization to anatomical regions unseen during training, within 0.15 SSIM of the fully supervised model, and compositional cross-dataset translation along paths never directly supervised. These results suggest that combining a strong volumetric prior with multitask training is a scalable route toward synthesis systems that generalize beyond their training distribution. Code is available at this https URL.
219. 【2608.08132】When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery
链接:https://arxiv.org/abs/2608.08132
作者:Y Huynh,Duc Thanh Nguyen,Thao Minh Le,Mohamed Abdelrazek
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:challenging research problem, computer vision, challenging research, research problem, problem in computer
备注:
点击查看摘要
Abstract:Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical information from viewpoints to complete 3D structures. Using an additional view may help to resolve the issue. However, there is no mechanism that can integrate the extra view into the single-view 3D reconstruction principle. We address this challenge by proposing ASV3D, a framework for adapting single-view 3D object reconstruction to test-time data with support from one additional image. We introduce two adaptation strategies: (i) a zero-shot adaptation scheme that leverages the auxiliary image to improve the reconstruction quality of an object without retraining, and (ii) an optimised adaptation scheme that further enhances visual fidelity and cross-view consistency via contrastive learning. We apply our ASV3D to improve two state-of-the-art single-view 3D reconstruction pipelines on both benchmark and real-world datasets. Results demonstrate that our approach consistently improves reconstruction accuracy and robustness under unconstrained multi-view inputs, outperforming the baselines in both quantitative metrics and human preference. We publish our code and the real-world object dataset in our project page at this https URL.
220. 【2608.08125】Staying True to the Origin: Continuous Image Stylization with Smooth Transitions
链接:https://arxiv.org/abs/2608.08125
作者:Rui Xu,Hanmo Zhang,Songhua Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:achieved remarkable performance, Recent advances, performance in text, advances in generative, achieved remarkable
备注: 14 pages, 12 figures
点击查看摘要
Abstract:Recent advances in generative models have achieved remarkable performance in text- and image-conditioned editing. However, preserving the content of a given image while referencing style patterns from another remains challenging, often leading to uncontrollable stylization results. In this paper, we approach image stylization from the perspective of continuous control, aiming to enable modern Diffusion Transformer (DiT)-based multi-reference editing models to (1) faithfully preserve the semantic structure of the content image, (2) render strong stylization effects, and (3) smoothly transition between the two. To this end, we propose a simple yet effective two-stage training strategy along with a style-strength-aware spline formulation. Specifically, in the first stage, the model is trained to produce strongly stylized outputs while preserving the content semantics as much as possible. In the second stage, with the base model frozen, we learn a set of anchor projectors that map various stylization strengths into the model parameter space. During inference, by performing style-strength-aware spline interpolation in a low-rank space, our method enables continuous control over stylization strength, even though the model is trained with only a few discrete strength levels. Extensive experiments demonstrate that our method supports precise and continuous manipulation of stylization strength while generating high-fidelity results with modern DiT models. Project page: this https URL.
221. 【2608.08115】SUMI: Scalable Unified Model for 3D Point Cloud Inference
链接:https://arxiv.org/abs/2608.08115
作者:Yanlong LI,Kanchana Thilakarathna
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:low-density coarse shape, target resolution, cloud completion commonly, Point cloud completion, coarse structural features
备注:
点击查看摘要
Abstract:Point cloud completion commonly follows a coarse-to-fine paradigm, where a low-density coarse shape is first predicted and then upsampled to the target resolution. Although recent methods have improved global structure recovery, the fine stage often remains limited by simple upsampling and insufficient interaction with coarse structural features, making local detail reconstruction challenging. We propose SUMI, a diffusion-enhanced refinement module for coarse-to-fine point cloud completion. Unlike prior diffusion-based completion methods that use diffusion as a standalone point generator, SUMI injects noisy geometric features into cross-attention with coarse structural features, enabling reverse denoising to refine local geometry while preserving global consistency. SUMI can also be integrated into existing coarse-to-fine models as a flexible refinement module. Experiments on PCN, ShapeNet-55/34, and MVP demonstrate consistent improvements over strong baselines. SUMI achieves the best overall CD and F1-score on PCN, reduces CD by up to 16.1% on ShapeNet-55, and obtains the best CD across all output densities on MVP.
222. 【2608.08106】SCTD 3.0: Sonar Common Target Detection in the Wild - A Large-Scale, Multi-Scene Dataset from Real Marine Surveys
链接:https://arxiv.org/abs/2608.08106
作者:Peng Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Synthetic Aperture Sonar, Synthetic Aperture, Aperture Sonar, core for wide-area, Sonar Common Target
备注:
点击查看摘要
Abstract:Synthetic Aperture Sonar (SAS) is core for wide-area detection of small underwater targets. However, large-scale, high-quality SAS datasets are scarce, hindering data-driven recognition. Existing benchmarks are small and limited to single scenarios, failing to reproduce complex acoustic scattering, diverse seabeds, and multi-pose imaging in real detection. To fill this gap, we introduce SCTD 3.0 - a large-scale real-measured dataset for Sonar Common Target Detection in the Wild in natural waters. It contains over 10,000 high-quality real SAS image snippets from multi-frequency systems (240 kHz, 450 kHz, and others), covering ten typical target categories across varied seabed geomorphologies, with multiple observation angles, detection ranges, and frequency bands. We establish a rigorous hierarchical annotation protocol that decouples labeling of intrinsic physical properties, deployment characteristics, and scattering phenomena - covering material, geometry, internal structure, burial state, shadow integrity, specular highlights, edge diffraction, and resonance effects. This enables fine-grained target characterization. We also construct a multi-task benchmark for object detection, fine-grained classification, and attribute prediction, evaluating mainstream deep learning models under cross-domain, cross-scene, cross-frequency, and cross-view generalization. SCTD 3.0 is expected to provide a critical data cornerstone for robust underwater target perception in open-water environments. SCTD 3.0 is available at this https URL.
223. 【2608.08075】Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
链接:https://arxiv.org/abs/2608.08075
作者:Sankalp Nagaonkar,Rohit Garg,Ankit Raj,Ashish Choithani,Ashutosh Trivedi
类目:Information Retrieval (cs.IR); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:return ranked files, video-retrieval systems assume, assume a bounded, return ranked, bounded corpus
备注: 33 pages, 5 figures, 17 tables. Technical report. Benchmark configurations and reproduction instructions: [this https URL](https://github.com/video-db/search-over-the-visual-world)
点击查看摘要
Abstract:Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; models interpret them at different temporal granularities; context must be selected without replaying the complete visual record; and results must stay connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer-defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability-declared indexes, distinguishing memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this model in production, exposed through a typed search surface spanning planned retrieval, stateful investigation, direct access, and grounded synthesis. We contrast this model-agnostic infrastructure, where segmentation, sampling, model choice, embeddings, and ranking are system decisions and live streams are first-class sources, with video-native foundation models offered as fixed APIs. In a semantic-retrieval comparison against a commercial video-native engine spanning 9,800+ queries over four public datasets, a pipeline of general-purpose components achieves higher macro-averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). Retrieval quality over the visual world is today governed more by system design than by video-specific pretraining, and visual-memory infrastructure can deliver it while keeping playable, source-grounded evidence first-class.
224. 【2608.08068】NeuroGuard: Neural Gradient Update Aware of Representation Damage
链接:https://arxiv.org/abs/2608.08068
作者:Taigo Sakai,Kazuhito Hotta
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Long-tailed class-incremental learning, Long-tailed class-incremental, imbalanced streams, streams while retaining, Long-tailed
备注: Accepted at HCV workshop on ECCV2026
点击查看摘要
Abstract:Long-tailed class-incremental learning (LT-CIL) must learn new classes from imbalanced streams while retaining old classes. Existing methods mainly change replay, classifiers, or losses. We study a different factor, namely how strongly the feature representation should be updated at each task boundary. We propose NeuroGuard, an update-control method added to DGR, a replay-based LT-CIL baseline, without adding learnable parameters. NeuroGuard preserves DGR's replay memory, classifier, and set of loss terms. Adaptive Gradient Scaling (AGS) converts teacher uncertainty into one task-wise gradient scale. Confidence-Ranked Knowledge Distillation Reweighting (CRK) gives larger knowledge-distillation weights to replay samples that the teacher predicts less decisively. Fragility-Blended Entropy Gate (FBE) adds old-memory leakage to the scale decision. Across five LT-CIL settings, NeuroGuard improves over DGR in every setting. In the four main benchmark comparisons, it achieves the best task-agnostic accuracy among the compared methods. The gains extend to both old- and new-class accuracy, while medium-frequency accuracy improves consistently across all five settings. Controlled comparisons show that the gain does not come from generic gradient suppression: AGS outperforms a matched fixed-scale control in all five settings, demonstrating that boundary-specific scaling is more effective than applying the same average scale throughout learning.
225. 【2608.08066】EvBS: Event-guided Blur Synthesis for Domain-adaptive Motion Deblurring
链接:https://arxiv.org/abs/2608.08066
作者:Junsik Jung,Seokryun Choi,Yoonki Cho,Woo Jae Kim,Andrew Jeong,Sung-Eui Yoon
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:achieved remarkable progress, real-world scenarios due, deep learning, achieved remarkable, remarkable progress
备注: Accepted to ACM Multimedia 2026 (ACM MM 2026)
点击查看摘要
Abstract:Motion deblurring has achieved remarkable progress with deep learning, yet pre-trained deblurring models often suffer from performance degradation in real-world scenarios due to the domain shift between training and testing distributions. To remedy this, we propose EvBS, an event-guided blur synthesis framework that generates diverse training pairs for calibrating pre-trained models to the target domain. While existing methods are constrained by the inherent entanglement between motion and visual content, our method leverages the high temporal resolution of event cameras to effectively decouple them. This enables us to utilize not only the intrinsic motion that is inherent to the given content but also extrinsic motion transferred from different sources within the target domain, thereby facilitating effective adaptation via fine-tuning. Specifically, EvBS comprises two complementary strategies: Intrinsic-Blur Synthesis, which blurs sharp contents with their own motion patterns, and Extrinsic-Blur Synthesis, which transfers motion from blurry patches to distinct sharp content. This approach generates a diverse set of training pairs that break the inherent constraints of naturally coupled motion and content, resulting in enhanced domain-adaptive deblurring performance. Extensive experiments on multiple benchmarks demonstrate that EvBS effectively enhances the robustness of existing deblurring models on unseen testing datasets.
226. 【2608.08060】ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models
链接:https://arxiv.org/abs/2608.08060
作者:Sajjad Ghiasvand,Yifan Yang,Mahnoosh Alizadeh,Ramtin Pedarsani
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:memory-constrained edge devices, typically requires backpropagation, CLIP typically requires, proprietary model deployments, Fine-tuning vision-language models
备注:
点击查看摘要
Abstract:Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as is common for memory-constrained edge devices and proprietary model deployments. Prior BP-free, zeroth-order prompt-tuning methods avoid this requirement but often tune prompts in a single modality or optimize over a search space large enough that convergence requires thousands of forward passes, which is impractical under realistic query budgets. We propose ZOMP (Zeroth-Order Multimodal Prompt tuning), a query-efficient, fully forward-only method that tunes deep prompts in both the vision and text branches of a frozen CLIP model using simultaneous perturbation stochastic approximation. ZOMP combines three ingredients: a cross-modal low-rank reparameterization that ties the two branches through a shared factor and keeps the effective search dimensionality small, a gradient-correction momentum term that stabilizes the noisy zeroth-order estimate, and a budget-indexed rank schedule that unlocks capacity as the query budget is spent. Across 13 vision-language benchmarks under a matched 5,000-query budget, ZOMP consistently outperforms prior BP-free prompt-tuning methods in both few-shot accuracy and query efficiency, and it generalizes better across base-to-new, cross-dataset transfer, and out-of-distribution settings. Our results show that jointly exploiting multimodality and low-rank structure is an effective route to practical, query-efficient BP-free prompt tuning.
227. 【2608.08053】PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets
链接:https://arxiv.org/abs/2608.08053
作者:Jie Huang,Xiaohe Li,Jiahao Li,Fangli Mou,Chen Qian,Yuqiang Fang,Junhao Fan,Kaixin Zhang,Zide Fan
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:central to robotics, robotics and embodied, Simulation-ready, geometry, intermediate physical states
备注:
点击查看摘要
Abstract:Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model that emits a serialized asset for a decoder to turn into geometry and physical fields, leaving the image-to-3D reasoning implicit. We argue the limiting factor is this output-centric view: part placement and local shape are entangled in one global-coordinate token stream, and the intermediate physical states are never exposed for supervision, conditioning, or verification. PhysX-CoT instead casts single-image asset generation as an explicit structured physical reasoning process, an ordered and machine-parseable trajectory of part-level states covering decomposition, 2D and 3D grounding, relations, coarse geometry, and surface cues that we separately supervise, use to condition geometry, and treat as reward targets. Geometry is factorized so that 3D boxes carry placement and local codes carry shape, and CoT-aligned GRPO optimizes parse validity, grounding, geometry, placement, and physical consistency. Under a unified protocol that retrains all learned baselines on the same backbone, data, and frozen decoder, PhysX-CoT outperforms the closest full-task baseline across geometry, scale, and physical-attribute metrics. Oracle, token-matched, and state-order controls show the explicit states are functional rather than cosmetic, and in Unreal Engine~5 the generated assets parse, collide, and articulate at high validity.
228. 【2608.08021】Evidence-RL: Towards Evidence-intensive Visual Reasoning
链接:https://arxiv.org/abs/2608.08021
作者:Haojie Huang,Xinlei Yu,Chengming Xu,Zhangquan Chen,Cheng Yang,Qingdong He,Yu Yang,Jiangning Zhang,Xiaobin Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:irrelevant visual context, Vision-Language Models, visual context, concrete image evidence, irrelevant visual
备注: 22 pages, 10 figures
点击查看摘要
Abstract:Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.
229. 【2608.08016】EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking
链接:https://arxiv.org/abs/2608.08016
作者:Jan Kulik,Bjarni Dagur Thor Karason,Yung-Hsu Yang,Boyang Sun,Marc Pollefeys,Xi Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:partial occlusions make, occlusions make building, make building structured, structured representations challenging, building structured representations
备注:
点击查看摘要
Abstract:Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static scenes, limiting their ability to capture complex dynamics. We introduce EgoTrack3D, a modular framework that reconstructs and maintains a dynamic 3D scene representation directly from egocentric RGB video. The framework lifts 2D segmentation masks into a global 3D coordinate frame, using a point-based motion scoring mechanism alongside a voxel-based merging heuristic to associate object tracks. EgoTrack3D maintains accurate representations over time, achieving an 11% improvement in percentage of correct locations (PCL) relative to the strongest baseline on the Aria Digital Twin (ADT) dataset, while addressing the more general setting of persistent 3D tracking for both static and dynamic objects. Furthermore, to demonstrate the system's robustness under degraded conditions that simulate real-world deployment constraints, we replace dense depth maps with sparse 3D bounding box estimation and integrate interaction-guided dynamic association, enabling EgoTrack3D to maintain accurate spatial representations despite noisy observations.
230. 【2608.08009】Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation
链接:https://arxiv.org/abs/2608.08009
作者:Yichun Yeh,Yiheng Li,Xiaobo Hu,Zhen Lei,Yang Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Grounding Multi-Modal Media, Detecting and Grounding, Fake news increasingly, Multi-Modal Media Manipulation, cross-modal image-text forgeries
备注: accepted by ACM MM 2026
点击查看摘要
Abstract:Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4). Existing methods produce black-box detection results without any decision rationale, limiting their reliability in forensic practice. Multi-modal Large Language Models (MLLMs) offer a natural path toward explainability, but applying them to DGM4 raises two difficulties. First, models tend to generate explanations disconnected from predicted evidence locations, producing unverified attribution. Second, enforcing evidence-conclusion consistency requires active optimization, yet uniform training signals fail to distinguish localization tokens from classification tokens, making multi-head joint training unreliable. We propose a multi-modal manipulation detector based on an Evidence-Grounded Forensic Reasoning (EFR) framework. EFR introduces an Anchor-and-Verify reasoning chain that enforces modality-isolated perception before cross-modal comparison, with conclusion coordinates as explicit anchors to which downstream evidence must spatially correspond. A verifiable reward system then enforces evidence-conclusion consistency during training, while a Modality-Decoupled Advantage (MDA) routing mechanism mitigats credit misassignment across prediction tasks. Experiments show that EFR achieves state-of-the-art performance while producing structured forensic reasoning records that explicitly bind explanations to evidence.
231. 【2608.07999】PE-Mamba: Bidirectional Selective Layer Aggregation for AI-Generated Image Detection
链接:https://arxiv.org/abs/2608.07999
作者:Kutub Uddin,Nusrat Tasnim,Khalid Malik
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:increasingly challenging due, AI-generated image, authentic content, increasingly challenging, challenging due
备注:
点击查看摘要
Abstract:AI-generated image (AIGI) detection has become increasingly challenging due to the rapid advancement of generative models and the diminishing gap between synthetic and authentic content. Existing vision transformer-based detectors commonly rely on weighted-sum strategies to aggregate intermediate representations across transformer layers, often overlooking the inherently ordered semantic progression of hierarchical features from shallow texture cues to deep semantic representations. In this work, we propose \textbf{PE-Mamba}, a novel framework built upon a pre-trained PE-Core vision transformer with lightweight LoRA adaptation that introduces three complementary components for cross-layer feature aggregation and fusion. First, a bidirectional selective aggregator (BSA) processes layer-wise classification tokens through forward and backward selective scans, where the forward scan progressively accumulates shallow-to-deep forensic evidence, and the backward scan performs deep-to-shallow contextual refinement to reinterpret low-level cues in light of high-level semantic context. Second, a softmax-weighted aggregator (SWA) computes a learned global summary of all layer tokens as a complementary aggregation path. Third, a sigmoid-gated blend (SGA) adaptively fuses the BSA and SWA outputs via a learnable scalar gate, allowing the model to dynamically balance directional sequential evidence and global layer-wise aggregation. Extensive experiments on UniversalFakeDetect (96.6\% mACC, 99.5\% mAP) and AIGCDetect (95.3\% mACC, 98.1\% mAP) demonstrate that \methodname{} outperforms 18 detectors with superior generalization across diverse generative models, while training only 1.3\% of total parameters (0.13\% for LoRA alone).
232. 【2608.07993】MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval
链接:https://arxiv.org/abs/2608.07993
作者:Fulong Liu,Liang Xu,Chengqun Yang,Yuhao Zhang,Yichao Yan,Xiaokang Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Human motion-text retrieval, assessing cross-modal alignment, Human motion-text, assessing cross-modal, Human
备注:
点击查看摘要
Abstract:Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor motions, imbalanced motion distributions, and oversimplified, repetitive texts, which hinder the reliable measurement of cross-domain and cross-granularity alignment. We thus introduce MRBench, a comprehensive motion-text retrieval benchmark featuring heterogeneous motions, broad and balanced category coverage, and reliable, discriminative, multi-granular descriptions. MRBench is constructed through a meticulously designed multi-stage data curation pipeline, which filters and balances candidates, verifies unambiguous semantic alignment, and generates motion-grounded descriptions at multiple granularities. The resulting benchmark contains 3,390 motions drawn from motion capture, in-the-wild videos, synthetic videos, and motion generative models, covering 118 fine-grained categories. Each motion is paired with concise, standard, and fine-grained descriptions, yielding 10,170 captions. Extensive evaluations of representative retrieval baselines on MRBench reveal a substantial cross-dataset generalization gap and pronounced sensitivity to query granularity. We propose a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model. LLM-based concise and fine-grained captions provide pseudo-supervision for extra-branch granularity-specific motion extractors and text adapters. For inference, granularity-aware score fusion integrates global and adapted similarities while strictly maintaining score comparability across all description levels. The resulting model improves mixed-granularity retrieval without compromising standard-caption performance. We believe that our MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.
233. 【2608.07987】Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence
链接:https://arxiv.org/abs/2608.07987
作者:Ling Lin,Yang Bai,Congcong Zhu,Jiangming Shi,Meng Wang,Yang Long,Jingrun Chen,Ling Shao,Huazhu Fu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multimodal large language, demonstrated significant potential, Multimodal large, large language models, complex spatial scene
备注:
点击查看摘要
Abstract:Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, their open-ended reasoning process is prone to decision errors and error accumulation, leading to instability in answer quality. To address this, we propose an advantage-guided gating framework that dynamically intervenes in and corrects deviations during the reasoning process. Specifically, we model step-by-step reasoning as a finite-horizon decision process and introduce Monte Carlo value evaluation on the reasoning tree to provide intermediate supervision signals. The framework includes Step-Advantage Gate and Trajectory-Advantage Gate, which dynamically select high-value reasoning steps and high-quality complete reasoning trajectories, respectively. During training, we perform supervised learning for the gates using reasoning trees generated via multi-branch sampling, and combine shared-parameter initialization with task-specific heads to achieve cross-task robustness and diversity. During inference, the model greedily selects high-value prefix reasoning steps while choosing the optimal reasoning head based on the problem type, thereby significantly improving the accuracy of the final answer. Furthermore, we constructed the Reasoning-Tree-160k dataset and performed two-stage learning on it. Extensive experiments demonstrate that this advantage-guided gating framework effectively enhances the performance of benchmark MLLMs in visual-based spatial understanding and reasoning tasks. The code is open to the public for research: this https URL.
234. 【2608.07984】AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining
链接:https://arxiv.org/abs/2608.07984
作者:Vasileios Tzouras,Paraskevas Pegios,Lazaros Nalpantidis
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Field-based agricultural computer, large pretrained models, precision agriculture, Field-based agricultural, important for precision
备注: Accepted at the Computer Vision in Plant Phenotyping and Agriculture (CVPPA) Workshop at the European Conference on Computer Vision (ECCV) 2026
点击查看摘要
Abstract:Field-based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation of large pretrained models. We introduce AgriField-40K, a field-centric dataset curated from 17 public resources and covering diverse crops, weeds, and field conditions. Building on this, we present AgriMAE, a parameter-efficient continual pretraining baseline that adapts a masked autoencoder pretrained on natural images by training only lightweight adapters. We further explore semantic feature reconstruction as an alternative pretraining objective and evaluate transfer across multiple tasks. AgriMAE consistently improves downstream performance and can match or even outperform full fine-tuning while using up to $9\times$ fewer trainable parameters, showing that AgriField-40K is a practical resource for continual pretraining in agricultural vision. Project page: this https URL
235. 【2608.07982】AdaDINO: Pair-Aware In-Backbone Adaptation of Frozen DINO for Efficient Remote Sensing Change Detection
链接:https://arxiv.org/abs/2608.07982
作者:Xu Zhang,Xinqing Li,Jianpeng Xie,Zeshuai Zhu,Xin He,Yun Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision foundation models, Vision foundation, detection requires reasoning, change detection requires, foundation models
备注:
点击查看摘要
Abstract:Vision foundation models (VFMs) such as DINO are pretrained for single-image representation, whereas remote sensing change detection requires reasoning over a bi-temporal pair. Existing VFM-based methods usually encode the two images independently and compare them only afterward, leaving the VFM backbone unaware of cross-temporal relations. To bridge this mismatch, we present AdaDINO, a pair-aware in-backbone adaptation framework that equips a frozen DINO encoder with bi-temporal interaction for efficient change detection. Its core component, Change-aware Gated Local Adaptation (CGLA), couples the two streams after selected frozen blocks and injects a shared temporal residual into them with opposite signs, enhancing genuine change responses while preserving the pair midpoint. Batch-Shared Chunk Selection (BSCS) further reduces feed-forward network (FFN) computation by retaining a batch-shared subset of channel chunks that can be executed as a compact dense FFN. A CGLA-Prior-Guided Refinement (CPGR) decoder reuses encoder-side change responses for coarse-to-fine prediction. Experiments on four remote sensing change detection benchmarks show that AdaDINO achieves competitive or superior performance against VFM-based baselines, with the largest gain on the category-agnostic SYSU-CD dataset. With 62.5% of the FFN hidden width removed, AdaDINO still achieves an F1 score of 85.29% on SYSU-CD while delivering a 1.41$\times$ throughput speedup. The code will be released.
236. 【2608.07981】Distilling Physical Priors into Streaming World Models
链接:https://arxiv.org/abs/2608.07981
作者:Liangliang Zhao,Junying Wang,Danni Yang,Yifan Chang,Bin Fu,Yu Qiao,Bowen Zhou,Yihao Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:predict future visual, models predict future, future visual states, maintaining physically coherent, long horizons
备注: 9 pages, 7 figures. Project page: [this https URL](https://lyongo.github.io/PhyS/)
点击查看摘要
Abstract:Streaming world models predict future visual states online while maintaining physically coherent dynamics over long horizons. However, their rollouts often violate basic physical constraints. A common approach distills pretrained bidirectional DiTs into few-step causal generators. However, this paradigm suffers from two fundamental limitations: generic bidirectional teachers acquire limited physical priors from visually oriented pretraining, and the limited priors suffer further loss during bidirectional-to-causal distillation. We present PhyS, a three-stage framework for distilling physical priors into streaming world models. To acquire physical priors from real-world interactions, we construct PhyS-120K, a dataset of 120K real-world physical-interaction videos spanning rigid-body dynamics, soft-body deformation, fluid phenomena, and phase transitions. Each video is annotated with structured descriptions of object properties and causal state transitions. Physics-aware supervised fine-tuning injects the physical priors into a bidirectional 14B DiT teacher, which we then distill into a lightweight 1.3B causal DiT for few-step autoregressive streaming generation. Finally, we use online reinforcement learning to incentivize the distilled model to generate physically plausible rollouts and further propose Temporal Credit Routing (TCR) to address temporal credit assignment. TCR evaluates physical consistency over overlapping temporal windows and routes the resulting group-relative advantages to temporally aligned denoising actions. On PhysicsIQ, PhyS improves the Wan2.1-14B teacher by 18.2\% and the Self Forcing, Rolling Forcing, and Causal Forcing by 23.7\%, 14.8\%, and 31.4\%, respectively. Results also improve the physics-aware video benchmarks VideoPhy, VideoPhy2, and PhyGenBench. The dataset, code, and more sample videos are available on our Project Page.
237. 【2608.07958】LIBAD: A Multimodal Anomaly Detection Benchmark for Li-Ion Battery Electrode Manufacturing
链接:https://arxiv.org/abs/2608.07958
作者:Wenbo Sui,Daniel Lichau,Harold Phelippeau,Zhao Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:strongly correlated RGB, sensing modalities underexplored, weakly correlated sensing, correlated sensing modalities, correlated RGB
备注:
点击查看摘要
Abstract:Multimodal industrial anomaly detection has largely focused on discrete products using strongly correlated RGB and 3D observations, leaving continuous process manufacturing and weakly correlated sensing modalities underexplored. We introduce LIBAD, the first multimodal anomaly detection benchmark for Li-ion battery electrode manufacturing. Collected from real roll-to-roll production lines, LIBAD provides aligned double-sided visible-light imaging, high-resolution X-ray radiography, and inline-compatible low-resolution X-ray radiography. Electrode patches in LIBAD exhibit highly homogeneous material appearance, while defect evidence can be strong in one modality but weak or absent in another, resulting in pronounced cross-modal anomaly inconsistency. Benchmarks of representative methods under the inline-compatible visible-light and low-resolution X-ray setting exhibit limited transferability and consistently high false-positive rates. We therefore propose DA-Core, a memory-based method that jointly considers feature-space coverage and local density of normal features during coreset selection, allowing compact memory banks to better preserve fine-grained normal variations. With a coreset ratio of 0.05, DA-Core reduces FPR95 from 60.4% to 54.3% compared with standard farthest point sampling. At this ratio, DA-Core also outperforms the best standard coreset result (obtained at 0.20) while reducing inference time by 43.9%. These results suggest that both the data distribution of normal features and the modality relationship itself require explicit consideration when designing anomaly detection methods for process manufacturing.
238. 【2608.07948】SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
链接:https://arxiv.org/abs/2608.07948
作者:Zhennan Chen,Tianxing Shi,Pengcheng Xu,Kepan Nan,Qian Wang,Zili Yi,Jian Yang,Ying Tai
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:gained widespread popularity, widespread popularity due, next-scale prediction paradigm, gained widespread, widespread popularity
备注: Accepted by ECCV 2026
点击查看摘要
Abstract:VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propagation and accumulation in VAR. To this end, we propose SynVAR, the first training-free enhancement framework specifically tailored for the VAR paradigm, which introduces a spatial-semantic collaborative control strategy to effectively suppress propagation error and improve generation quality. SynVAR comprises three key components: (1) Global guidance to ensure reasonable spatial structure in the early stages, (2) Receptive field constraints to mitigate early-stage semantic confusion, (3) High-frequency compensation to recover fine-grained details. Extensive quantitative and qualitative experiments demonstrate the significant improvements in the ability of SynVAR to enhance the VAR's capability for complex scene modeling.
239. 【2608.07941】LAD-COD: Language-Aligned Dense Perception for Camouflaged Object Detection
链接:https://arxiv.org/abs/2608.07941
作者:Shangye Song,Tianzhi Zhu,Syed Ariff Syed Hesham,Xin He,Yun Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Camouflaged object detection, reduces foreground-background discriminability, Camouflaged object, exhibit high visual, high visual similarity
备注:
点击查看摘要
Abstract:Camouflaged object detection (COD) aims to segment objects that exhibit high visual similarity to their surroundings, which reduces foreground-background discriminability and weakens boundary evidence across appearance, texture, and structure. Such limitations motivate the use of instruction-conditioned semantics as top-down guidance for identifying which weak visual cues are relevant to the target. Recent segmentation systems built on large multimodal models (LMMs) demonstrate this possibility through instruction-conditioned target embeddings that guide mask decoding. However, in this language-to-mask paradigm, the generated target embedding conditions mainly the mask decoder, leaving the dense visual features that must preserve low-contrast boundaries and fine local structure without explicit guidance. We propose Language-Aligned Dense perception for COD (LAD-COD), a framework that aligns top-down semantic target guidance with bottom-up hierarchical visual features. Instead of fully adapting a large generic image encoder, LAD-COD learns a trainable hierarchical visual branch that captures camouflage-sensitive texture, boundary, and contextual information. To align these features with the target embedding, LAD-COD applies Language-Aligned Dual Visual Fusion (LADVF), which extends the embedding beyond sparse prompting to query patch-level language-aligned features and to gate their residual integration with the hierarchical features. This design allows semantic information to guide localization while preserving the fine structural details needed for camouflage segmentation. Experiments on CAMO, COD10K, and NC4K show that LAD-COD obtains the best reported value in all 12 dataset-metric comparisons.
240. 【2608.07937】FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence
链接:https://arxiv.org/abs/2608.07937
作者:Amir Sabbaghziarani,Hanting Ye,Maria Gorlatova,Yi Ding
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:object-centric multi-view image, multi-view image collections, feed-forward framework, jointly trained geometry, object-centric multi-view
备注: Accepted to the British Machine Vision Conference (BMVC) 2026. Code: [this https URL](https://github.com/amir-sbg/FlexSplat)
点击查看摘要
Abstract:We present FlexSplat, a feed-forward framework for novel view synthesis (NVS) from uncalibrated, object-centric multi-view image collections. A recent line of query-based methods reconstructs a compact set of 3D Gaussians by treating them as transformer queries that are refined with multi-view deformable attention; these methods, however, assume that camera poses are given. FlexSplat removes this assumption: a geometry transformer is trained jointly with the Gaussian decoder to predict per-image camera parameters and depth, which in turn ground a depth-guided Gaussian parameterization and a multi-view deformable cross-attention that aggregates evidence across all input views into a single, view-consistent set of primitives. An uncertainty-weighted depth-consistency objective lets the jointly trained geometry adapt to the reconstruction task, while the cross-view consensus formed during decoding absorbs the residual error of the estimated cameras and depth. The representation uses a compact Gaussian budget that is decoupled from the input resolution - unlike pixel-aligned methods, the primitive count does not grow with the image grid - and is not dictated by the number of views. On ShapeNet-SRN and Google Scanned Objects (GSO), FlexSplat matches or approaches posed state-of-the-art reconstructors while requiring neither camera poses nor ground-truth depth, and matches the best perceptual (LPIPS) quality among the compared methods on GSO. Our results indicate that a jointly trained geometry front-end is sufficient to bring calibration-free operation to query-based Gaussian reconstruction while staying within 0.7 dB PSNR of posed methods and matching their perceptual quality.
241. 【2608.07932】SportsGrounder: Proposal-Aided Interleaved Grounding for Dense Sports Video Reasoning
链接:https://arxiv.org/abs/2608.07932
作者:Yizhi Li,Jiawei Jiang,Guanhong Wang,Yingcai Wu,Gaoang Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Sports video analysis, Dense sports video, sports video reasoning, Sports video, broadcasting enhancement
备注: ACMMM 2026
点击查看摘要
Abstract:Sports video analysis is crucial for athletic analytics and broadcasting enhancement. Dense sports video reasoning, however, demands a fine-grained understanding of numerous small-scale, highly interactive, and visually homogeneous entities (e.g., players sharing identical uniforms, the ball) across long temporal contexts. Current Large Multimodal Models (LMMs) inherently struggle with such dense visual complexities. Due to the lack of fine-grained visual details, these models often over-rely on textual priors to guess answers, especially when distinguishing visually similar actions and players. To address this, we propose \textbf{SportsGrounder}, a framework that leverages an open-vocabulary visual expert to aid interleaved grounding specifically for dense sports video reasoning. To achieve precise spatial localization, we extract domain-guided object proposals and introduce an Interleaved Grounding Fusion (IGF) mechanism. The IGF frame-by-frame integrates explicit bounding box coordinates and implicit visual semantics with global grid features. This design preserves strict temporal alignment and prevents sequence length explosion. Furthermore, we design an Action-Aware Supervision (AAS) module that directly regularizes the model's hidden states, forcing the network to learn accurate motion representations rather than relying on language bias. Optimized with Mixed Preference Optimization (MPO) to better distinguish deceptive distractors, our extensive experiments on newly curated dense sports VQA datasets (derived from SoccerNet and FineSports) demonstrate that SportsGrounder significantly improves fine-grained reasoning and achieves state-of-the-art accuracy.
242. 【2608.07923】SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange
链接:https://arxiv.org/abs/2608.07923
作者:Jaemo Jeong,Junho Yoon,Hyunju Kim,Dongman Lee
类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)
关键词:Audio-visual event perception, Audio-visual event, AVEP, Audio-visual, occur
备注: 17 pages, 6 figures
点击查看摘要
Abstract:Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual features with text-encoded event names. However, related labels share evidence. An incorrect label can then score at least as high as a correct one. We call this a false co-activation (FCA). No scalar cutoff can reject the incorrect label while keeping every correct one. Class-specific thresholds may prevent that label from becoming a final prediction, but the FCA remains in the underlying score vector. We introduce SCoPE, a training-free framework in which all queried labels compete for shared evidence and each modality guides event selection in the other. We derive an exact condition for when this competition removes an FCA in a two-label fit. With identical frozen CLIP+CLAP backbones on LLP, SCoPE improves Type@seg by 7.45 points and Event@seg by 5.04 points compared with the reported AV$^2$A values. The same fixed configuration transfers unchanged to OV-AVEBench and VGGSound-AVEL100k.
243. 【2608.07920】Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification
链接:https://arxiv.org/abs/2608.07920
作者:Yang Shu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:LLM judge panels, Multimodal LLM judge, Multimodal LLM, LLM judge, quoted peer judgment
备注:
点击查看摘要
Abstract:Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-level attack surface in vision-language model (VLM) panels. Quoting independent visual judgments creates large anchoring gaps (19--26 percentage points) under both self and peer framing. A matched-content, label-only control changes the broken rate by only $-0.17$pp (95\% CI $[-0.68,0.35]$), showing that the self/peer label itself does not explain the effect. Under our tested construction, deliberately generated, concise wrong quotes overturn originally-correct verdicts 1.5--2.7$\times$ more often than naturally occurring wrong peer statements, with bootstrap 95\% CIs excluding parity across two datasets and seven VLM judges. Because the two statement populations differ in selection and form, this ratio measures differential damage under the tested attack rather than a provenance-only causal effect. We then introduce panel-consensus verification, which cross-checks a quote against independently collected blind votes. It blocks 84.9\% of fabricated attacks, cuts their net harm by 97.5\%, and preserves the positive but statistically inconclusive point estimate for genuine peer information under leave-one-out re-verification. These results identify a low-cost attack surface and a concrete defense for safer multimodal collaborative evaluation.
244. 【2608.07916】SegDem: Segmentation helps Demosaicing
链接:https://arxiv.org/abs/2608.07916
作者:Ping Chen,Xiangming Wang,Yongyong Chen,Jiezhang Cao,Kai Zhang,Jingyong Su,Jie Liu,Haijin Zeng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:color filter array, incomplete color measurements, color measurements produced, Image demosaicing reconstructs, incomplete color
备注: 20 pagess
点击查看摘要
Abstract:Image demosaicing reconstructs a full-color image from incomplete color measurements produced by a sensor covered with a color filter array (CFA). Most existing methods formulate demosaicing as pixel-level reconstruction and mainly rely on local textures, cross-channel correlations, and low-level image statistics. Our core insight is that reconstruction and visual understanding can be viewed as complementary views of shared scene structure: both are grounded in the same underlying physical world, and therefore the structural and physical information inferred from an image should remain consistent across the two tasks. We instantiate this idea with instance segmentation and propose \emph{SegDem}, a cross-task decoder representation transfer framework for demosaicing. SegDem first learns region- and boundary-aware representations through instance-aware structural pretraining and then transfers the decoder to RAW-conditioned reconstruction. Segmentation- and demosaicing-conditioned features are further anchored to a shared frozen DINOv2 representation space to preserve structural organization across tasks. We instantiate SegDem with convolutional, Transformer-based, and state-space backbones for unified Single- and Quad-Bayer demosaicing. Extensive experiments on synthetic, external, and challenging datasets demonstrate consistent improvements across different architectures and CFA layouts.
245. 【2608.07904】DeCo: Zero-Shot Industrial Anomaly Generation through Decoupling and Recoupling
链接:https://arxiv.org/abs/2608.07904
作者:Shilei Zeng,Xurui Li,Yaohan Tang,Yu Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Industrial anomaly inspection, Industrial anomaly, Zero-shot industrial anomaly, industrial anomaly generation, real anomalous
备注:
点击查看摘要
Abstract:Industrial anomaly inspection is severely hindered by the scarcity of real anomalous data. Zero-shot industrial anomaly generation addresses this by generating anomalies on specific products without requiring any of their real anomalous images. However, existing methods suffer from two critical limitations, i.e., inaccurate anomaly information acquisition and uncontrolled anomaly-product fusion. To overcome these challenges, we propose DeCo, which decouples the anomaly structure from its source product, and explicitly recouples it with the normal textures of the target product. During anomaly information acquisition, Dual-Routing Flow (DR-Flow) binds the texture-invariant anomaly structure to an abnormal token, while a parallel constraint, Product-Invariant Flow (PI-Flow), prevents the abnormal token from binding the source product. During anomaly-product fusion, we propose a hybrid injection to recouple the acquired anomaly structure with the target product, and Product Compatibility Correction (PCC) to compensate for the incompatibility between the acquired anomaly structure and the product. Extensive experiments demonstrate that DeCo establishes a new state-of-the-art. Training downstream detection models on our generated data yields massive pixel AP improvements of 5.1% on MVTec AD and 8.2% on VisA. Code is available at this https URL.
246. 【2608.07886】Vision-Language Grounding as Bidirectional Concept Correspondence
链接:https://arxiv.org/abs/2608.07886
作者:Jieyu Zhang,Ziqi Gao,Luke Zettlemoyer,Ranjay Krishna
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:unidirectional localization problem, existing formulations reduce, grounding connects language, visual content, connects language
备注:
点击查看摘要
Abstract:Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region. This setup assumes that the relevant linguistic unit is already known, overlooking a more basic challenge in grounded communication: determining which parts of the text are visually referential and how they correspond to entities in the image. We formulate grounding as $\textit{bidirectional concept correspondence}$ over an image-text pair. Given an image and its paired text, the goal is to recover all correspondences between visually referential text spans and instance-level image segments, without assuming that the relevant text spans are provided. This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem. To address this task, we introduce $\textbf{ConCor-1}$, a grounding model built on top of a pretrained vision-language model. It uses learnable $\textit{bridge tokens}$ to represent candidate image-text correspondences and predicts, for each token, a text mask, an image mask, and a correspondence presence score. To train and evaluate this task, we convert diverse grounding and segmentation datasets into a unified correspondence format. Experiments show that $\textbf{ConCor-1}$ consistently outperforms baselines, improving correspondence F1 by 48% on the long-caption dataset and by 29% on zero-shot LVIS, where the large category list serves as the text input.
247. 【2608.07864】UniScale: Arbitrary-Scale Industrial Anomaly Generation
链接:https://arxiv.org/abs/2608.07864
作者:Shilei Zeng,Linxin Guan,Xurui Li,Yaohan Tang,Yu Zhou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:major challenge due, real-world anomaly samples, anomaly inspection faces, Industrial anomaly inspection, inspection faces
备注:
点击查看摘要
Abstract:Industrial anomaly inspection faces a major challenge due to the lack of real-world anomaly samples. While generative models are used to create anomaly data, existing methods still struggle when handling small-scale anomalies. This failure occurs because extreme downsampling in diffusion models causes the information of small anomalies to be lost in the latent space. To address this, we introduce UniScale, a unified training and inference framework for high-fidelity industrial anomaly generation across arbitrary scales. During training, we introduce an Error-Suppressed Multi-Scale Training (EMT) strategy, which enables the model to learn the rich location-aware textures of anomalies, while suppressing upsampling-induced interpolation errors in texture acquisition, ensuring the model is capable of learning small-scale anomalies, while remaining effective for regular scale anomalies. For inference, we propose Generation-then-Fusion Denoising. It decouples anomaly generation from background integration, preventing small anomalies from being overwhelmed. Extensive experiments demonstrate that our method outperforms state-of-the-art competitors in both anomaly generation quality and downstream detection performance. It achieves a relative IS(a) improvement of 45.86% (from 1.81 to 2.64) on VisA and 37.70% (from 1.22 to 1.68) on MVTec AD 2, while also improving the downstream pixel-level IoU by 4.22% on VisA and AUROC by 6.55% on MVTec AD 2. Code is available at this https URL.
248. 【2608.07863】LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering
链接:https://arxiv.org/abs/2608.07863
作者:Qian Yao,Jun-Jie Huang,Yongjun Wang,Luming Yang
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:Driven by advances, AI-generated image, AI-generated image detection, AI-generated, AI-generated images
备注:
点击查看摘要
Abstract:Driven by advances in diffusion models and autoregressive models, the fidelity and resolution of AI-generated images now rival those of real images. However, existing AI-generated image detection methods often downsample the images, inevitably overlooking critical low-level texture details in high-resolution AI-generated images, therefore limiting their detection performance. In addition, the ceaseless emergence of unknown generative models makes large-scale pre-training datasets inaccessible. To address these challenges, we propose a novel high-resolution AI-generated image detector, termed LHSDet. Specifically, we formulate the AI-generated image detection task as a Visual Question Answering problem, leveraging a fine-tuned vision-language framework to fully exploit the complementary information between visual and textual modalities. Recognizing that the default visual encoder of existing vision-language models is not tailored for AI-generated image detection, we redesign a visual encoder to better capture both the low-level and high-level artifacts inherent in AI-generated images. Furthermore, we incorporate a semantic-level textual branch to enable multi-modal feature fusion and detection. Consequently, LHSDet employs a triple-branch architecture to extract complementary multi-modal features: a low-level visual branch that aggregates non-overlapping patches for local texture cues, a high-level visual branch based on SigLIP2 for global perception feature extraction, and a semantic-level textual branch that generates captions using BLIP-2. Extensive experimental results demonstrate that LHSDet achieves high detection accuracy and robust performance across diverse generative models, including both diffusion and autoregressive models.
249. 【2608.07861】How Much Does It Cost to Answer My Question? Benchmarking Cloud VLM-based VQA Systems
链接:https://arxiv.org/abs/2608.07861
作者:Henri Vanhuynegem,Weitao Xu,Yiran Shen,Guohao Lan
类目:Computer Vision and Pattern Recognition (cs.CV); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR); Multimedia (cs.MM)
关键词:visual question answering, answer users' questions, question answering, Vision-language models, enabling smartphones
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) are becoming a practical backend for mobile visual question answering (VQA) systems, enabling smartphones and smart glasses to answer users' questions about the physical world. Since modern VLMs remain difficult to run on mobile and edge devices, VQA systems increasingly offload inference to cloud-based VLMs. This gives mobile devices access to stronger computation, but it also makes visual input preparation a key system variable: how the image is prepared before offloading affects not only answer quality but also payload size, token cost, and system latency. Proprietary APIs expose little control over model internals or serving behavior, leaving client-side preprocessing as the main practical optimization space for downstream developers. Many such techniques have been proposed for visual offloading, yet their cost-quality impact on commercial cloud VLMs has never been studied. To fill this gap, we present VQABench, the first systematic benchmark that treats client-side input preprocessing as a controlled variable for cloud-VLM-based VQA. We evaluate 12 preprocessing techniques across three VQA datasets and four commercial VLMs from three providers, totaling 95,168 API calls. Our results show that preprocessing is not universally beneficial: its effectiveness depends on the target model, API paradigm, provider token-accounting rule, and task formulation. A poorly selected preprocessing strategy can increase deployment cost or latency while degrading answer accuracy. Overall, our benchmark clarifies when preprocessing helps, when it fails, and why, providing insights to guide future research and real-world deployment of VQA systems.
250. 【2608.07857】Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction
链接:https://arxiv.org/abs/2608.07857
作者:Fakrul Islam Tushar,Stephen Adamo,Geoffrey D. Rubin
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Foundation models provide, models provide transferable, Foundation models, predictions based directly, difficult to interpret
备注: Submitted to SPIE Medical Imaging 2027 conference
点击查看摘要
Abstract:Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. We developed concept bottleneck models that map two frozen CT foundation-model representations to eight radiologist-defined pulmonary-nodule attributes and predict malignancy from the estimated concepts and nodule size. The models included CT-FM, a whole-CT self-supervised encoder using a 96^3-voxel nodule-centered patch, and FMCIB, a nodule-focused contrastive encoder using a 50-mm crop. Eight ridge-regression concept heads were trained on 2,610 LIDC-IDRI nodules. Malignancy models were trained on LUNA25 and evaluated on a held-out internal test set and the external DLCS cohort. Concept fidelity was assessed using five-fold cross-validated R^2, and malignancy discrimination was assessed using AUROC with 95% confidence intervals estimated by patient-grouped bootstrap resampling. Concept fidelity was modest but higher for FMCIB than CT-FM for subtlety (R2, 0.24 vs. 0.11), spiculation (0.17 vs. 0.08), texture (0.17 vs. 0.07), and lobulation (0.15 vs. 0.05). Internally, the CT-FM and FMCIB concept+size models achieved AUROCs of 0.86 (95% CI, 0.80-0.92) and 0.86 (0.79-0.92), respectively. Externally, AUROCs were 0.72 (0.68-0.75) and 0.73 (0.70-0.76), compared with 0.73 for nodule size alone and 0.60 and 0.67 for the corresponding embedding only probes. Additive predictions could be decomposed into feature-level contributions and modified through controlled concept interventions. Concept bottlenecks provided transparent malignancy predictions with discrimination similar to nodule size alone, while differences in concept fidelity suggest that concept recovery depends on the underlying foundation-model representation.
251. 【2608.07848】IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility
链接:https://arxiv.org/abs/2608.07848
作者:Zhuangfan Huang,Chusheng Fang,Xiaosong Li,Yang Liua,Xiaoqi Cheng,Haishu Tan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:conditions requires fused, requires fused imagery, Robust perception, low-visibility conditions requires, Polarization Texture Injector
备注:
点击查看摘要
Abstract:Robust perception under low-visibility conditions requires fused imagery that jointly preserves infrared thermal saliency and polarization-derived structural details. However, existing infrared-polarization image fusion (IPIF) methods often overemphasize dominant infrared responses, causing weak yet informative polarization textures in dark regions to be suppressed. To address this issue, we propose IRPol-Fuse, an energy-structure coordinated IPIF framework for challenging low-visibility scenarios. The proposed framework contains three key modules: Polarization Attention Fusion for adaptive infrared-polarization allocation, Infrared Highlight Injector for highlight-guided infrared preservation, and Polarization Texture Injector for polarization texture restoration and fine-detail recovery. We further construct LI-PI, a dedicated infrared-polarization evaluation dataset for low-visibility and visually concealed scenes. Experiments on LI-PI and the public LDDRS dataset demonstrate that IRPol-Fuse achieves favorable performance in thermal target preservation, structural detail recovery, and visual naturalness. Region-aware evaluation and downstream object detection further verify that the proposed energy-structure coordination strategy effectively preserves both infrared target saliency and polarization-derived structural information. Code is available at this https URL .
252. 【2608.07835】SeqLoc: Beyond the Single Frame for Cross-View Geo-Localization in Feature-Sparse Scenes
链接:https://arxiv.org/abs/2608.07835
作者:Junwei Zheng,Yun Huang,Ruize Dai,Ruiping Liu,Yufan Chen,Kunyu Peng,Kailun Yang,Jiaming Zhang,Guangming Wang,Olaf Wysocki,Rainer Stiefelhagen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:structure-rich urban environments, Cross-View Geo-Localization, aligned OSM maps, structure-rich urban, urban environments
备注:
点击查看摘要
Abstract:Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes such as rural roads. To study this failure mode, in this work, we introduce CV-FSS, a benchmark that pairs sequential panoramas from five rural regions with aligned OSM maps, on which single-frame methods degrade drastically. We then propose SeqLoc, an online test-time sequence aggregation mechanism that recursively maintains a log-belief volume with three key components: (1) Entropy-Tempered Uncertainty (ETU) tempers each incoming pose likelihood volume by its normalized entropy; (2) Map-Guided Relocalization (MGR) mixes a map-shaped recovery distribution into the belief so that a suppressed true pose can recover; (3) Peak-Anchored Smoothing (PAS) derives the final pose at sub-grid precision. Extensive experiments on CV-FSS and CV-RHO demonstrate that SeqLoc outperforms single-frame localization by a large margin, improving both position and orientation recall by over 50%. The benchmark and source code will be made publicly available.
253. 【2608.07797】Drone-Assisted UAV-UGV Collaboration for Autonomous Navigation in Snow-Covered Terrain
链接:https://arxiv.org/abs/2608.07797
作者:Shreyam Gupta(1),P. Agrawal(2),Priyam Gupta(3),R. Gautam(1) ((1) Robotics Research Group, Indian Institute of Technology (BHU), Varanasi, India, (2) University of Colorado, Boulder, USA, (3) Intelligent Field Robotic Systems (IFRoS), University of Girona, Spain)
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:conventional methods ineffective, unstable ground render, ground render conventional, render conventional methods, collaborative UAV-UGV navigation
备注:
点击查看摘要
Abstract:This paper presents a collaborative UAV-UGV navigation framework for high-altitude, snow-covered terrain, where reduced visibility and unstable ground render conventional methods ineffective. We introduce a custom efficient U-Net architecture that falls under the computational constraints for real-time road segmentation, utilizing a novel synthetic snow data augmentation technique to achieve 96.5% segmentation accuracy. For UAV localization, we implement an Extended Kalman Filter (EKF) fusing onboard GPS and IMU data, achieving a maximum observed positional error of +-0.5 meters. The UGV position is determined via a visual tracking pipeline using YOLOv5 and depth data from the UAV's RGB-D camera. A dynamic path planning algorithm utilizes this segmentation to adjust for snow drifts, enabling successful navigation in obscured test environment with minimal deviation.
254. 【2608.07770】From Benchmark Performance to Tool Deployment: Human-in-the-Loop Anomaly Detection
链接:https://arxiv.org/abs/2608.07770
作者:Mike Szklarzewski,CJ George,Gavin Smithson,Christopher Stokes,Dakota Fulp,William M. Jones,Benjamin Wynn,Alexander Ur,Agit Yesiloz,Clint Kallenbach,Mark Swartz,Nathan DeBardeleben,Sharmistha Chakrabarti
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:Automated anomaly detection, real-world industrial conditions, anomaly detection methods, curated academic benchmarks, report strong performance
备注: 8 pages, 8 figures, 6 tables. Accepted as a regular paper at the 25th International Conference on Machine Learning and Applications (ICMLA 2026)
点击查看摘要
Abstract:Automated anomaly detection methods often report strong performance on curated academic benchmarks, but their behavior under real-world industrial conditions is less clear. In this work, we evaluate 19 unsupervised anomaly detection models on the BowTie dataset, a challenging manufacturing dataset with reflective surfaces, subtle defects, and profile-specific variation. In contrast to benchmark results, we observe that model performance is less stable than typically reported on standard benchmarks such as MVTec AD, highly sensitive to preprocessing, and inconsistent across conditions, with no single approach emerging as uniformly robust; a consensus audit further indicates that nominal-data quality affects deployment. Motivated by these findings, we developed and initially deployed a unified human-in-the-loop framework for manufactured-part inspection that combines image annotation, AI-assisted defect detection, and an integrated validation engine, replacing a prior manual visual inspection and documentation workflow. The system supports heatmap-guided defect review, SAM-refined candidate regions for inspector acceptance, rejection, or boundary adjustment, mask evaluation where annotations exist, and review history for inspector consistency and onboarding. Together, the results highlight the gap between benchmark performance and deployment reality, and provide a practical framework for addressing it.
Comments:
8 pages, 8 figures, 6 tables. Accepted as a regular paper at the 25th International Conference on Machine Learning and Applications (ICMLA 2026)
Subjects:
Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
ACMclasses:
I.2.6; I.2.10; I.4.8
Reportnumber:
LA-UR-26-23692
Cite as:
arXiv:2608.07770 [cs.LG]
(or
arXiv:2608.07770v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2608.07770
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
255. 【2608.07767】DINO-3DRA: Leveraging 2D Foundation Model Semantics for 3D Cerebral Aneurysm Segmentation
链接:https://arxiv.org/abs/2608.07767
作者:Jiayang Lu,Fengming Lin,Alejandro F. Frangi,Ali Sarrami-Foroushani
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:extreme class imbalance, Accurate aneurysm segmentation, rotational angiography, class imbalance, morphological similarity
备注: Accepted by MICCAI 2026
点击查看摘要
Abstract:Accurate aneurysm segmentation in 3D rotational angiography (3DRA) is hindered by extreme class imbalance, morphological similarity to vessels, and absent large-scale 3D pretraining. 2D vision foundation models encode dense structural priors from 1.7 billion images, yet naïve slice-wise transfer fragments anatomical continuity and destabilises optimisation. We propose DINO-3DRA, a dual-path framework achieving effective cross-dimensional semantic transfer by injecting frozen DINOv3 features into a 3D U-Net backbone via Room-Lite spatial mixing and calibrated residual fusion. On multi-centre 3DRA data, DINO-3DRA achieves state-of-the-art aneurysm segmentation (Dice: 0.758; HD95: 2.75 mm; +13% over nnU-Net) with only 5.72M trainable parameters. Ablation studies confirm that gains arise from structured cross-dimensional transfer rather than loss design alone, with bridged foundation features improving anatomical continuity between aneurysms and parent vessels. Without fine-tuning on CADA and SHINY-ICARUS, DINO-3DRA eliminates all catastrophic failure cases observed in baseline architectures, demonstrating robust generalisation across heterogeneous imaging protocols.
256. 【2608.07763】Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation
链接:https://arxiv.org/abs/2608.07763
作者:Anna Kołos,Grzegorz Statkiewicz,Karolina Seweryn,Katarzyna Kowol,Karolina Piosek,Wojciech Kusa
类目:Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:achieved strong performance, visual question answering, question answering, achieved strong, strong performance
备注: 28 pages. Preprint under review
点击查看摘要
Abstract:Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generation. However, they are predominantly trained on English-centric data, which limits their ability to handle culturally grounded visual understanding and leads to failures in interpreting region-specific meanings, symbolic content, and context-dependent visual cues. Existing benchmarks for cultural competence are often template-driven and focused on surface-level recognition, making them insufficient for evaluating deeper linguistic and pragmatic understanding in culturally situated settings. We introduce PoVisLE, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context. The dataset contains 1,117 images and 2,366 manually annotated VQA pairs. Overall, our dataset provides a controlled and challenging resource for assessing culturally grounded vision-language understanding beyond surface-level recognition.
257. 【2608.07760】XClipGS: Exact Half-Space Clipping for Medical Volume Gaussian Splatting
链接:https://arxiv.org/abs/2608.07760
作者:Zhongpai Gao,Benjamin Planche,Meng Zheng,Anwesa Choudhuri,Chaoyi Zhou,Terrence Chen,Ziyan Wu
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:Gaussian-splatting proxies enable, volumetric medical scans, proxies enable interactive, Gaussian-splatting proxies, plane exposes anatomy
备注:
点击查看摘要
Abstract:Gaussian-splatting proxies enable interactive rendering of volumetric medical scans, but a clipping plane exposes anatomy not constrained by external-view training and intersects primitives that conventional splatting can only keep or drop whole. We present XClipGS (eXact Clipping), which treats these as two separate problems: the render-time clip operator and supervision of the hidden interior. Under the local affine model used by EWA splatting, the ray integral of a half-space-restricted Gaussian factorizes exactly into its ordinary 2D footprint and a conditional Gaussian CDF whose argument is affine in pixel coordinates. The resulting closed-form per-pixel operator introduces no learned clipping parameters or auxiliary network and remains differentiable with respect to the primitive and plane. We use multi-distance reference views with varied clipping-plane axes and offsets to supervise the interior through the same operator. We also introduce a paired clipped/unclipped cut-face protocol with difference-referenced cut error (CDE) and culled-side leakage (Leak), because global image metrics dilute errors near the plane. On eight CT and MRI volumes with plane offsets not used for training, XClipGS attains the highest PSNR on every volume (33.56 versus 32.34 dB for ClipGS) while rendering at over 650 FPS, far above real time, versus 278 FPS. On voxel-axis cut-face views, it raises average band SSIM from 0.809 to 0.860 and leaks roughly 40 times less. Without retraining, it also achieves the best average across all four metrics on arbitrary-normal planes; on a fixed interior, it matches RaRa's face fidelity with about 16 times less leakage. Project page: this https URL
258. 【2608.07757】Rethinking 3D Segmentation from Individual LiDAR Scans: Incidence-Aware Sampling on the SIP Benchmark
链接:https://arxiv.org/abs/2608.07757
作者:Seongyong Kim,Jingdao Chen,Yong Kwon Cho
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:fully reflect real, site sensing conditions, reflect real site, real site sensing, sensing conditions
备注:
点击查看摘要
Abstract:3D scene understanding is increasingly important in construction, yet most methods are developed on curated datasets that do not fully reflect real site sensing conditions. In many workflows, individual LiDAR scans provide rapid local updates rather than complete scene representations, producing limited surface coverage, acquisition-driven density variation, and severe imbalance between dominant planar surfaces and sparse construction elements. Because large point clouds must be downsampled, sampling resolution and point allocation directly affect the balance between geometric detail and spatial context. This study evaluates these effects under a fixed per-fragment point budget and introduces an incidence-aware sampling strategy for individual LiDAR scans. The method maps points to a geometry-normalized manifold space for voxel-based selection while preserving original Euclidean coordinates for downstream learning. It requires only point coordinates and normals and no backbone modification. Using the Site in Pieces (SIP) benchmark, experiments with Point Transformer and PointNeXt show improved resolution-averaged segmentation performance, especially for non-planar elements and ladders, while reducing sensitivity to sampling resolution. The results show that acquisition-aware sampling can provide a more stable geometric representation and should be treated as an active component of individual-scan 3D segmentation rather than generic preprocessing.
259. 【2608.07750】Multi-Task Consistency-based Detection of Adversarial Attacks
链接:https://arxiv.org/abs/2608.07750
作者:Cong Chen,Jean-Philippe Monteuuis,Jonathan Petit
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Deep Neural Networks, Deep Neural, Neural Networks, found successful deployment, found successful
备注:
点击查看摘要
Abstract:Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to adversarial attacks has prompted concerns regarding their practical applications, specifically in the context of autonomous driving. Existing defenses often suffer from cost inefficiency, rendering their deployment impractical for resource-constrained applications. In this work, we propose an efficient and effective adversarial attack detection scheme leveraging the multi-task perception within a complex vision system. Adversarial perturbations are detected by the inconsistencies between the inference outputs of multiple vision tasks, e.g., object detection and instance segmentation. To this end, we developed a consistency score metric to measure the inconsistency between vision tasks. Next, we designed an approach to select the best model pairs for detecting inconsistencies effectively. Finally, we evaluated our defense against PGD attacks across multiple vision models on the BDD100k validation dataset. The experimental results demonstrated that our defense achieved a ROC-AUC performance of 99.9% detection within the considered attacker model.
260. 【2608.07749】LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks
链接:https://arxiv.org/abs/2608.07749
作者:Saed Moradi,Benyamin Ghojogh,M. Hadi Sepanj,Yimin Yang,Ashirbani Saha
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Parameter-efficient fine-tuning enables, limited computational resources, Parameter-efficient fine-tuning, narrow parameter subspace, computational resources
备注:
点击查看摘要
Abstract:Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resources, but a single low-rank update can constrain all task-specific changes to one narrow parameter subspace. This restriction may prevent the model from simultaneously representing globally shared task structure and localized residual directions required for generalization to unseen imaging domains. We introduce LoRSA, a global--residual adaptation framework that jointly learns a dense low-rank component and a dynamically structured-sparse low-rank component. The dense component captures globally coordinated task adaptation, while the structured component provides complementary residual corrections whose support evolves during training. We characterize the representational capacity, approximation properties, rank structure, and singular-subspace complementarity of this decomposition. We evaluate LoRSA for four-class breast-density classification using DINOv3-Base, with VinDr-Mammo as the source domain and MammosighTR and RSNA as unseen external domains. LoRSA remains competitive on the internal validation set and achieves the best external macro-F1 on both target datasets, improving upon the strongest competing method by 2.15 percentage points on MammosighTR and 3.09 percentage points on RSNA. Weight-matrix analysis further shows that approximately $92\%$ of the energy of each adaptation component lies outside the bilateral singular subspace of the other, indicating that the two components learn largely complementary update directions. These results suggest that organizing adaptation capacity into distinct global and residual paths can improve the external-domain generalization of parameter-efficiently adapted biomedical vision models.
261. 【2608.07742】BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language Reasoning
链接:https://arxiv.org/abs/2608.07742
作者:Saim Rehman,Muhammad Shafique
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:real-world situations due, Visual-language models, varying-quality input images, input images, frequently struggle
备注:
点击查看摘要
Abstract:Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this paper, we aim at analyzing VLMs' robustness by applying perturbations and distortions to the input images, such as blur or low contrast. Toward this goal, we propose BRUCE (Benchmarking Robustness Under Corruption Escalation, a multimodal reasoning fragility framework for scientific vision-language reasoning. State-of-the-art evaluation frameworks/studies primarily focus on clean-task accuracy and rarely analyze how reasoning stability degrades across robustness dimensions. Besides varying over a wide-range of input perturbations, BRUCE employs two novel metrics -- Robustness Corruption Index (RCI) and Traversal-RCI (T-RCI) -- to quantify how rapidly multimodal reasoning performance deteriorates in VLMs as visual corruption severity increases under progressive perturbation scaling. We evaluate BRUCE across chemistry and mathematical reasoning tasks for multiple datasets, while analyzing corruption-induced prediction failures in terms of four high-level reasoning domains: OCR-dependent reasoning, spatial reasoning, symbolic reasoning, and semantic failures, with each containing fine-grained corruption specific failure subtypes, thereby enabling an interpretable failure analysis.
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2608.07742 [cs.CV]
(or
arXiv:2608.07742v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.07742
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
262. 【2608.07735】Ghost Features and Spooky Transfer Learning for Hypercomplex-Valued Neural Networks
链接:https://arxiv.org/abs/2608.07735
作者:Guilherme Vieira Neto,Marcos Eduardo Valle
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Hypercomplex numbers extend, Hypercomplex numbers, additional imaginary components, introducing additional imaginary, numbers extend
备注: 2026 IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2026)
点击查看摘要
Abstract:Hypercomplex numbers extend the concept of complex numbers by introducing additional imaginary components. Besides increasing dimensionality, operations on the imaginary parts provide algebraic and geometrical properties that can be beneficial for solving machine learning problems. In this paper, we show how to create hypercomplex-valued neural network layers where the real part corresponds to the output of a traditional real-valued layer. The additional imaginary parts of these hypercomplex-valued layers produce what we call ``ghost features,'' which contain enhanced information that is not present in the output of the real-valued layer. Moreover, ghost features can be effectively integrated into a trained neural network through a process we refer to as ``spooky transfer learning.'' This approach allows us to harness the richness of ghost features, leading to more efficient neural networks. The source code and Jupyter Notebook are available at this https URL.
263. 【2608.07726】Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties
链接:https://arxiv.org/abs/2608.07726
作者:Ali Bahri,Hongliang Li,Soufiane Lamghari,Jie Chuai,Zhitang Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:including Young modulus, Estimating volumetric mechanical, Poisson ratio, including Young, Young modulus
备注:
点击查看摘要
Abstract:Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from vision alone, as visually similar objects may have substantially different material compositions and physical behavior. Existing approaches predict these properties independently across voxels, overlooking the piecewise-constant material structure of real objects and producing noisy or inconsistent estimates for voxels that share the same material, while lacking an explicit mechanism to resolve visual ambiguity. We introduce ViWi (Vision Meets WiFi), an object-centric framework for volumetric mechanical-property estimation. ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions. To complement visual appearance, ViWi incorporates a compact RF descriptor generated through WiFi-band electromagnetic simulation using permittivity and conductivity. The RF descriptor conditions the material slots with global composition cues that may be unavailable from images, while visual features preserve voxel-level spatial localization. Across volumetric mechanical-property and mass-estimation benchmarks, ViWi improves over the prior state of the art on four of six per-voxel metrics, while its vision-only variant improves all mass-estimation metrics. These results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone.
264. 【2608.07713】okenizer Generator Coupling in Medical Image Generation
链接:https://arxiv.org/abs/2608.07713
作者:Liam Chalcroft
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Latent medical image, fixed preprocessing, Latent medical, medical image generators, shared latent grid
备注:
点击查看摘要
Abstract:Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST study at 64x64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. In this controlled setting, rankings depend jointly on the tokenizer, generator, and sampler: the best quantizer changes with the generator, and validation-based sampler selection changes the apparent generator ranking. We retrain the vocabulary-1024 interaction block at three seeds and the interaction survives (6 of 9 pairwise quantizer comparisons exceed three seed standard deviations), and we scope the wider single-seed grid accordingly. Reconstruction PSNR alone is not a reliable selection criterion; we instead introduce a generator-free statistic, neighbour-conditional predictive gain, that separates the quantizer families by downstream generation quality (rank-AUC 1.00) where reconstruction PSNR and marginal token entropy do not. On LFQ-1024, retuning D3PM and SE-D3PM (selected on a held-out validation split) moves them from default FID-192 0.44/0.41 to 0.09/0.10 at lower NFE, replicated across seeds; the continuous references were not given an equivalent sampler sweep. We report FID-192 as an internal ranking metric; it ranks consistently with standard FID-2048 (Spearman 0.80) and with a label-free classifier two-sample test (0.78). We interpret these results through a rate-distortion-modelability framing, where modelability is conditional on the generator, sampler, and inference budget. All experiments are at 64x64 on low-resolution medical-style images, unconditional, and evaluated with non-clinical FID-based metrics, and we scope every claim to that setting. Code: this https URL and this https URL.
265. 【2608.07712】SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models
链接:https://arxiv.org/abs/2608.07712
作者:Ziqiao Yu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:predictive model receives, receives a self-supervised, self-supervised signal, predictive model, model receives
备注: 14 pages, 2 figures, 4 tables. Code: [this https URL](https://github.com/Oooorca/SpikeWorld)
点击查看摘要
Abstract:A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment is difficult when dynamics and semantics share parameters: freezing prevents adaptation, whereas weight updates require optimizer state and may alter the learned representation. Here we introduce SpikeWorld, a 1.45M-parameter sparse spiking model jointly trained for heterogeneous sensory prediction, semantics, image-text binding and action-conditioned dynamics. At deployment, all trained parameters are frozen. Delayed next-state residuals update two external paths: cumulative fixed-bank losses select the bounded action correction, while route-specific residual matrices refine next-state prediction. Neither path uses labels, teacher outputs, rewards, success signals or the true shift value. Joint optimization improves action next-state MSE by 17.10\% while also improving multimodal prediction, semantic accuracy and image-text retrieval. On held-out shear and attenuation streams, the combined external state improves aggregate prediction by 5.48\% and 30.01\%; its fixed-bank action path improves tracking by 24.20\% and 3.94\%, respectively. In a six-arm study comprising 450 new Meta-World trajectories (75 per arm), SpikeWorld raises frozen-policy reward by 7.90 (95\% CI [2.48, 14.06]); the 13.33-point success difference is descriptive (CI [0, 40]). For identical sensory inputs, model parameters and inherited semantic outputs remain bitwise unchanged. A 16-byte RLS estimator obtains the highest non-oracle reward on linear attenuation, showing that the contribution is not superior linear identification, but its integration with a frozen multimodal spiking checkpoint. Reference code is publicly available at this https URL.
266. 【2608.07693】CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
链接:https://arxiv.org/abs/2608.07693
作者:Quang Minh Dinh,Tuan Kiet Doan
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Generative traffic video, temporally coherent future, coherent future videos, traffic video forecasting, short observation history
备注: Accepted at ECCVW 2026
点击查看摘要
Abstract:Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation history and textual descriptions. In this paper, we present CosmosAlign, a generative traffic video forecasting framework built upon the pretrained Cosmos3-Nano world foundation model. Our approach is motivated by the observation that successfully adapting large pretrained world models to downstream forecasting tasks depends primarily on distribution alignment rather than increased model capacity. To this end, we propose a two-stage LoRA adaptation strategy that first aligns the conditioning-mode distribution with the target forecasting task, and then aligns the training captions with the model's native structured prompting interface through an LLM-based re-captioning pipeline. During inference, we further improve prediction quality using a fully training-free procedure consisting of consensus-based medoid sample selection and motion-adaptive blending of static scene regions. CosmosAlign achieves a final score of 76.49 on the AI City Challenge 2026 Track 5 benchmark, ranking first on the final leaderboard. Our code is publicly available at this https URL.
267. 【2608.07663】Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
链接:https://arxiv.org/abs/2608.07663
作者:Yeeun Choi,Youngbeom Yoo,Joon-Young Lee,Hyolim Kang,Seon Joo Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Large Language Models, Multi-modal Large Language, current Multi-modal Large, Language Models, Multi-modal Large
备注: Accepted to ECCV 2026 (Oral). Project Page: [this https URL](https://choi-yeeun.github.io/MERIT/)
点击查看摘要
Abstract:When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Models (MLLMs). This ultra-long setting necessitates a two-stage paradigm: query-agnostic memory construction followed by retrieval-based inference. Prior work invests in complex memory construction to pre-model high-level relations in videos, despite not knowing the downstream query at build time. We instead prioritize high-recall retrievability during memory building, and defer query-specific, high-level relation composition to inference time. To this end, we propose MERIT(Multi-key Episodic Retrieval with Inference-time Temporal expansion), a simple yet effective agentic framework for ultra-long video understanding. First, we formulate an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism. Second, we introduce a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction. This is achieved by expanding the temporal scope exclusively around the retrieved segments at inference time. By leveraging simple key-matching with this on-demand temporal expansion, MERIT achieves state-of-the-art performance across three long-video benchmarks: EgoLifeQA, LVBench, and Video-MME (Long).
268. 【2608.07651】An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
链接:https://arxiv.org/abs/2608.07651
作者:Jalil Jalili,Hossein Taghizad,Anuwat Jiravarnsirikul,Christopher Bowd,Akram Belghith,Raheleh Kafieh,Christopher A. Girkin,Sally L. Baxter,Robert N. Weinreb,Linda M. Zangwill,Mark Christopher
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Large language models, Large language, show promise, suffer from hallucination, interpretation but suffer
备注:
点击查看摘要
Abstract:Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency. We developed and validated an agentic AI framework integrating LLMs with specialized deep learning tools for glaucoma detection from fundus photography. The workflow had three steps: (1) LLM initial assessment; (2) function calling to invoke specialized tools for image quality (QAModel, FundaQ-8), glaucoma classification (SwinV2-Tiny), and optic disc/cup segmentation (SegFormer-B0); and (3) LLM reflection integrating the initial impression with tool outputs. Two LLMs (Gemini 2.5 Flash, GPT-5.4 mini) were evaluated on two public datasets (ORIGA, n=100; RIM-ONE-v3, n=100) under uncropped and cropped fields of view; all images were independently graded by a masked fellowship-trained glaucoma specialist. The agentic workflow improved classification accuracy by 16 to 47 percentage points across all conditions, reaching within 6 points of the specialist; on RIM-ONE-v3 the best configurations matched the specialist accuracy of 88%. LLM-alone approaches failed in two ways: GPT-5.4 mini showed positive bias (sensitivity 95-100%, specificity 0-5%), while Gemini 2.5 Flash varied stochastically between runs; the agentic workflow corrected both. Cup-to-disc ratio error fell 15-50% (MAE 0.156-0.228 to 0.104-0.132), and correlation with specialist grading rose from weak (r=0.12-0.39) to moderate-strong (r=0.59-0.84). Run-to-run consistency rose from near-random (kappa as low as -0.01) to near-perfect (kappa up to 0.96). Integrating LLMs with specialized tools addressed key limitations of LLM-alone approaches, including over-diagnosis and run-to-run variability. Gains held for both LLMs, suggesting generalizability across backbones, and may signal a shift from monolithic models toward orchestrated multi-agent systems in medical AI.
269. 【2608.07643】Data collection from highways: a geometric, class-agnostic approach to embedded vehicle counting
链接:https://arxiv.org/abs/2608.07643
作者:Lucas Gouveia Omena Lopes,William W. M. Lira,Alexandre M. Lima,Thales M. A. Vieira
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Traffic data collection, Single Board Computers, deep object detectors, missing in practice, collection is dominated
备注:
点击查看摘要
Abstract:Traffic data collection is dominated today by deep object detectors followed by tracking-by-detection, a pipeline that presupposes what is often missing in practice: a detector already trained on the class one wants to count. We revisit a purely geometric traffic-sensing pipeline for Single Board Computers in which detection is class-agnostic: moving objects come from background subtraction and thresholding, and counting is decided by a geometric rule on an imaginary line across the road, a software inductive loop detector. With no object model, training set or per-object trajectory, it runs faster than real time on Raspberry Pi class hardware. Two counting rules are described: a constant average speed rule, whose expected accuracy is derived analytically as about 86% under a Gaussian speed distribution, and a self-calibrating pre-calibration rule that recovers the lane geometry from blob statistics and counts edges of lane occupancy, additionally yielding per-vehicle average speed at no extra cost. Over four videos the latter counts with 83.3%-100% accuracy; in a field deployment it reaches 91% against 37.5% for a blob-tracking baseline under the same compute budget. We report the observations of that period in detail: the resolution floor below which accuracy collapses, the frame rate floor at which vehicles alias past the counting line, the gap between short curated clips and long uncontrolled footage, and the trade-off between Python (easier to tune, 100% CPU) and C++ (40% CPU, thermally viable). These are properties of the sampling geometry, not of the hardware of the time, and still constrain edge deployments. We close by arguing where motion-based, class-agnostic detection remains the right tool: open-set classes with no annotated data, tight power budgets, privacy-constrained installations, and the cold start of mining training crops to bootstrap a learned detector.
270. 【2608.07640】HeatCast: A Benchmark for Neighborhood-Scale LST Forecasting across 124 U.S. Cities
链接:https://arxiv.org/abs/2608.07640
作者:Jesus Guerrero,Isaac Corley,Leon Najafirad,Maryam Tabar,Paul Rad
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Land Surface Temperature, urban surface heat, Land Surface, Surface Temperature, surface heat
备注:
点击查看摘要
Abstract:Land Surface Temperature (LST) is a widely used satellite-derived measure of urban surface heat, but there is no shared benchmark for forecasting it at 30 m. Prior studies usually cover one to three cities, use kilometer-scale products, or do not release data and code. We introduce HeatCast, a Landsat-based benchmark for monthly LST forecasting across 124 U.S. cities from 2013 through June 2025. HeatCast contains 30 m monthly tiles with LST, elevation, surfacereflectance RGB, three spectral indices, broadband albedo, quality masks, and Local Climate Zone (LCZ) labels, together with a fixed temporal split, LCZ-stratified metrics, and a reference evaluation harness. We evaluate a CNN+LSTM and Earthformer on next-month forecasting, where Earthformer reaches 7.74 K RMSE against 10.42 K for the CNN+LSTM. Forecasting from the eight nonLST channels alone reaches 7.72 K, against 8.15 K from LST history and 8.68 K from RGB. The data, code, and weights are released under MIT at this https URL.
271. 【2608.07636】Adversarial Attacks on Deep OCR Systems
链接:https://arxiv.org/abs/2608.07636
作者:Wenbo Sun,Hongzong LI,Yanyun Wang,Jiahao MA,Shuxin Zhuang,Rong Feng,Shiqin Tang,Zi Liang
类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:optical compression medium, advances document recognition, enabling long-context OCR, low token cost, advances document
备注:
点击查看摘要
Abstract:Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black-box adversarial attack against a generative OCR vision-language model, where only the decoded string can be queried and no gradients, logits, or model internals are available. We recast the attack as a zeroth-order optimization problem driven by a bounded scalar loss defined directly on the string output via sequence similarity, and estimate the gradient with a random-direction finite-difference scheme whose query cost is independent of the image dimension. An Adam update with ell_infinity projection yields imperceptible perturbations for both untargeted and targeted objectives. Pilot experiments on Deep-OCR validate the string-only attack and evaluation pipeline and expose severe qualitative decoder failures, including repetition, truncation, and prompt leakage. They also show that controlled targeted rewriting remains substantially harder than untargeted degradation; we avoid claiming targeted success until the pre-registered evaluation is complete.
272. 【2608.07621】CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
链接:https://arxiv.org/abs/2608.07621
作者:Hsu-kuang Chiu,Stephen F. Smith
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:recently achieved impressive, achieved impressive performance, cooperative autonomous driving, autonomous driving agent, individual single autonomous
备注:
点击查看摘要
Abstract:Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primarily designed for an individual single autonomous driving agent with limited support for cooperative perception, reasoning, and planning. We present Cooperative Multi-agent Unified Driving with Reasoning (CMU-Drive), a closed-loop end-to-end benchmark for evaluating cooperative autonomous driving with multiple connected autonomous vehicles (CAVs) operating in safety-critical driving scenarios with background traffic participants. We further propose Vehicle-to-Vehicle Vision-Language-Action (V2V-VLA), a cooperative VLA model that integrates cooperative driving into a single forward pass by jointly generating driving actions, future waypoints, language reasoning, and communication policies. Experiments on CMU-Drive establish the first benchmark and baseline for cooperative VLA driving and provide a foundation for future research on multi-agent, closed-loop, end-to-end cooperative autonomous driving. Our code, benchmark, and model checkpoint will be publicly released to facilitate open-source research.
273. 【2608.07620】FlowErase-OPD: Multi-Concept Erasure via Anchored On-Policy Distillation in Flow Matching Models
链接:https://arxiv.org/abs/2608.07620
作者:Yi Sun,Yimin Zhou,Xinhao Zhong,Zhiqi Zhang,Junhao Li,Bin Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:raised increasing safety, increasing safety concerns, safety concerns due, flow matching models, Recent advances
备注:
点击查看摘要
Abstract:Recent advances in flow matching models have substantially improved the quality of text-to-image generation, but have also raised increasing safety concerns due to their potential to generate harmful or undesirable content. Existing concept erasure methods for flow matching models predominantly focus on removing individual concepts, while effectively erasing multiple concepts simultaneously remains challenging. We propose FlowErase-OPD, a framework for multi-concept erasure based on on-policy distillation (OPD). Our approach first distills multiple single-concept erased models into a unified LoRA module and introduces Anchored Multi-Teacher Distillation (AMTD), which incorporates a retention teacher to mitigate the trade-off between concept erasure and preservation of generative capabilities. To further improve the coordination of multiple erasure objectives, we develop Adaptive Retention Control (ARC), which dynamically adjusts the sampling frequency and loss weight of each erasure teacher, together with the relative contribution of erasure and retention teachers throughout training. Extensive experiments on nudity, object, and artistic-style erasure demonstrate that FlowErase-OPD consistently improves the trade-off between erasure effectiveness, image quality, and semantic alignment, achieving state-of-the-art performance across diverse multi-concept erasure settings. Furthermore, the resulting models exhibit strong robustness against adversarial attacks. These results highlight the potential of on-policy distillation as a principled framework for safe and controllable generation in flow matching models.
274. 【2608.07616】HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers
链接:https://arxiv.org/abs/2608.07616
作者:Dong Liu,Yanxuan Yu,Renata Borovica-Gajic,Ying Nian Wu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Vision transformers face, transformers face significant, face significant computational, significant computational overheads, Vision transformers
备注:
点击查看摘要
Abstract:Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficiency but sacrifices local context modeling. We propose \textbf{HSMLA (Hierarchical Softmax Multi-scale Linear Attention)}, which combines ReLU-based linear attention for global context, selective softmax refinement for critical local features, and multi-scale token representations via depthwise convolutions. HSMLA achieves superior accuracy-efficiency trade-offs: up to $4.2\times$ inference-time speedup across dense prediction tasks, $87.3%$ Dice with $3.2\times$ speedup on CT organ segmentation, and $94.2%$ AUC with $4.1\times$ speedup on pathology WSI.
275. 【2608.07598】NewtonGS: Physics-Structured Object-Level Neural Newtonian Dynamics for Gaussian Scene Animation
链接:https://arxiv.org/abs/2608.07598
作者:Lianlei Shan,Feiyang Ye,Yan Chen,Yong Wu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Gaussian scene requires, explicit object-level dynamic, Animating objects, Existing dynamic Gaussian, dynamic Gaussian methods
备注: 40 pages, 15 figures
点击查看摘要
Abstract:Animating objects in a static 3D Gaussian scene requires an explicit object-level dynamic state and a controllable model of object motion. Existing dynamic Gaussian methods primarily reconstruct time-varying scenes or simulate deformation, rather than provide compact object states for direct control. To address this gap, we present NewtonGS, a physics-structured framework for object-level state rollout and Gaussian scene animation. NewtonGS represents each object with a 22-dimensional state covering pose, linear and angular velocity, anisotropic scale and its rate, mass, and contact properties. Its Gaussian Neural Newtonian Dynamics (Gaussian-NND) model combines analytic translation, quaternion kinematics, gravity, damping, and scale-restoration dynamics with learned continuous and contact residuals. A discrete event map handles floor contact. Predicted poses and scales define a shared affine transformation that updates the means and covariances of all Gaussians associated with each object. We construct two procedurally generated datasets: State-32 for state-rollout evaluation and Gaussian-32 for state-to-Gaussian transformation. On both the in-distribution and velocity-range-shift splits of State-32, NewtonGS achieves lower trajectory RMSE, final displacement error, and velocity RMSE than five analytic baselines. Experiments on Gaussian-32 further demonstrate effective conversion from predicted states to animated Gaussian objects.
276. 【2608.07596】LIRA: Local Cross-Layer Information Routing for Vision-Language-Action Decoding
链接:https://arxiv.org/abs/2608.07596
作者:Zhewei Zhang,Puyue Wang,Guanren Qiao,Yijie Weng,Jiawei Hu,Guo Li,Lujia Wang,Junyan Wang,Tao Gu,Hongliang Lu,Guiliang Liu,Hong Jia,Xinhu Zheng
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:pretrained vision-language models, decoders remains underexplored, models transform representations, routes intermediate VLM, LIRA Query features
备注: 9 pages, 4 figures. Code and model checkpoints will be released upon acceptance of the paper
点击查看摘要
Abstract:Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Existing designs either expose only a narrow part of the representation hierarchy or rigidly match each decoder block to one VLM layer, restricting access to complementary task evidence across depths. We introduce LIRA, a local cross-layer action-conditioning mechanism that formulates VLM-to-action conditioning as depth-aware information routing. LIRA operates on task-token features and LIRA Query features derived from intermediate VLM states, then assigns each Parallel Fusion Block a depth-aligned local window centered on its corresponding VLM layer. Parallel Fusion Blocks aggregate neighboring LIRA Query features and integrate them with task-token features and proprioceptive inputs before action prediction. This routing interface leaves the backbone architecture, action decoder, and supervised training recipe unchanged. Across LIBERO, LIBERO-Plus, CALVIN ABC$\rightarrow$D, and real-world manipulation, LIRA improves the principal aggregate metrics over the VLA-Adapter baseline under the same 0.5B-parameter configuration. In zero-shot transfer to LIBERO-Plus, LIRA increases average success from 59.1% to 78.0%, an 18.9-point gain indicating improved robustness under controlled distribution shifts.
277. 【2608.07586】MAGIC-SSCIL: Manifold Anchoring and Geometric Incremental Calibration for Semi-Supervised Class Incremental Learning
链接:https://arxiv.org/abs/2608.07586
作者:Yousef Abdi,Mohammad Asadpour,Yousef Seyfari
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Semi-supervised Class Incremental, Class Incremental Learning, Incremental Learning, Semi-supervised Class, Geometric Incremental Calibration
备注:
点击查看摘要
Abstract:Semi-supervised Class Incremental Learning (SSCIL) is a severe challenge for neural networks, and it is hardest in the exemplar-free setting where no past data may be stored. Existing methods forget catastrophically due to feature drift, and their pseudo-labels become increasingly unreliable as the label space grows. In this paper, we propose MAGIC (Manifold Anchoring and Geometric Incremental Calibration), a framework that stabilizes plasticity without storing exemplars. MAGIC's design centers on two components. The first is Soft-Weighted Geometry Calibration (SWGC), which uses graph-based label propagation on the learner's plastic feature space to weight and calibrate class means and variances computed on the frozen backbone; from these calibrated Gaussians, we sample phantom features that stand in for data from previous tasks. The second is a Geometric Structural Alignment (GSA) objective that preserves representation topology by matching the relational structure of student and teacher heads and aligning feature anchors with the fixed classifier prototypes, locking the orientation of the feature space. Together, these constraints keep the adapter from drifting, so geometric relations between classes remain stable as new classes arrive. We implement MAGIC with a frozen ResNet-18 backbone and a learnable plastic adapter. Across CIFAR-100, CUB-200, and ImageNet-R, at label ratios of 1%, 5%, and 10%, MAGIC improves average incremental accuracy over most of the supervised CIL methods equipped with FixMatch and native SSCIL baselines; the largest gains occur in the fine-grained, low-label setting, where confidence thresholding fails most clearly.
278. 【2608.07585】LAVE: Latent Visual Evidence-Enhanced Planning for Video Tool-use Agents
链接:https://arxiv.org/abs/2608.07585
作者:Zijian Wang,Junnan Zhu,Rongzhen Li,Xiao Liu,Guohui Xiang,Quan Lu,Lijia Liu,Yining Wang,Jiang Zhong,Kaiwen Wei
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
关键词:Long-video understanding requires, understanding requires models, Long-video understanding, redundant video streams, visual
备注: 16 pages, 6 figures, 9 tables. Includes appendix
点击查看摘要
Abstract:Long-video understanding requires models to efficiently acquire and reuse sparse visual evidence from long and redundant video streams. Recent video tool-use agents address this challenge by iteratively invoking visual Tools at different temporal scales, but their Tool-Planner communication typically relies on textual observations. Such text-only interfaces provide lossy summaries of Tool computations, causing previously computed visual evidence not verbalized to be discarded and unavailable for subsequent planning. We identify this limitation as the Tool observation bottleneck and propose Latent Visual Evidence-Enhanced Planning (LAVE), a training-free framework for reusing latent visual evidence from completed Tool calls. LAVE introduces a dual-channel observation interface: the visible channel preserves the original textual trajectory, while the latent channel stores pre-verbal visual updates with their Tool roles, source-frame timestamps, and visual locations. During planning, LAVE retrieves evidence relevant to the current Planner state but not covered by textual observations, and integrates it through bounded timestamp-aligned latent updates with entropy-constrained frame-time routing. This enables video agents to reuse existing visual computation without additional training, frame replay, or modifications to the original orchestration. Extensive experiments on Video-MME, LongVideoBench, and CG-Bench show that LAVE consistently improves video tool-use agents across backbones. Under a comparable frame budget, LAVE improves the Video-MME overall score by 3.76 points over the strongest baseline, demonstrating the effectiveness of latent visual evidence reuse for multi-step video-agent planning.
279. 【2608.07584】ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making
链接:https://arxiv.org/abs/2608.07584
作者:Ningxin Pan,Hanyu Li,Yehui Tang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:increasingly support real-world, support real-world tasks, Vision-language models, made rapid progress, rapid progress
备注:
点击查看摘要
Abstract:Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such tasks, however, require more than rec- ognizing what an image contains: a model must use visual evidence to make a complete decision whose parts jointly satisfy global constraints. We introduce COMPLEXITYWORLD, a benchmark of 390 tasks across 39 domain-inspired visual worlds and 29 decision categories. Each task is generated from a hidden structured specification, rendered as a visual scene, and scored by an exe- cutable verifier that accepts any feasible solution. Under direct inference, all evaluated models ex- cept GPT-5.6-Sol remain below 40% verifier ac- ceptance rate (VAR), while GPT-5.6-Sol reaches 75.6%. Performance improves substantially when the same decision information is made explicit in structured form, yet varies sharply across equiva- lent visual presentations. Agent scaffolds provide smaller, model-dependent gains. Together, these results reveal a persistent visual-to-decision bot- tleneck that additional inference alone does not remove.
280. 【2608.07582】Predictive Failure Detection in Network Hardware Using Thermal Imaging and Deep Learning with Sensor Fusion
链接:https://arxiv.org/abs/2608.07582
作者:Ashly Joseph
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Unplanned network hardware, Unplanned network, malfunctions can interrupt, interrupt services, expensive downtime
备注:
点击查看摘要
Abstract:Unplanned network hardware malfunctions can interrupt services and result in expensive downtime in data centers. A deep learning-based predictive maintenance strategy is presented that utilizes thermal imaging and power sensor data to detect early indicators of equipment breakdown in routers, switches, and servers. A simulated dataset was generated comprising annotated thermal pictures and power readings indicative of three operating states: Normal, Warning, and Critical. Three ImageNet-pretrained convolutional neural network (CNN) models ResNet-50, InceptionV3, and VGG16 were assessed together with a multi-modal CNN-LSTM fusion model that integrates visual and sensor time-series information. Experiments were performed with and without pre-processing procedures, including region-of-interest (ROI) extraction and normalization. In the absence of pre-processing, CNNs attained moderate accuracy (e.g., ResNet-50 at 52%), but ROI-based pre-processing significantly enhanced performance (ResNet-50 accuracy reaching 91%). The CNN-LSTM model attained the greatest accuracy of 94%, with precision and recall approaching 95%, illustrating the effectiveness of multi-modal fusion. The results validate that domain-specific pre-processing and sensor fusion substantially improve early failure prediction, providing a potential foundation for proactive maintenance of network hardware through non-intrusive monitoring.
281. 【2608.07581】Multi-Branch Policy Optimization for Multimodal Large Language Models
链接:https://arxiv.org/abs/2608.07581
作者:Shuai Lyu,Yuning Gong,Ruiling Gao,Xiaoran Shang,Zhonghong Ou,Ping Zong,Yifan Zhu,Yuan Sun,Yang Qin,Peng Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Group-based reinforcement learning, large language models, language models typically, models typically rely, multimodal large language
备注: 10 pages,8 figures
点击查看摘要
Abstract:Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that applies a single advantage to all tokens in a response. However, multimodal reasoning involves substantially higher perceptual uncertainty than text-only settings, where the model must repeatedly re-examine visual information to verify intermediate interpretations, and different visual groundings can lead to divergent reasoning paths, making such uniform credit assignment particularly inadequate and causing relative advantages to progressively degenerate toward zero. To address these challenges, we propose Multi-Branch Policy Optimization (MBPO), a tree-based framework that constructs reasoning trees at vision-language decision boundaries, enabling sibling branches to explore diverse visual hypotheses and assigning segment-level credit through branch-relative advantages. We further introduce a temporal replay buffer to reuse informative segments while controlling policy staleness. Experiments on several multimodal reasoning benchmarks show that MBPO outperforms representative baselines, improving both learning signal quality and optimization efficiency. The code is publicly available at this https URL.
282. 【2608.07580】Real-time physics inversion for retrieval of sub-pixel wildfire temperatures from VSWIR imaging spectroscopy
链接:https://arxiv.org/abs/2608.07580
作者:William R. Keely,Philip G. Brodrick,Katherine Mistick,Adam Chlus,Robert O. Green,Philip E. Dennison
类目:Computer Vision and Pattern Recognition (cs.CV); Instrumentation and Methods for Astrophysics (astro-ph.IM); Machine Learning (cs.LG)
关键词:NASA Airborne Visible, Airborne Visible Infrared, VSWIR imaging spectroscopy, NASA Airborne, Airborne Visible
备注: In Review in Remote Sensing of Environment
点击查看摘要
Abstract:In this work, we present a wildfire temperature retrieval framework for VSWIR imaging spectroscopy data, employed on data from NASA's Airborne Visible Infrared Imaging Spectrometer (AVIRIS-3). The retrieval framework utilizes a full-physics approach in which a forward model is employed to resolve both solar and emitted radiance derived from a temperature distribution and utilizes the full spectral range in the residual fit. To optimize the forward model retrieval, we use state-of-the-art nonlinear least squares methods implemented for fast convergence on the on-board GPU, allowing for estimation of effective fire temperature within flight cadence. We verify the forward model assumptions on simulated spectra with an injected thermal signature and find good agreement with an RMSE of $41.8$ Kelvin (K). We apply the retrieval over the full 2025 FireSense AVIRIS-3 campaign, totaling 168 overflights with probable active fire spectra, and demonstrate a residual radiance fit of $\leq 10\%$ across bands in the short-wave infrared (SWIR). Lastly, we verify the applicability of the retrieved posterior fire temperature parameters to generalize to space-borne imaging spectrometers such as EMIT, by retrieving at coarsened spatial resolution. We find that the posterior distribution exhibits good coverage of the underlying sub-pixel temperature range with an absolute error of $30$ K across quantiles and a mean absolute error of $27.16$ K between spatial resolutions.
283. 【2608.07579】Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real
链接:https://arxiv.org/abs/2608.07579
作者:Abdullah Naeem,Anav Katwal,Ayon Dey,Noman Khan,Md Tamjidul Hoque
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:City Challenge, large indoor warehouses, perception in large, training and validation, large indoor
备注:
点击查看摘要
Abstract:The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real. The first is a geometry-first pipeline: YOLO11x detection, homography lifting to the world frame, class-level 3D size priors, multi-camera fusion, world-coordinate tracking, and offline tracklet stitching. The second is estimated-depth pseudo-LiDAR: monocular depth (D4RT, Metric3D~v2) back-projected into a fused point cloud and passed to a 3D detector (V-DETR), mirroring prior point-cloud winners that used depth. The gap is decisive: geometry-first reaches 13.0 3D HOTA (51.6 LocA), whereas pseudo-LiDAR collapses to 0.12 (9.2 LocA). We trace the collapse to cross-view inconsistency of monocular depth---scale correction is necessary but not sufficient---which domain-adaptation fine-tuning does not repair within budget. Within the geometry pipeline, offline stitching is the only intervention that helps; SAHI detection, appearance Re-ID, learned lifting, RT-DETR ensembling, test-time augmentation, and domain randomization all fail to beat the baseline detector. The bottlenecks are complementary: detection quality bounds the geometry route (DetA), localization consistency bounds pseudo-LiDAR (LocA). We release a complete, reproducible RGB-only pipeline and ablation.
284. 【2608.07577】Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
链接:https://arxiv.org/abs/2608.07577
作者:Felix Schaller
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Robotics (cs.RO)
关键词:wrong specific label, confident wrong specific, wrong specific labels, autonomous driving, driving must assign
备注: 6 pages, 4 figures, 1 table. Third paper in a series; v1 archived at Zenodo, doi: [https://doi.org/10.5281/zenodo.21593472](https://doi.org/10.5281/zenodo.21593472) . Code: [this https URL](https://github.com/freshNfunky/IE2025-Research-Paper)
点击查看摘要
Abstract:A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label or drop the object. Prior work in this series replaced the flat label set with a hierarchical taxonomy and a runtime abstraction rule, but evaluated it only on the boxes a closed detector already produces. This paper takes the layer open-world: we place taxonomic abstraction on top of class-agnostic region proposals so objects the closed detector never boxes can still be classified or flagged; we report a feasibility study of three open-world signals (class-agnostic segmentation, appearance-based out-of-distribution scoring, monocular depth) that shows why no single 2D cue suffices and how they compose; and we run the evaluation the earlier papers could not, a ground-truth leave-classes-out benchmark on real annotated objects. Holding out seven COCO classes and classifying their 235 ground-truth crops, a flat closed head emits a confident wrong specific label 100% of the time (37% of them in the wrong super-category, e.g. an animal named as a vehicle), whereas the hierarchical layer emits zero confident wrong specific labels and safely handles 94% of the objects (a correct super-category, or an explicit UNKNOWN OBSTACLE). We are explicit that this is a safety result, not a specificity one: the correct super-category is recovered only 26% of the time and the remaining 69% are conservatively flagged unknown. The contribution is an open-world perception layer that never makes a confident categorical mistake on an out-of-vocabulary object, together with an honest account of its cost.
285. 【2608.07575】Beyond Isotropic Assumptions: Continuity-Constrained Segmentation and GPU Morphometry for Nanoscale GBM Analysis
链接:https://arxiv.org/abs/2608.07575
作者:Arash Fatehi,Robin Ebbestad,Linus Butt,Hans Blom,Sigrid Lundberg,Hannes Olauson,Hjalmar Brismar,David Unnersjö-Jess,Thomas Benzing,Katarzyna Bozek
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:swelled tissue resolves, tissue resolves complex, under-sampled axial direction, Confocal microscopy, automated quantitative analysis
备注:
点击查看摘要
Abstract:Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic: along the under-sampled axial direction the structure can appear discontinuous, hampering reconstruction and automated quantitative analysis. The usual remedy upsamples the axial dimension to an isotropic volume before training a segmentation model, which requires dense annotations in the upsampled space, a prohibitive labeling burden. We present an end-to-end, GPU-accelerated framework that overcomes this without additional annotations. The model is trained on the native acquisition volume; random rotation of training patches leverages the well-resolved lateral plane to supply the missing axial information, and a z-axis continuity loss keeps neighboring slices consistent. We adapt both a convolutional (3D U-Net) and a transformer (SwinUNETR) backbone, aggregate overlapping patches by Gaussian consensus, and compute point-spread-function-corrected membrane thickness by ray-surface intersection on the GPU. We apply the method to the glomerular basement membrane (GBM), a thin, highly convoluted part of the kidney's filtration barrier that grows more irregular in disease. Segmentation accuracy matches inter-expert agreement. Continuity-aware training improves reconstruction smoothness and suppresses a periodic terracing artifact at minimal accuracy cost. We quantify GBM thickness across the reconstructed 3D surface and capture disease-related thickening, enabling fully automated anisotropic 3D morphometry of biological structures without dense volumetric labels or image restoration.
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:
arXiv:2608.07575 [cs.CV]
(or
arXiv:2608.07575v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.07575
Focus to learn more
arXiv-issued DOI via DataCite</p>
286. 【2608.07574】Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion
链接:https://arxiv.org/abs/2608.07574
作者:Nethmi Pathirana,Isuru Munasinghe,Dileeka Alwis
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:plays an important, important role, role in supporting, supporting the early, early diagnosis
备注:
点击查看摘要
Abstract:Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, automated analysis remains challenging due to class imbalance, inter-class similarity, and intra-class variability in dermoscopic images. This paper proposes a multimodal classification framework that combines Swin Transformer-based image features with structured clinical metadata to improve diagnostic performance through integrated visual-context learning. Experiments on a publicly available dataset show that the proposed model achieves a test accuracy of 92.55% and a macro F1-score of 91.33%, with strong performance across minority classes. Temperature scaling is applied as a post-hoc calibration method, resulting in a reduction in expected calibration error and improving prediction reliability, while uncertainty estimation is incorporated to further assess the confidence of model predictions. Qualitative explainability analysis further shows that the model focuses on lesion regions during inference. Therefore, the results demonstrate that multimodal fusion, combined with calibration and interpretability analysis, provides an effective and trustworthy approach for automated skin lesion classification.
287. 【2608.07572】BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference
链接:https://arxiv.org/abs/2608.07572
作者:Jinlong Yang,Jinke Wu,Lizilin,Yao Zhou
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Diffusion Transformers, demonstrated exceptional performance, video generation, demonstrated exceptional, exceptional performance
备注: 10 pages, 6 figures, 5 tables. Project page: [this https URL](https://youngkinlon.github.io/BRACE-Taming-Sharp-Irregularities-via-Barycentric-Rational-Forecasting-for-Fast-DiT-Inference/)
点击查看摘要
Abstract:Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their massive computational overhead, temporal feature caching has been proposed to bypass redundant computations. However, existing cache-then-forecast methods driven by derivative-based polynomials often cause severe quality degradation under high acceleration due to unstable long-step predictions. To address this bottleneck, we propose Barycentric Rational Forecasting with Chebyshev Enhancement (BRACE). Motivated by the observation that DiT feature trajectories are globally smooth yet frequently exhibit sharp irregularities and local non-smoothness, BRACE shifts the paradigm from derivative-driven polynomial extrapolation to feature-driven rational forecasting. Specifically, it maintains a local sliding window to cache sparse historical features and leverages adapted Chebyshev weights to formulate a barycentric rational function, directly aggregating these raw features to ensure numerical stability. Extensive experiments demonstrate that BRACE achieves state-of-the-art quality-efficiency trade-offs across various DiT architectures with negligible computational overhead.
288. 【2608.07571】A Review of Vision-Based Vehicle Detection for UAV-Based Traffic Monitoring: Experimental Insights and Future Directions
链接:https://arxiv.org/abs/2608.07571
作者:Jianlin Ye,Christos Kyrkou
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:based surveillance offers, Intelligent Transportation System, unmanned aerial vehicle, Intelligent Transportation, data collection capabilities
备注:
点击查看摘要
Abstract:In Intelligent Transportation System (ITS), unmanned aerial vehicle (UAV)-based surveillance offers an innovative solution to traffic surveillance with wide coverage and real-time data collection capabilities. In comparison to fixed ground-based infrastructure, UAVs are able to respond to dynamic traffic but present challenges such as vehicle detection at varying altitudes, compensation for motion-induced image variations and efficient processing of high-resolution images. Deep learning has been largely beneficial on improving the detection accuracy; however, for practical deployment, a critical assessment of the accuracy, latency, and harmonization with current transportation systems needs to be carefully considered. This survey reviews recent advancements in the UAV-based traffic monitoring, with a primary focus being deep neural network models for traffic analytics in various urban settings. Three main challenges identified in the literature are ensuring compatibility with traffic control systems, achieving real-time processing to optimize traffic flow, and maintaining robust detection in different environmental conditions. Existing solutions often lack comprehensive frameworks for utilizing UAV captured data to respond to incidents and manage traffic effectively. Future research should focus on optimal detection models, edge processing, and adaptive control integration to improve the responsiveness of urban traffic management.
289. 【2608.07570】COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping
链接:https://arxiv.org/abs/2608.07570
作者:Rui Yang,Wei Zhou,Dingyong Gou,Xiaohui Cui,Cong Li,Yinyin Gong,Yipo Huang,Jiliang Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Explainable aesthetic image, visually pleasing crop, Explainable aesthetic, aesthetic image cropping, localizing a visually
备注:
点击查看摘要
Abstract:Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. To support this setting, we introduce COMEX, a new benchmark built through image expansion and an IO-reversal pipeline. COMEX contains 33,161 quadruples, each consisting of an expanded image, a crop box, a composition category, and a composition-grounded explanation, enabling joint learning of crop localization, composition understanding, and explanation generation. We further propose a two-stage SFT+GRPO framework, where supervised fine-tuning establishes the structured output protocol and basic cropping ability, and GRPO further improves crop quality, composition prediction, and explanation faithfulness. We benchmark 15 large vision-language models and existing cropping methods on COMEX, establishing a comprehensive testbed for composition-grounded explainable aesthetic cropping. Experiments on both COMEX and prior benchmarks demonstrate the effectiveness and transferability of our framework, with strong performance across evaluation metrics.
290. 【2608.07569】Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators
链接:https://arxiv.org/abs/2608.07569
作者:Bowen Xue,Jiafeng Xiong,Xin Quan
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:control noise, editing in video-VAE, LFV, Direct spectral editing, frequency content
备注:
点击查看摘要
Abstract:Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reencode pass. However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent edits can disrupt VAE round-trip dynamics. We introduce \emph{latent-frequency validity} (LFV), which learns a compact VAE-specific spectral response and deploys it only when it improves decoded-target fidelity without worsening round-trip drift. LFV follows a validation-selected path from a diagonal per-frequency calibrator (C1) to full channel mixing (CM), making cross-channel capacity a controllable per-edit resource. Across 544 VAE--edit cells spanning six spectral families, LFV emits 423 cheap operators: 277 are handled by C1, while 146 (34.5\% of emitted operators) require channel mixing. On the primary 120-cell radial sweep, 99/100 emitted operators pass source-video-grouped held-out evaluation. Across five additional filter families, all 323 emitted operators pass held-out evaluation. Fully frozen OpenVid-fitted operators, including the validation-selected path coefficient, pass all 20 tested CogVideoX and HunyuanVideo generated-domain cells without adaptation. The selected response matches direct latent-filter latency and is about $3\times$ faster than pixel filter--reencode. The resulting maps reveal distinct VAE regimes, including strongly channel-coupled CogVideoX responses and a sharp Open-Sora high-band stability frontier.
291. 【2608.07567】mporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark
链接:https://arxiv.org/abs/2608.07567
作者:Marios Petrov,Sahana Vinayak,Targol Bakhtiarvand,Moses Smith Guddah,Adham Atyabi,Frederick Shic,Kevin A. Pelphrey
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Functional near-infrared spectroscopy, autism spectrum disorder, temporally aligned evaluation, existing approaches assume, approaches assume temporally
备注:
点击查看摘要
Abstract:Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing approaches assume temporally aligned evaluation. In practice, the optimal observation window varies across subjects due to differences in hemodynamic delay and neurovascular coupling, creating a temporal distribution shift that degrades performance. We formalize this as a \textit{cross-time-window transfer problem}, introducing a protocol that varies window length (2.5--10\,s) and offset within biological motion trials. Using topographic map representations of fNIRS recordings, we benchmark three vision architectures under two zero-shot baselines and eight adaptation strategies under leave-one-subject-out cross-validation ($N{=}124$). Key findings: (1) zero-shot cross-window accuracy is near chance (54--69\%); (2) ${\approx}5\%$ subject-specific fine-tuning recovers 90--96\%, while a subject-specific upper bound reaches 97--100\%, identifying inter-subject variability as the dominant barrier; (3) domain-adversarial and self-supervised strategies achieve 78--90\% without target-subject data; and (4) discriminative information is recoverable from windows as short as 2.5\,s. These findings provide a practical roadmap for deploying fNIRS-based ASD classifiers under realistic temporal variability.
292. 【2608.07565】What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
链接:https://arxiv.org/abs/2608.07565
作者:Zhijing Zhang,Jinpeng Yu,Xin Song,Bingnan Li,Chuyue Li,Changhui Du,Xiaolin Fang,Jiaming Liu,Ruihua Huang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Conversational assistants increasingly, assistants increasingly recommend, Conversational assistants, increasingly recommend follow-up, assistants increasingly
备注:
点击查看摘要
Abstract:Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p0.05).
293. 【2608.07562】Mechanistic Interpretability-Guided Selective Fine-Tuning of Vision-Language Models for Centimeter-Level Flood Depth Estimation
链接:https://arxiv.org/abs/2608.07562
作者:Nafis Fuad,Xiaodong Qian,Dongxiao Zhu
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Urban flooding poses, Urban flooding, street-level flood-depth estimates, transportation infrastructure, system provides real-time
备注: This Paper is accepted in International Conference on Machine Learning and Application (ICMLA) 2026
点击查看摘要
Abstract:Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution. This paper presents three vision-language models fine-tuned for continuous flood-depth estimation from street-level imagery: FloodLlama-Dense, a fully fine-tuned QLoRA baseline, and FloodLlama-MI5 and FloodLlama-MI6, interpretability-guided sparse variants that fine-tune only the top five and six causally relevant cross-attention layers identified through mechanistic interpretability analysis, respectively. Training uses an approximately 610,000-image subset of a 2.81-million-image synthetic corpus generated in Unreal Engine 5. The dataset combines single-vehicle subsets with 5 cm depth increments and mixed-vehicle subsets with 1 cm depth increments, spanning seven vehicle types, four weather conditions, and flood depths from 0 to 40 cm. FloodLlama-Dense achieves an MAE of 0.40 cm, an RMSE of 1.97 cm, and an Acc@5cm of 97.59%. Mechanistic interpretability analysis combining linear probing, logit lens, centered kernel alignment (CKA), and cross-attention entropy reveals a two-stage adaptation pattern: layers L13-L22 restructure visual representations, while depth first becomes linearly decodable at layer L23. FloodLlama-MI5 and FloodLlama-MI6 leverage this insight by fine-tuning only five or six of the eight cross-attention layers, achieving an 86-88% reduction in trainable parameters (6.55-7.86 million versus 54.4 million) with minimal accuracy loss. On a real-world benchmark, FloodLlama-MI6 achieves 98.62% accuracy, compared with 86.61% for the published STURM-FloodDepth baseline.
294. 【2608.07561】XEns-CKD: An Explainable Ensemble-Based Approach for Chronic Kidney Disease Stage Detection
链接:https://arxiv.org/abs/2608.07561
作者:Rehan Ahmad,Gousia Habib,Muhammad Shaban,Ishfaq Ahmad Malik
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Chronic kidney disease, CKD, silent disease, CKD progression, Chronic kidney
备注:
点击查看摘要
Abstract:Chronic kidney disease (CKD) is a silent disease. Its progression may not significantly hamper a person's daily routine. Human kidney function can be classified as normal or as one of the five stages of CKD. Early detection of the CKD stage can help patients understand the functional status of their kidneys and follow medical advice to slow CKD progression. In this paper, we propose XEns-CKD, a novel ensemble vision transformer-based scheme for CKD stage classification using ultrasound images. Three ViTs were trained on a private ultrasound image dataset using different training parameters. The performance of each ViT was evaluated using macro sensitivity, macro specificity, macro precision, macro F1-score, macro Youden index, the Matthews correlation coefficient (MCC), and macro balanced accuracy. The ensemble model achieved an overall classification accuracy of 86.36%. This work also emphasizes identifying and interpreting kidney regions affected by CKD progression. Explainable artificial intelligence techniques, including LIME, LRP, Attention-Min, and Attention-Max, were used to improve model transparency and clinical trust. An attention map combining the Attention-Min and Attention-Max results effectively identified and interpreted kidney regions affected during CKD progression from one stage to another. The attention map also highlighted the effects of CKD progression in these regions. Compared with existing methods, the proposed method classified the five CKD stages and normal kidney status with a 4% improvement in accuracy.
295. 【2608.07559】MVMD: A Multi-View Approach for Enhanced Mirror Detection
链接:https://arxiv.org/abs/2608.07559
作者:Yidan Shen,Yu Wen,Chen Zhang,Xin Fu,Renjie Hu
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:introduce significant challenges, mirrors introduce significant, fragmented spaces, resulting in inaccurate, inaccurate and unreliable
备注: This work has already published at WACV 2025, just want more accessibility
点击查看摘要
Abstract:In 3D reconstruction, mirrors introduce significant challenges by creating distorted and fragmented spaces, resulting in inaccurate and unreliable 3D models. As 3D reconstruction typically relies on multi-view images to capture different perspectives of a scene, detecting and labeling mirrors in multi-view images before reconstruction can effectively address this issue. However, existing methods focus solely on single-image detection, overlooking the rich information provided by multi-view setups. To overcome this limitation, we propose MVMD, a novel Multi-View Mirror Detection method, along with the first database specifically designed for mirror detection in multi-view scenes. The design of MVMD is grounded in the inherent associations between objects seen from different views and those reflected inside and outside of mirrors. These relationships are learned through cross- and self-attention mechanisms. MVMD consists of three key blocks: the Inter-Views Block tracks the shifts of objects within mirrors caused by changes in viewpoint; the Intra-View Block detects object reflections inside mirrors; and the Refinement Block sharpens mirror boundaries and enhances detected details. Experimental results show that our method improves accuracy by up to 2.6% and IoU by up to 11.1%, compared to single-image mirror detection techniques. This substantial improvement makes MVMD particularly effective for computer vision tasks, especially in enhancing the accuracy of 3D reconstruction in mirror-dense environments.
Comments:
This work has already published at WACV 2025, just want more accessibility
Subjects:
Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:
arXiv:2608.07559 [cs.CV]
(or
arXiv:2608.07559v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2608.07559
Focus to learn more
arXiv-issued DOI via DataCite
Journalreference:
2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
Related DOI:
https://doi.org/10.1109/WACV61041.2025.00904
Focus to learn more
DOI(s) linking to related resources</p>
296. 【2608.07558】Learning Physical Interaction: A Survey of Tactile- and Force-aware Robot Learning
链接:https://arxiv.org/abs/2608.07558
作者:Shilin Shan,Chuhao Zhou,Ruize Wang,Xinyan Chen,Xiangyu Chen,Xinyu Zhou,Boyu Ma,Iris Yuxuan Hu,Jingliang Li,Celeste Yuxuan Hu,Geng Li,Guohao Chen,Tianrui Zhu,Zhe Li,Yanjie Ze,Haoran Geng,Zhiyang Dou,Jianxin Bi,Yuejiang Liu,Jianshu Zhou,Jiachen Li,Paul Liang,Tatsuya Harada,Robert Katzschmann,Harold Soh,Na Li,Edward Johns,Danica Kragic,Jan Peters,Wojciech Matusik,Masayoshi Tomizuka,Jitendra Malik,Jianfei Yang
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Physically grounded robot, grounded robot intelligence, robot intelligence requires, intelligence requires robots, Physically grounded
备注: 53 pages, 7 figures
点击查看摘要
Abstract:Physically grounded robot intelligence requires robots to perceive, reason about, and regulate their interactions with the physical world. This capability is particularly critical in contact-sensitive manipulation, where successful task execution depends not only on visual perception and motion generation, but also on force regulation and adaptive control. In this context, recent robot learning methods have made substantial progress by integrating force, tactile, vision, language, and proprioceptive sensing into learned manipulation policies. In parallel, many systems adopt multi-phase architectures that combine high-level policies, action-refinement modules, and low-level controllers to bridge semantic task understanding with reactive physical execution. Despite these advances, existing surveys have not explicitly reviewed force- and tactile-aware robot learning from a unified perspective that jointly captures multimodal sensing and multi-phase system design. This survey addresses this gap by proposing TF-ART, a Tactile/Force-Aware Robot learning Taxonomy for multimodal and multi-phase frameworks, which maps individual methods into a unified hierarchical structure. The framework characterizes how recent works organize observation modalities, encode and fuse heterogeneous sensory inputs, generate and refine actions across multiple phases, and connect learned policies to reactive robot-end control. Building on this methodological view, we further examine the task settings and infrastructure requirements of physical interaction, thereby integrating both algorithmic and practical perspectives on force- and tactile-aware robot learning.
297. 【2608.07557】AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization
链接:https://arxiv.org/abs/2608.07557
作者:Peng Xu,Chengcheng Wang,Shaohua Wan
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Unmanned Aerial Vehicles, Navigation for Unmanned, Vision-Language Navigation, requires rapid, control in complex
备注: 7 pages, 3 figures, 4 tables
点击查看摘要
Abstract:Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.
298. 【2608.07554】Impact of Dataset Composition on Embedded Real-Time UAV Wildfire Detection Using Compact YOLO Models
链接:https://arxiv.org/abs/2608.07554
作者:Eduardo de los Santos,Andre S. Kelbouscas,Ricardo B. Grando,Bruna V. Guterres
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:unmanned aerial vehicles, diverse real-world training, UAV wildfire detection, wildfire detection systems, vision-based wildfire detection
备注: Paper accepted at the ICCAS 2026
点击查看摘要
Abstract:The development of vision-based wildfire detection systems for unmanned aerial vehicles is constrained by the limited availability of diverse real-world training images. This paper investigates the impact of dataset composition on embedded real-time UAV wildfire detection using compact YOLO models as a controlled validation family. Four training configurations were evaluated: real non-augmented, real augmented, hybrid non-augmented, and hybrid augmented, where the hybrid sets combine real wildfire images with AI-generated samples. The objective is to determine whether synthetic data mixing and image augmentation improve practical detection performance under resource-constrained deployment conditions. Experimental results show that the best overall operating point was obtained with the real non-augmented dataset, which achieved the strongest balance between recall and mean average precision for UAV-based wildfire detection. The results also show that neither hybridization with synthetic data nor augmentation produced a better final deployment choice. These findings suggest that, for embedded UAV wildfire detection, dataset realism and domain alignment are more valuable than increasing training set size through synthetic expansion.
299. 【2608.07550】Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions
链接:https://arxiv.org/abs/2608.07550
作者:Pengyang Yu,Yiou Wang,Zhongping Dong,Sahraoui Dhelim,Chun-Mei Feng,M. Tahar Kechadi
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:return structured chest-radiograph, models return structured, Vision-language models return, structured chest-radiograph findings, individual judgment
备注: 10 pages, 3 figures, 3 tables
点击查看摘要
Abstract:Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.
300. 【2608.07549】P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization
链接:https://arxiv.org/abs/2608.07549
作者:Zhenhong Sun,Haozhe Liu,Yifu Wang,Xibin Song,Senbo Wang,Huadong Mo,Daoyi Dong,Hongdong Li,Pan Ji
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Triangle meshes provide, topology connectivity makes, irregular topology connectivity, geometric sampling problem, accurate surface geometry
备注:
点击查看摘要
Abstract:Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geometric sampling problem: how to sample and organize geometric evidence into compact, structured and learnable tokens. Beyond field-centric volumetric sampling and edge-intersection surface sampling, we retarget mesh tokenization as \textit{local surface evidence sampling}: identifying the minimal geometric evidence inside each active voxel that is sufficient for deterministic surface recovery. To this end, we introduce \textbf{P2Voxel}, a pyramid pivot voxelization framework for compact and reconstruction-aware mesh tokenization. P2Voxel is built on three key innovations. Under the \textit{Local Planarity} assumption, Pivot Voxelization represents each active voxel with a surface pivot and an orientation sign, providing minimal local evidence that can induce the corner values required for deterministic reconstruction. Under the \textit{Spatial Complexity} assumption, Pyramid Pivot Voxelization exploits the spatial non-uniformity of real surfaces by allocating finer pivot tokens to geometrically complex regions while keeping smooth regions coarse and compact. Under the \textit{Block Reconstructability} assumption, a Pyramid VAE learns compact multi-resolution latent codes over locally reconstructable pivot blocks, avoiding the need to model the entire high-resolution voxelized shape as a dense global field. Together, these designs convert meshes into compact, structured, and learnable pyramid pivot tokens, enabling efficient mesh reconstruction for downstream 3D tasks.
301. 【2608.07548】SC$^{2}$-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments
链接:https://arxiv.org/abs/2608.07548
作者:Xuan Yao,Yuze Zhu,Junyu Gao,Zongmeng Wang,Changsheng Xu
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:Continuous Environments, make fine-grained navigation, requires agents, partial observability, agents to make
备注: Accepted by ICML 2026
点击查看摘要
Abstract:Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC$^{2}$-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. Our method derives feedback from world-model foresight to perform state-level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world-aware adaptation, which enables model-level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN-CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at this https URL.
302. 【2608.07547】Learning an Interior Layout Policy in a Domain Specific Language Action Space
链接:https://arxiv.org/abs/2608.07547
作者:Yuhao Lu,Weichen Zhang,Wenyi Xiao,Haohui Chen,Yiyun Fei
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Indoor scene layout, Indoor scene, scene layout generation, Indoor, challenging task
备注:
点击查看摘要
Abstract:Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room conditions to coarse 3D bounding boxes and neglecting structural elements such as doors and windows. More fundamentally, many prior approaches formulate spatial reasoning as direct coordinate prediction, thereby casting interior layout design as continuous regression over raw geometric parameters, which hinders the model from learning the underlying reasoning logic of intelligent layout design. We propose \textbf{LayoutDSL}, a novel LLM-based framework for learning an interior layout policy in a domain-specific language (DSL) action space. The DSL provides an explicit symbolic representation of layout information and serves as a structured action space for layout reasoning, where each action corresponds to an interpretable design decision. Under this DSL-based policy learning paradigm, we construct 3D-FrontDSL, a dataset of room-structure annotations paired with synthetic DSL action sequences for supervised fine-tuning. To promote a more generalizable and scalable policy with verifiable feedback, we design rewards grounded in interior design principles and physical plausibility, and optimize the policy via reinforcement learning. Extensive experiments demonstrate that LayoutDSL substantially improves spatial plausibility and design logicality over strong baselines and existing methods.
303. 【2608.07543】Performance of large language models in the optical diagnosis of colorectal polyps
链接:https://arxiv.org/abs/2608.07543
作者:Joshua C. Vences,William T. Tran,Nikko Gimpaya,Catharine M. Walsh,Rishad J. Khan,Robert Bechara,Asher C. Wiggins,Celine N. Rousan,Kaitlyn V.G.L. Morgado,Angie Ibrahim,Kevin H. M. Kuo,Daniel von Renteln,Alexander Hann,Dennis L. Shung,Michael A. Scaffidi,Charles Ménard,Joshua Landy,Samir C. Grover
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Accurate optical diagnosis, guides resection strategy, multimodal large language, Accurate optical, Study Aims
备注: 22 pages, 1 figure, 5 tables
点击查看摘要
Abstract:Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were 0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.
304. 【2608.07541】NeuroPilot: An Agent-Driven Smart Pipeline for Processing, Quality Control, and Managing Neuroimages
链接:https://arxiv.org/abs/2608.07541
作者:Yiyao Chen,Yucheng Li,Jungong Tong,Shaoqi Wang,Kunhao Zhou,Ziquan Wei,Monica Murea,Marissa DiPiero,Tingting Dan,Guorong Wu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
关键词:Transforming raw neuroimage, analysis-ready derivatives relies, Transforming raw, raw neuroimage archives, modality-specific preprocessing
备注: 21 pages, 6 figures
点击查看摘要
Abstract:Transforming raw neuroimage archives into analysis-ready derivatives relies on three brittle stages: data standardization, modality-specific preprocessing, and quality control (QC). While individual neuroimaging tools are well developed, their orchestration requires project-specific scripts, environment-adaptive tuning, and labor-intensive manual QC. To address this, we introduce NeuroPilot, a multi-agent system that digitalizes the expertise of neuroimage processing, QC, and data management into three LLM-invocable skills: dcm2bids-skill, neuroimage-pre-skill, and qc-agent-skill. The LLM-driven agent autonomously orchestrates workflows, generalizing various infrastructure settings into a single configuration to achieve the highest scalability. Demonstrating the system's generalizability, we deployed NeuroPilot across 17 cohorts (123,000 subjects) spanning infant to aging populations and multiple MRI modalities (structural, diffusion, functional). In practice, after standardizing data via the dcm2bids-skill, the agent dynamically routes datasets to the optimal neuroimage-pre-skill based on available modalities and cohort traits (e.g., dispatching T1w and fMRI data to fMRIPrep, or selecting specialized pipelines for infant cohorts). The qc-agent-skill then drives an evidence-based, semi-automated QC via a 3-D browser dashboard, utilizing a multi-tiered verification system to optimize failed cases and escalate complex issues for supervisor inspection. Quantitatively, our QC agent screened 558 production subjects, validating its automated flags against FreeSurfer's topology-defect metrics. The infant processing pipeline achieved a 100% (201/201) completion rate on QC-validated inputs. Importantly, NeuroPilot compresses the traditional 2--3 month timeline for training staff and processing complete datasets into a single week. NeuroPilot is deployed in this https URL.
305. 【2608.07510】World Simulator: Queer Erotica and the Absurdity of AI Video Models That Promise the World
链接:https://arxiv.org/abs/2608.07510
作者:Adam Cole,Mick Grierson
类目:Multimedia (cs.MM); Computer Vision and Pattern Recognition (cs.CV); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
关键词:model infinite realities, suggesting their ability, infinite realities, model infinite, World Simulator
备注: Accepted to Creativity and Cognition (CC '26), July 13-16, 2026, London, United Kingdom. 5 pages, 5 figures
点击查看摘要
Abstract:Increasingly, AI video models are marketed as "world simulators," suggesting their ability to model infinite realities. Despite such claims, these models systematically exclude significant aspects of embodied human experience, particularly sexuality. World Simulator is a video installation exploring the poetic friction between these universal claims and the models' inherent blindness. To do so, the work feeds explicit gay erotica into an AI video-to-video pipeline. Lacking the training data to recognize these images, the system hallucinates surreal alternatives, transforming intimate acts into banal scenes of kitchen appliances, strange architectures, and abstract flesh. By visualizing the limits of synthetic knowledge, the work challenges the hubris of the "world simulator" label, asking how a system, trained primarily on large filtered video datasets, can claim to simulate the world while remaining structurally blind to the body. Beyond this critique, we question the value of simulation itself, asking what forms of sensual representation might offer more expansive, life-affirming possibilities.
306. 【2608.07478】PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings
链接:https://arxiv.org/abs/2608.07478
作者:Jagpal Singh Jhala
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:official languages create, critical accessibility barrier, documentation exists exclusively, medical documentation exists, ASHA workers
备注: 7 pages 4 images/figure
点击查看摘要
Abstract:India's 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the patients who most urgently require this information - rural populations, ASHA workers, and patient families - are functionally excluded from understanding it. This paper presents PragyaDoc, a Universal Document Intelligence Framework that addresses this gap through a four-layer pipeline: a parallel ensemble OCR extraction layer, a geometric-lexical fusion layer, a deterministic domain structuring layer, and a dual-LLM medical reasoning and localization layer
307. 【2607.23755】DAP-Pose: Deep Temporal Alignment and Physics-aware Cross-modal Sensor Fusion for Robust Pose Estimation
链接:https://arxiv.org/abs/2607.23755
作者:Jianhan Lin,Yuchu Qin,Jiateng Yuan,Wenbo Zhang,Shuai Gao
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:mobile robotic systems, accurate pose estimation, multi-modal pose estimation, Bi-level Cross-modal Fusion, robust multi-modal pose
备注:
点击查看摘要
Abstract:Robust and accurate pose estimation with multi-modal sensors is fundamental for autonomous vehicles and mobile robotic systems in complex environments. In this paper, we propose DAP-Pose, a unified end-to-end model for robust multi-modal pose estimation. DAP-Pose introduces a Bi-level Cross-modal Fusion (BCF) module that captures complementary semantic and geometric motion cues from visual, inertial, and GNSS measurements. To handle temporal offsets, we designed a Deep Temporal Alignment (DTA) module that explicitly aligns asynchronous streams in latent space, enabling coherent motion modeling without strict hardware synchronization. Furthermore, we incorporate physics-aware constraints via manifold geometry and GNSS-guided absolute metric scale, enforcing motion consistency and mitigating drift. Experiments upon the public KITTI benchmark dataset were conducted to evaluate the performance of DAP-Pose against existing methods. DAP-Pose achieved the state-of-the-art performance, with the lowest average translation error ($t_{rel}$) of 1.31% and rotation error ($r_{rel}$) of 0.46$^{\circ}$. Furthermore, it accurately estimates poses and maintains robust performance under severe artificially injected temporal misalignment.
308. 【2608.09630】Unsupervised Domain Adaptation for Multitask Image Analysis in Realistic Context with Extreme Label Shift; Application to the CTAO first Large Sized Telescope
链接:https://arxiv.org/abs/2608.09630
作者:Michaël Dell'aiera,Thomas Vuillaume,Alexandre Benoit
类目:Instrumentation and Methods for Astrophysics (astro-ph.IM); Computer Vision and Pattern Recognition (cs.CV)
关键词:related unlabeled target, Unsupervised domain adaptation, labeled source domain, unlabeled target domain, Unsupervised domain
备注: This is the accepted version of the article published in Astronomy and Computing
点击查看摘要
Abstract:Unsupervised domain adaptation is a widespread set of methods that leverages the knowledge of a labeled source domain to train a model to perform well on a related unlabeled target domain. They generally introduce an auxiliary adaptation-related task that can be integrated into the multitask paradigm, which aims to merge multiple single-task models into a unified architecture. In this paper, we propose to associate domain adaptation and multitask balancing in the realistic context of an extreme class imbalance. Therefore, we propose a combined framework to cover and validate these approaches, and evaluate its performance in the physics-based context of the Cherenkov Telescope Array Observatory (CTAO). Along with a comparative study of some relevant adaptation techniques, we highlight the impact of extreme label shift and extend the investigations on importance weighting to rectify it. The complete code and results are published and available as open-source resources on Zenodo.
309. 【2608.09053】Diagnosing as Cardiologists Do: ECG Agents with Doctor-Grounded Priors for Clinical Reasoning Across Diseases and Populations
链接:https://arxiv.org/abs/2608.09053
作者:Hongxiang Gao,He-yang Xu,Yuwen Li,Minghui Zhao,Zhipeng Cai,Xingyao Wang,Chenxi Yang,Jianqing Li,Chengyu Liu
类目:ignal Processing (eess.SP); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Cardiologists interpret electrocardiograms, Cardiologists interpret, localizing waveform components, measuring rhythm, interval patterns
备注:
点击查看摘要
Abstract:Cardiologists interpret electrocardiograms by localizing waveform components, measuring rhythm and interval patterns, and translating these structured observations into diagnostic evidence. Whether this expert reading process can serve as an effective prior for ECG agents remains unclear. To address this question, we introduce LuminaECG, a clinically structured ECG reasoning framework that reformulates ECG interpretation as measurement-grounded visual reading. ECG signals are rendered on standard electrocardiographic grid paper to preserve the spatial and scale cues used in clinical reading. P-wave, QRS-complex, and T-wave boundaries are explicitly delineated, and color-coded segmentation decomposes the waveform into discrete visual measurement primitives. A general 2B vision-language backbone is then trained with low-rank supervised fine-tuning to associate these primitives with diagnostic reasoning, without architectural modification. Across open, proprietary, and ECG-specialist zero-shot baselines, LuminaECG improves both waveform measurement and diagnostic recovery. It reaches a clinically meaningful reader tier on the CODE-test benchmark, transfers across geographically diverse ECG datasets without retraining, and generates reports whose structure contains an emergent prognostic signal. These findings suggest that effective ECG agents require not only larger models, but supervision that preserves the alignment between measurable waveform evidence and clinical knowledge.
310. 【2608.08424】ARC: Augmented-Rank Conformalization for Changepoint Localization --- Finite-Sample Validity and Distribution-Robust Efficiency
链接:https://arxiv.org/abs/2608.08424
作者:Chenchen Peng,Mixia Wu,Qijing Yan,Zhiqi Shen,Jie Zhang
类目:Machine Learning (stat.ML); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Conformal changepoint localization, changepoint localization turns, Conformal changepoint, changepoint localization, localization turns
备注:
点击查看摘要
Abstract:Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage. Coverage is universal; efficiency is not. The oracle score is a likelihood ratio, so practical scores estimate density ratios, and set length deteriorates under heavy tails, skewness, and distribution shift, where no length guarantee applies. We propose ARC (Augmented-Rank Conformalization), a family of scores depending on the data only through within-segment ranks: rank-CUSUM location and scale channels, their fixed combinations, and a lightweight neural score frozen after synthetic training. Every ARC score inherits finite-sample coverage for every frozen weight configuration, including random initialization and mistraining. The main result is an efficiency transfer theorem: the entire ARC confidence set is almost surely invariant under strictly increasing marginal transforms, so the set length distribution depends on the data pair only through its rank structure, and lengths certified once hold verbatim across its monotone orbit, whereas a plug-in score's length changes with every re-expression. Across different rank structures lengths do change, and are reported as such. Classical rank-test theory positions ARC as targeting the optimal invariant score at bounded cost. Simulations confirm nominal coverage for all scores, including sabotaged networks, identical sets under monotone transforms where plug-in scores inflate, and smooth degradation where plug-in sets become vacuous; on the well-log benchmark ARC localizes annotated shifts to three to five candidates and flags misfit by an empty set. Two boundaries are stated rather than hidden: serial dependence destroys exactness, and trend-type alternatives lie outside the piecewise-exchangeable model.
311. 【2608.08226】High-Capacity Generalized Hopfield Networks
链接:https://arxiv.org/abs/2608.08226
作者:Victor Galitski
类目:atistical Mechanics (cond-mat.stat-mech); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Quantum Physics (quant-ph)
关键词:Riemannian manifold, introduced where memories, Generalized Hopfield networks, Hopfield networks, Phys. Rev
备注: 14 pages, 9 figures
点击查看摘要
Abstract:Generalized Hopfield networks are introduced where memories and neurons are continuous variables that lie on a Riemannian manifold. We explicitly focus on symmetric spaces associated with the special unitary groups SU(d), and use both numerical and analytical (replica) techniques to demonstrate an almost order of magnitude enhancement in critical capacity over the vector networks starting with d=3 and further rapidly growing with d. To circumvent the non-linear geometric constraints, we use a Lie algebraic method [following V. Galitski, Phys. Rev. A 84, 012118 (2011)] to exactly describe the classical neural network in terms of linear algebra in an auxiliary Hilbert space. It is shown that in contrast to the traditional Hopfield networks, memory recall in SU(d) Hopfields corresponds to neuron alignment along a top eigenvector of a spiked matrix, which is less susceptible to random matrix crosstalk than other models with continuous neuron variables. Physical platforms to realize SU(d) Hopfields are briefly discussed and physical (in addition to algorithmic) recall mechanism is demonstrated, where memory recovery occurs naturally through generalized Landau-Lifshitz-Gilbert dynamics. To illustrate SU(3) memory recall, we introduce a color (RGB) image encoding/decoding protocol and explicitly run image recovery on corrupted cues. Finally, we quantize the generalized Hopfields which are shown to reduce to Sachdev-Ye glassy type of models. Their many-body spectra generally feature two types of dark and memory bands, where the latter exhibits chaotic Wigner-Dyson level statistics that hides Hebbian data.
312. 【2608.08211】Retrieval-Augmented Generation-Based Color Restoration for Low-Light Image Enhancement
链接:https://arxiv.org/abs/2608.08211
作者:Li-Wei Lu,Shaou-Gang Miaou
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent low-light image, warm-tinted white objects, Recent low-light, structural fidelity close, exhibit systematic color
备注: 29 pages, 11 figures
点击查看摘要
Abstract:Recent low-light image enhancement (LLIE) methods have driven brightness and structural fidelity close to that of normally-exposed images, yet their outputs still exhibit systematic color shifts such as greenish skies, yellowish faces, and warm-tinted white objects. We attribute this to end-to-end LLIE training coupling brightness, structure, and color within a single network, leaving the color channels weakly supervised. We recast color restoration as an independent sub-problem and decouple it from brightness enhancement, realizing it as a general-purpose post-processing module built on retrieval-augmented generation (RAG). Rather than relying solely on parametric color priors learned during training, the module dynamically retrieves a reference image from an external high-quality color knowledge base and injects its color distribution into a color-restoration network to correct residual bias. The design has three components: (i) a dual-index FAISS retriever built on intermediate VGG19 features, capturing textural and structural similarity through global mean and variance statistics; (ii) GlobalSPHistAdaIN, which reduces the reference spatial-preserving color histogram to a global color vector and modulates network features via adaptive instance normalization, removing dependence on pixel-level correspondence; and (iii) a residual formulation that predicts a color correction over the front-end output. Across LOLv1, LOLv2-Real, and LOLv2-Synthetic, the module consistently improves color-specific metrics, and it remains effective when the front end is swapped among CPGA-Net++, LLFormer, FLIGHTNet, and IAT, confirming cross-front-end generality. Ablations show that a VGG19 dual index outperforms CLIP-based retrieval, indicating that color restoration depends on textural and structural similarity rather than high-level semantics.
313. 【2608.07632】JUMP-lite: Compact, reproducible benchmarking of cell representations
链接:https://arxiv.org/abs/2608.07632
作者:Alán F. Muñoz,Johan Fredin Haslum,Runxi Shen,Anne E. Carpenter,Shantanu Singh
类目:Quantitative Methods (q-bio.QM); Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:profiling captures rich, Image-based profiling captures, captures rich phenotypic, rich phenotypic signatures, JUMP Cell Painting
备注: Submitted to WACV 2027
点击查看摘要
Abstract:Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting now provide millions of images for systematic study. However, the scale of these resources, 115 TB for JUMP alone, and fragmented evaluation practices make systematic comparison of representation methods intractable for many researchers. Here we present Nahual, an open-source framework for reproducible model deployment, and JUMP-lite, a curated 116 GB subset of JUMP that is 1000 times smaller while preserving phenotypic diversity through careful selection of perturbations with high-confidence annotations and a storage reduction via lossy JPEG XL compression. With these, we benchmark five representation methods, including classical features (CellProfiler) and deep learning models (MorphEM, OpenPhenom, SubCell, DINOv2), and demonstrate that compression preserves downstream signal while standardized phenotypic activity and consistency metrics reveal meaningful performance differences across methods. Together, JUMP-lite and Nahual provide a foundation for accessible, reproducible benchmarking of image-based cell representations.
314. 【2608.07564】Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth
链接:https://arxiv.org/abs/2608.07564
作者:Sho Mitarai,Hikaru Kayo,Hisashi Ozaki,Yuichiro Imai,Megumi Nakao
类目:Image and Video Processing (eess.IV); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:dental surface geometry, integrating internal bone, internal bone structure, high-resolution dental surface, oral surgery
备注: 9 pages, 2 figures, 1 table
点击查看摘要
Abstract:In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating internal bone structure with high-resolution dental surface geometry. However, this registration is challenging because the two modalities share only a limited region in common, and their true correspondence is generally unknown. This uncertainty has prevented rigorous quantitative evaluation of registration accuracy. In this study, we propose a pseudoIOS evaluation framework in which a point cloud emulating an intraoral scan is generated from CT data within the same coordinate frame, so that the transformation between them is known by construction and can serve as a true ground truth. Using this framework, we propose a coarse-to-fine registration method that combines a domain-generalizable local descriptor (GeDi) for initialization-independent global alignment with the iterative closest point (ICP) algorithm for local refinement, and also evaluate the influence of metal artifacts on registration. In experiments on seven cases, ICP alone frequently converged to local minima from a perturbed initial position, whereas GeDi$+$ICP maintained submillimeter mean absolute error (MAE) across all evaluated jaw and artifact conditions (0.55-0.69 mm). A three-way repeated-measures analysis confirmed that GeDi$+$ICP was significantly more accurate than GeDi alone. Metal artifacts had a statistically detectable overall effect, but their absolute impact on GeDi$+$ICP was small.

