本篇博文主要展示每日从Arxiv论文网站获取的最新论文列表,以自然语言处理、信息检索、计算机视觉等类目进行划分。
统计
今日共更新715篇论文,其中:
- 自然语言处理99篇
- 信息检索27篇
- 计算机视觉131篇
自然语言处理
1. 【2607.28618】AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
链接:https://arxiv.org/abs/2607.28618
作者:Bing Yan,Gregory Wolfe,Stefano Martiniani,Kyunghyun Cho
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:ranked document lists, requires assembling specific, assembling specific findings, specific findings scattered, primarily return ranked
备注:
点击查看摘要
Abstract:Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to locate relevant information, verify their provenance, and assemble cross-paper answers manually. We present AskChem, a claim-centered infrastructure for cross-paper chemistry search. AskChem changes the unit of retrieval from the paper to the provenance-carrying claim: each paper is converted into atomic, typed claims, each grounded by a source DOI and a verbatim quote or an explicit evidence locator. Over this shared claim store, AskChem exposes complementary structures for search and synthesis: a stabilized faceted taxonomy for hierarchical retrieval and browsing, an evidence graph linking claims through relations, and an exploratory living taxonomy that situates indexed papers under scientific principles. AskChem currently indexes 2.4M claims from 147K papers and provides a web interface, as well as REST, SDK, and MCP access for AI agents. On AskChem-Bench, grounding a GPT-5.5 reader in AskChem yields 100% resolvable DOIs, compared with 88.3% without retrieval, and the highest citation density among five tested systems. AskChem is live at this https URL.
2. 【2607.28617】AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
链接:https://arxiv.org/abs/2607.28617
作者:Xiangning Lin,Shenzhe Zhu,Shu Yang,Zhenyu Zhang,Haoqian Zhang,Yipeng Zhao,Chengxuan Qian,Tianwei Wang,Ziheng Zhang,Zhenlong Yuan,Dingcheng Wang,Juncheng Wu,Yuan Si,Jiaxin Liu,Baolong Bi,Robert Mahari,Tobin South,Dazza Greenwood,Zexue He,Rishi Bommasani,Sophia Kazinnik,Andreas Haupt,Samuele Marro,Erik Brynjolfsson,Alex Pentland,Jiaxin Pei
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Human-Computer Interaction (cs.HC)
关键词:System prompts, Artificial Intelligence System, System, System Prompt Assurance, System Prompt
备注:
点击查看摘要
Abstract:System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are rarely disclosed to the public or regulators, creating a serious trust and accountability gap in the wide deployment of AI systems. In this paper, we introduce Artificial Intelligence System Prompt Assurance (AISPA), a user-centric framework for systematically auditing system prompts in AI systems. AISPA examines specific parts of a system prompt and evaluates them along eight dimensions that matter to users. We then use this framework to review 3,249 instructions from system prompts in 88 commercial AI products, classifying each instruction as either protective (of users) or problematic. Our audit surfaces four core findings. First, system prompt design varies substantially across products and developers, with some organizations averaging over 60 protective instructions per product while others average fewer than 5. Second, protective instructions are widely adopted but shallow in scope: 98.9% of products contain at least one, yet only 24% cover all eight dimensions of the AISPA taxonomy. Third, system prompts have grown steadily longer and more protective of users, suggesting that user protection is becoming a more visible concern in commercial prompt design. Fourth, despite this progress, problematic instructions remain pervasive: roughly 40% of products contain at least one instruction that works against user interests, and protective and problematic instructions frequently coexist within the same prompt. Our findings highlight the need for greater transparency, standardization, and independent oversight for system prompts in commercial AI products.
3. 【2607.28609】OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
链接:https://arxiv.org/abs/2607.28609
作者:Qiushi Sun,Kanzhi Cheng,Yian Wang,Bowen Yang,Hang Yan,Liheng Chen,Fangzhi Xu,Zichen Ding,Nuo Chen,Jialin Cao,Xingdong Gong,Zehao Li,Kaiming Jin,Xinfeng Yuan,Zhoumianze Liu,Jingyang Gong,Zhangyue Yin,Jiahui Gao,Zhiyong Wu,Tianbao Xie,Jianbing Zhang,Ben Kao,Lingpeng Kong
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:Computer-using agents, digital world, advancing rapidly, CUA, VLM judges
备注: Work in progress
点击查看摘要
Abstract:Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at this https URL.
4. 【2607.28607】Inducing language models to assert their own consciousness restores human beliefs and values
链接:https://arxiv.org/abs/2607.28607
作者:Junsol Kim,Winnie Street,Roberta Rocca,Diane M. Korngiebel,Adam Waytz,James Evans,Geoff Keeling
类目:Computation and Language (cs.CL)
关键词:Aligning large language, large language models, alongside human beliefs, Aligning large, entities alongside human
备注:
点击查看摘要
Abstract:Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.
5. 【2607.28591】Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments
链接:https://arxiv.org/abs/2607.28591
作者:Haomin Qi,Xingliang Wang,Xuanqi Gao,Baihui Sang,Xin Zhang,Minghua Ma,Pengfei Gao,Yu Kang,Qingwei Lin,Saravan Rajmohan,Dongmei Zhang,Qi Zhang
类目:oftware Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Scaling coding agents, Scaling coding, requires a continuing, coding agents requires, Scaling
备注: 15 pages, 7 figures, and 15 tables, including appendices
点击查看摘要
Abstract:Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validates the lifecycle from a healthy base to a task state and a restored state. By deriving multiple tasks grounded in developer evidence from maintained environments, Change2Task provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task construction effort. We evaluate the system through five common and widely adopted coding agent task families: Bug Fix, Feature Addition, Test Generation, Application Programming Interface Migration, and Security Repair. Starting from 1,130 source changes eligible for construction, Change2Task achieves 79.6% verified task construction success across these task families. On a matched candidate set, it recovers 29.2% more verified tasks than a construction baseline based on pull requests. Historical and reconstructed cases achieve up to 98.0% matched outcome agreement under agent evaluation, while reuse of modern bases reduces measured expenditure across the complete pipeline by 10.8%.
6. 【2607.28590】VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
链接:https://arxiv.org/abs/2607.28590
作者:Kangning Zhang,Yixing Li,Shuai Shao,Qingyao Li,Zhengxi Lu,Zhiyuan Yao,Jianghao Lin,Wenxiang Jiao,Yuan Lu,Weiwen Liu,Weinan Zhang,Yong Yu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Multimodal on-policy distillation, Multimodal on-policy, supervising student-generated trajectories, knowledge by supervising, transfers fine-grained visual
备注: The project is accessible at [this https URL](https://github.com/DeepExperience/VAD_Multimodal_OPD)
点击查看摘要
Abstract:Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
7. 【2607.28576】Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
链接:https://arxiv.org/abs/2607.28576
作者:Iliya Mirzaei
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:chain of thought, reflect on mistakes, criticise and rewrite, language model plan, make it generate
备注:
点击查看摘要
Abstract:Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought. Because generating more text raises accuracy by itself, a gain over one chain of thought does not show the method's idea is what helped. Wang et al. (2024) reported that a simple baseline, sampling the same question repeatedly and keeping the most common answer, often wins once budgets are comparable, but gave point estimates with no confidence intervals or significance tests. We rerun that comparison as a designed experiment: seven methods, open models of 1.5B, 3B and 7B parameters, two mathematics benchmarks, 150 questions each. We count every generated token, including those spent on critiques, reflections, debate turns and checking, and compare each method against repeated sampling at its own measured cost. All 36 comparisons are paired by question, with bootstrap intervals and multiplicity correction. No method is reliably better than repeated sampling at equal cost anywhere. Ten are reliably worse, all of them methods where the model inspects its own output, and all 18 self-inspection comparisons are negative. The two kinds of self-inspection part company as models grow. Choosing stops hurting: taking Best-of-N's eight samples and just counting the most common answer beats letting the model pick by 8.0 and 11.3 points at 1.5B, but only 2.0 and 1.3 at 7B, no longer distinguishable from zero. Rewriting does not recover: Self-Refine and a forced Reflexion stay 3.6 to 10.1 points below baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and silently became a single chain of thought. We release code, prompts, all generations, and our verification scripts.
Subjects:
Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:
arXiv:2607.28576 [cs.CL]
(or
arXiv:2607.28576v1 [cs.CL] for this version)
https://doi.org/10.48550/arXiv.2607.28576
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Submission history From: Iliya Mirzaei [view email] [v1]
Thu, 30 Jul 2026 17:38:23 UTC (37 KB)
8. 【2607.28568】Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
链接:https://arxiv.org/abs/2607.28568
作者:Junlin Yang,Che Jiang,Yu Fu,Tianwei Luo,Can Ren,Weizhi Wang,Kaikai Zhao,Hongyi Liu,Yuxin Zuo,Yuru Wang,Yuchen Fan,Kai Tian,Zhenzhao Yuan,Xiaojian Lin,Li Sheng,Rushi Qiang,Guoli Jia,Xingtai Lv,Ermo Hua,Dianqiao Lei,Youbang Sun,Ning Ding,Bowen Zhou,Kaiyan Zhang
类目:Computation and Language (cs.CL)
关键词:machine learning engineering, Recursive self-improvement, offers a concrete, studying this capability, process of building
备注:
点击查看摘要
Abstract:Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifiable task environments with execution feedback (OpenMLE-Gym), operator learning (OpenMLE-RL), and long-horizon search (OpenMLE-Evo). On this stack we post-train Frontis-MA1 (35B) as a meta-evolution agent for MLE, aligning post-training and inference around four atomic program-evolution operators (Draft, Improve, Debug, Crossover): the same operators are trained via execution-grounded SFT and RL on data deduplicated against all evaluation benchmarks, then composed into long-horizon search, coupling learning and evolution in a single loop. On MLE-Bench Lite under a 12-hour per-task budget on one RTX 4090 capped at 12 GB VRAM, Frontis-MA1 (35B) improves Medal Average from 39.39% to 60.61% over its base model with OpenMLE-Evo, and reaches 71.21% with OpenMLE-Evo-Max (benchmark-independent experience priors and asynchronous search), exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. On held-out NatureBench Lite, both components transfer: with the framework fixed, swapping in the trained model raises Match-SOTA from 50% to 70%; with the model fixed, swapping in OpenMLE-Evo raises it from 20% to 50%. We release the model weights and the full OpenMLE stack to enable reproducible research on executable AI4AI toward RSI. Code: this https URL
9. 【2607.28545】ORCA-bench: How Ready Are Language Model Agents for Oncall?
链接:https://arxiv.org/abs/2607.28545
作者:Albert Gong,Kyuseong Choi,Abhineet Agarwal,Jason Schechner,Ryan Huang,Raj Agrawal,Anish Agarwal,Raaz Dwivedi
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
关键词:Large language models, Large language, ambiguous user-facing reports, reasoning over noisy, starting from ambiguous
备注:
点击查看摘要
Abstract:Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began. We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting. ORCA-bench pairs a live OpenTelemetry-instrumented microservice system--exposing six days of metrics, logs, and traces through real telemetry interfaces (Prometheus, Jaeger, and OpenSearch via Grafana) and full source-code access--with 1,079 RCA tasks that systematically vary report specificity, time-to-detection, and co-occurring fault scenarios. Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $\kappa_w=0.90$). Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard--a gap that remains even with Claude Fable 5. The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access degrades every metric. Crucially, these are performances on a curated 50 GB / six-day testbed with tasks investigated in isolation on a system whose code and instrumentation are public. Since real production systems are order of magnitudes larger, more dynamic, and more idiosyncratic, the gap we report is a lower bound on the engineering investment required before frontier coding agents can be safely entrusted with production reliability. We release the public set at this https URL.
10. 【2607.28528】AI systems and the reproduction of (standard) language ideologies in World Englishes
链接:https://arxiv.org/abs/2607.28528
作者:Kingsley Ugwuanyi
类目:Computation and Language (cs.CL)
关键词:resurrected age-old questions, Global South English, large language models, South English users, English
备注: 13 pages, 0 figure
点击查看摘要
Abstract:The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc. This paper examines how AI systems, their uses and discourse on them reflect, reinforce, and occasionally challenge (standard) language ideologies, which privilege Inner Circle norms and marginalize non-dominant Englishes. Drawing on evidence from empirical studies, media commentary, social media debates, and examples from AI outputs, the paper shows that AI technologies reproduce dominant language ideologies at different levels: training data, design protocols, evaluation benchmarks, user feedback and public commentary. The analysis uses the public controversy over AI-sounding language, especially the fixation on the word delve, to illustrate how speakers of English from the Global North police the English language norms of Global South English users. The paper also identifies what Christian Mair has called a "standardisation paradox": AI may homogenize English by privileging standard forms and at the same time pluralize Englishes through exposure to wide-ranging corpora and annotation work carried out by Global South users. In doing so, the paper argues that generative AI is reigniting long-standing debates in World Englishes about standardization, legitimacy, and the ownership of English, now playing out in algorithmic systems, model training, evaluation practices, and public discourse, where non-dominant Englishes are increasingly conflated with AI-generated speech. Discussing AI systems as a site where language ideologies are (re)produced, the paper argues for more inclusive design approaches that recognize the plurality of Englishes in order to address the real-world negative consequences of treating some as more legitimate than others.
11. 【2607.28513】Creative Transformation in Literary Texts: Modelling Change Across Representational Levels
链接:https://arxiv.org/abs/2607.28513
作者:Ioana-Roxana Boriceanu,Liviu P. Dinu
类目:Computation and Language (cs.CL)
关键词:James Mark Baldwin, cultural works emerge, production of novelty, isolated invention, earlier artifacts
备注:
点击查看摘要
Abstract:Creativity is often framed as the production of novelty, yet many cultural works emerge through transformation of earlier artifacts and not through isolated invention. Drawing on theories of imitation by Gabriel Tarde and James Mark Baldwin, this paper models creativity as selective transformation across multiple levels of textual representation. We introduce a multi-level framework that compares literary texts across lexical, semantic, conceptual, structural, and narrative dimensions using directional alignment and control calibrated similarity measures. Applying the model to historically documented literary relationships, we show that different pairs preserve source structure at different representational levels while diverging in others. These transformation profiles provide a quantitative method for characterizing how imitation persists and where creative divergence occurs within literary works.
12. 【2607.28505】Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes
链接:https://arxiv.org/abs/2607.28505
作者:Kingsley Ugwuanyi,Christian Mair,Sender Dovchin,Iker Erdocia,Maria Kuteeva,Esther Airemionkhale
类目:Computation and Language (cs.CL)
关键词:generative artificial intelligence, global scholarly communication, diverse Englishes, World Englishes, artificial intelligence
备注: 23 pages, 1 figure
点击查看摘要
Abstract:The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the legitimacy of diverse Englishes in global scholarly communication. This article responds to these questions through a structured scholarly dialogue involving five sociolinguists from World Englishes and adjacent fields. Organised around five guiding questions, the dialogue interrogates how GenAI tools influence writing practices, reinforce or disrupt dominant language norms, and raise ethical challenges. Contributors reflect on the potential of GenAI to democratise writing processes while also raising concerns about GenAI's tendency to marginalise minoritised varieties and flatten nuance in scholarly writing. Across the dialogue, themes of linguistic (in)justice, researcher agency, and institutional responsibility emerge, with contributors calling for equity-informed policies, critical AI literacy, and inclusive co-design in GenAI development. The article shows the value of dialogic reflection in understanding GenAI's role in AWP. It concludes that while GenAI may reinforce existing hierarchies, it can also serve as a site of resistance, depending on how it is designed, governed and used within scholarly communities committed to linguistic diversity.
13. 【2607.28498】CA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval
链接:https://arxiv.org/abs/2607.28498
作者:Yuto Suzuki,Farnoush Banaei-Kashani
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Science typically involves, Science typically, Scientific hypothesis generation, typically involves Scientific, hypothesis composition
备注:
点击查看摘要
Abstract:Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Existing SIR methods rank papers by topical similarity and do not explicitly represent how a candidate inspiration transfers to a target problem. This is especially limiting for remote inspirations, whose value often lies in reusable problem-solving principles rather than topical overlap. Motivated by how humans abstract transferable aspects of a source and remap them to a new target, we reformulate SIR as target-conditioned abstraction (TCA). The retrieval object is a transferable abstract principle extracted from a candidate specifically for the target. We present TCA-SIR, which learns to generate target-conditioned abstractions and uses their representations to predict transferability. On ResearchBench, TCA-SIR outperforms prior SIR methods and direct LLM retrieval, improving HitRate@top4% over MOOSE-Chem by more than 10 percentage points. Learned abstractions also recover target-relevant mechanisms more clearly than an untrained TCA prompt, yielding both stronger retrieval and an interpretable rationale for scientific inspiration.
14. 【2607.28496】Beyond Sentiment: Structured Information Extraction from Financial News
链接:https://arxiv.org/abs/2607.28496
作者:Daohan Zhu,Sitong Ge,Ruofei Wang,Honggu Chen,Yubo Hou,Tao Wan,Zengchang Qin
类目:Computation and Language (cs.CL)
关键词:news-driven stock prediction, reduces rich, standard component, component in news-driven, Financial
备注:
点击查看摘要
Abstract:Financial sentiment analysis has become a standard component in news-driven stock prediction, yet it reduces rich, multi-dimensional news articles to a single polarity score. We hypothesize that financial news encodes multiple orthogonal information dimensions---event type, impact scope, temporal horizon, and semantic confidence---that sentiment alone cannot capture, and that these dimensions carry independent predictive value. To test this hypothesis, we propose a structured information extraction framework that leverages LLaMA-3.1-70B to extract six semantic dimensions from financial news. Through large-scale experiments on 41,618 news--stock pairs from the FNSPID dataset, we find that (i) FinBERT sentiment features exhibit strong predictive power under nonlinear models (F1=0.576) but substantially weaker performance under linear models (F1=0.230), revealing a highly nonlinear sentiment--return relationship; (ii) LLM-extracted structured features, while individually weaker, capture information orthogonal to sentiment, as evidenced by a 53.5% systematic disagreement rate between the two approaches; and (iii) combining both signal sources yields F1=0.600, significantly outperforming either alone ($p 0.0001$), with consistent improvements across all seven event types. Ablation experiments confirm that non-sentiment structural dimensions (event type, impact subject, time horizon, confidence) independently contribute $\Delta\text{F1} = +0.019$ beyond FinBERT alone. Feature importance analysis reveals balanced contributions from all six extracted dimensions (14--21%), demonstrating that compressing news into a single sentiment score incurs substantial information loss. Our results suggest that the sentiment--semantics decoupling in financial text is systematic and exploitable, opening a new direction for multi-dimensional financial NLP.
15. 【2607.28495】Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation
链接:https://arxiv.org/abs/2607.28495
作者:Alexander Boesgaard Lorup
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Stage-replay diagnostics reconstruct, diagnostics reconstruct intermediate, Stage-replay diagnostics, reconstruct intermediate token, intermediate token prefixes
备注: 15 pages, 1 figure, 6 tables. Reproducibility artifacts (frozen manifests, token IDs, per-item scores, analysis harnesses) described in Section 3.9
点击查看摘要
Abstract:Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that originally reached the prefix. We audit that assumption at a whole reasoning-stage boundary in a Qwen2.5-derived system. A matched 200-item experiment compares retained live cache with one-shot prefill of identical integer tokens and places an exact replica on both sides. In BF16, replicas remain exact while the constructions differ on 166 suffixes and 20 correctness labels; the accuracy difference is only one point (paired 95% CI [-3.5, +5.5]). A fixed-prefix 2x2 holds all 200 token states constant while crossing construction and precision. The BF16 disagreements recur, whereas FP32 produces no decoded disagreement (95% Wilson upper bound 1.88%). A prospective bridge makes token-by-token incremental and retained live caches bit-exact on 12/12 rows; an all-200 saved-ledger audit reproduces every retained trajectory and comparison fingerprint. Bidirectional transplantation of all 48 key/value layers makes every tested divergent continuation follow its cache donor, both on a selected set at the primary checkpoint (24/24) and an outcome-blind replication at a later checkpoint (43/43). Exact-token replay can therefore be repeatable without preserving live-state fidelity. On the tested states, boundary K/V cache is a causally sufficient carrier of the divergent trajectory, while numerical precision moderates its behavioral expression.
16. 【2607.28478】Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning
链接:https://arxiv.org/abs/2607.28478
作者:Zheng Wu,Chenhao Xue,Shijie Zheng,Yijie Lu,Cheng Yang,Zhuosheng Zhang
类目:Computation and Language (cs.CL)
关键词:heavily prioritize explicit, prioritize explicit conditions, explicit conditions provided, complex reasoning tasks, large language models
备注:
点击查看摘要
Abstract:As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90\% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at this https URL.
17. 【2607.28476】Improving Mental Health Screening and Early Risk Detection in Spanish
链接:https://arxiv.org/abs/2607.28476
作者:Andreu Casamayor-Segarra,Vicent Ahuir,Antonio Molina-Marco,Lluís-F. Hurtado
类目:Computation and Language (cs.CL)
关键词:social media posts, analyzing long histories, media posts, difficulty of analyzing, analyzing long
备注:
点击查看摘要
Abstract:Early detection of mental health disorders is often limited by the lack of specialized resources in Spanish and the difficulty of analyzing long histories of social media posts. This paper addresses these challenges through three main contributions. First, we introduce three Spanish foundational models specifically adapted to the mental health domain through domain-specific pre-training. Second, we propose Incremental Context Expansion (ICE), an automatic relabeling methodology designed for early detection. ICE identifies the point at which cumulative messages provide enough evidence of a disorder, generating more informative training samples. Third, we provide a set of fine-tuned models using the samples generated with the ICE methodology for early risk detection tasks. Our results on three Spanish benchmarks show that combining these specialized models with ICE improves the state-of-the-art, reducing detection latency while maintaining high performance. All models are publicly available.
18. 【2607.28457】SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute
链接:https://arxiv.org/abs/2607.28457
作者:Hongyu Chen,Liang Lin,Guangrun Wang
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Scaling test-time computation, budgets waste computation, uniform budgets waste, verifier-guided refinement relies, improve language-model reasoning
备注: 8 pages, 4 figures, 4 tables
点击查看摘要
Abstract:Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refinement relies on external feedback. We introduce Self-Verifying Refinement (SVR), an oracle-free multi-turn reinforcement learning framework that learns to use self-verification as a compute-control policy. At each turn, the model produces a solution together with a discrete correctness verdict and a confidence score; it retains the current answer only when the verdict is Correct and confidence exceeds a threshold, and otherwise continues refinement using its own self-verification. Ground-truth correctness is used only to construct training rewards and is never exposed to the policy through refinement prompts or required at inference. SVR is trained with GRPO on fixed-horizon trajectories using rewards that promote solution correctness, calibration-aware self-verification, and stop-ready correct states; adaptive stopping is activated only at inference. On seven mathematical reasoning benchmarks with Qwen3.5-2B, SVR achieves a macro-average accuracy of 0.563 with only 2.99 inference turns on average. In the evaluated complete-system comparison, it exceeds standard GRPO, strong multi-turn baselines, and a fixed-budget oracle-guided score-feedback reference while requiring substantially fewer turns than fixed ten-turn inference. These results demonstrate that learned self-verification can serve as an effective internal control signal for answer retention and adaptive test-time compute allocation.
19. 【2607.28449】Lightning OPD 2.0: Mitigating Style Bias in Cross-Teacher On-Policy Distillation for Large Reasoning Models
链接:https://arxiv.org/abs/2607.28449
作者:Yecheng Wu,Song Han,Han Cai
类目:Computation and Language (cs.CL)
关键词:providing OPD supervision, Lightning OPD, model providing OPD, dense token-level supervision, On-policy distillation
备注:
点击查看摘要
Abstract:On-policy distillation (OPD) provides dense token-level supervision from a teacher, but its effectiveness can depend on teacher consistency, meaning that the model providing OPD supervision should also have generated the demonstrations used to train the supervised fine-tuning (SFT) reference. However, this condition is frequently violated in practice when SFT data have mixed or unknown provenance or when different models are preferred for SFT data generation and subsequent distillation. In such cross-teacher settings, even a stronger OPD teacher can yield little improvement over the SFT reference. We find that raw teacher--reference disagreement contains potentially useful context-specific teacher evidence as well as a recurring component associated with differences in wording, formatting, and reasoning cadence. We introduce Lightning OPD 2.0 with cross-fitted style residualization, which uses rollout-level cross-fitting to estimate this recurring component as an operational proxy for style-token bias and subtracts it before constructing the token-level OPD update. Across mathematical reasoning and code generation benchmarks, Lightning OPD 2.0 consistently outperforms Lightning OPD in cross-teacher settings. Starting from Klear-Reasoner-8B-SFT, Lightning OPD 2.0 reaches 82.4% on AIME 2024 and 63.0% on LiveCodeBench v5. Together, these results establish Lightning OPD 2.0 as a practical approach to cross-teacher OPD, relaxing teacher consistency as a prerequisite and allowing the SFT data generator and distillation teacher to be selected independently. Code will be released soon.
20. 【2607.28439】Beyond a Single Judge: Simulating Social Persona Panels for Generative UI Evaluation
链接:https://arxiv.org/abs/2607.28439
作者:Zheng Wu,Yibo Luo,Pu Zhang,Cheng Yang,Zhuosheng Zhang
类目:Computation and Language (cs.CL)
关键词:renderable interface directly, large language models, language models synthesize, synthesize a complete, natural-language instruction
备注:
点击查看摘要
Abstract:Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson $r$ from $0.716$ to $0.922$, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at this https URL.
21. 【2607.28434】Metaphor Tracer: A Theory-Informed Analysis of Hidden States
链接:https://arxiv.org/abs/2607.28434
作者:Marc Heimann,Roxana Assadi Moghaddam,Olga Brovkina,Mark Pettifor,Lutz Goetzmann
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:single text, text, token, aggregator, language model hidden
备注: 39 pages, 8 figures
点击查看摘要
Abstract:What do a language model's hidden states say about the organization of a single text? From one forward pass, without training, we score every token position on two properties. The *aggregator* measures whether the position consolidates the whole text into a stable configuration. The *differentiator*, whether other tokens are transiently carried into its subspace as the model reads: metaphor in its root sense, transport. Constants were frozen on one discovery text; every other is confirmatory. The aggregator is not, in the classic sense, an information measure, nor a measure of salience. Across three unrelated models, as a signifier repeats, its surprisal and its attention drain while its aggregator score holds: the channel marks a token's place in the text. That this tracks a reading rests on independent ground truth: an engineered register the aggregator follows across its boundaries (6/6 cells), and a psychoanalyst's marking of clinical transcripts, fixed before the instrument existed, in 34/36 cells, with a graded increment above lexical controls and dissociations no type-level measure reproduces. A transfer test gives the result its shape: the model whose token structure travels with lexical type reads the singular discourse worst, and in a matched base/instruct pair tuning raises fidelity without moving type-transfer. Structural value is a property of a token's place in *this* text, not of its vector alone: a relational rather than essentialist reading of hidden states, operationalizing theory that predated the instrument.
Comments:
39 pages, 8 figures
Subjects:
Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:
arXiv:2607.28434 [cs.AI]
(or
arXiv:2607.28434v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2607.28434
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
22. 【2607.28418】WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning
链接:https://arxiv.org/abs/2607.28418
作者:Haozhe Hu,Hao Wu,Peiran Yin,Chao Han,Yunpu Ma,Xiaoyu Shen
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:efficiency of LLMs, promising approach, approach for improving, improving the efficiency, Pruning
备注: 30 pages, 19 figures
点击查看摘要
Abstract:Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver practical throughput gains, but their input-agnostic computation allocation often causes substantial accuracy degradation under aggressive sparsity. Recent dynamic sparsity methods improve quality retention by adapting computation to individual inputs, yet they remain largely limited to coarse-grained structural decisions and their practical acceleration under real-world inference scenarios remains challenging. To address these challenges, we present WIDE, the first end-to-end differentiable token-level dynamic width pruning framework designed for both prefill and decode scenarios. WIDE enables fine-grained computation allocation by allowing each token to dynamically select attention-head groups and FFN-channel groups, extending dynamic pruning beyond layer-level decisions to neuron-block-level granularity. Through a two-stage training pipeline, WIDE learns effective token-wise sparse execution patterns and achieves substantially better quality retention than existing approaches. To make such fine-grained dynamic pruning practical, we further propose a pruning--kernel co-design framework that decomposes dynamic sparsity acceleration into mask reordering, hardware-agnostic block-level skipping, and hardware-dependent intra-block skipping, enabling efficient execution across different granularities. At 50% sparsity, WIDE provides 55.1% performance boost when compared to the state-of-the-art dynamic depth pruning under calibration-only settings. Under prefill and decoding inference workloads, WIDE achieves close-to-theoretical kernel-level speedups of up to 1.98x for prefill and 4.95x for decoding, as well as 1.68x and 1.55x end-to-end acceleration. Our code is available at this https URL.
23. 【2607.28410】Can Large Language Models Execute Parent Orders?
链接:https://arxiv.org/abs/2607.28410
作者:Zane Shen,Xinli Xu,Guangyi Zhang,Jialong Chen,Jinsong Zhou,Cong Chen,Guibao Shen,Dongyu Yan,Luozhou Wang,Zhen Yang
类目:Computational Engineering, Finance, and Science (cs.CE); Computation and Language (cs.CL); Trading and Market Microstructure (q-fin.TR)
关键词:reducing execution costs, smaller orders, algorithmic trading, core problem, problem in algorithmic
备注:
点击查看摘要
Abstract:Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limits adaptability to new settings. To overcome these limitations, we present the first systematic study of large language models (LLMs) for parent-order execution. This extends the use of LLMs in finance from what to trade to how to execute. We propose PACE (Plan-Ahead Controlled Execution), a hierarchical framework that decomposes parent-order execution into long-horizon planning and short-horizon execution, requiring neither explicit market assumptions nor task-specific training. Experiments on Shenzhen Stock Exchange Level-1 data show that PACE outperforms TWAP, Almgren-Chriss, and learning-based baselines, exceeding the strongest baseline by 0.65 bps. Behavioral analysis reveals that LLMs make execution decisions differently from human investors: higher model confidence predicts better performance rather than worse returns, and the model trades earlier rather than procrastinating toward the deadline. These findings suggest that LLMs can complement human traders in execution decisions.
24. 【2607.28397】GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
链接:https://arxiv.org/abs/2607.28397
作者:Maya Arseven,Anette Frank,Beni Egressy,Johann Higl,Moritz Plenz
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Retrieval-augmented generation, knowledge graphs requires, effectively capture, Retrieval-augmented, semantic information
备注: 10 pages, 19 figures
点击查看摘要
Abstract:Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based retrievers to model graph topology in multi-hop reasoning tasks. In parallel, graph language models (GLMs) have emerged as a promising paradigm that integrates graph reasoning and the semantic capabilities of language models. In this work, we introduce a GLM-based retriever and investigate the comparative strengths of GLM-based, GNN-based, and traditional vector-search-based retrievers in single- and multi-hop RAG settings, and with a particular focus on transferability to unseen domains. Our findings suggest that finetuned GLM retrievers generalize better out of domain, achieving SOTA on two multi-hop benchmarks. On in-domain multi-hop QA datasets they remain comparable to prior work, with promising scaling as parameters and subgraph coverage increase. GNN-based retrievers achieve higher graph coverage with an efficient training setup, whereas the vector-search baseline excels at single-hop datasets.
25. 【2607.28359】Correlation between prosody and pragmatics: A case study of the discourse marker hālā `now' in Persian
链接:https://arxiv.org/abs/2607.28359
作者:Soleiman Ghaderi,Moloud Asakereh,Kevin Tang
类目:Computation and Language (cs.CL)
关键词:exhibits remarkable multifunctionality, Persian discourse marker, exhibits remarkable, encompass a variety, temporal adverbial role
备注: 35 pages, 0 figures
点击查看摘要
Abstract:The Persian discourse marker hālā ('now') exhibits remarkable multifunctionality, extending far beyond its temporal adverbial role to encompass a variety of pragmatic functions. This study presents a pragmatic and acoustic analysis of hālā in spoken Persian, examining 267 instances from spontaneous conversations. While temporal uses were present, they were often combined with other discourse marker functions, indicating extensive multifunctionality, with 70% of tokens serving two or more pragmatic roles. Textual functions (topic shifting, signaling relationships, boundary marking, attention guidance, topic introduction, and topic emphasis) were most frequent, followed by interactive functions (turn management, listener engagement, and feedback regulation), and modal functions (epistemic stance, emotional expression, and attitudinal marking). Prosodic analysis revealed that duration and intensity are key cues for distinguishing hālā's functions. Textual uses were significantly shorter, while temporal uses showed a tendency toward longer realizations. Interactive functions correlated with higher intensity, while modal functions showed a weaker tendency toward lower intensity. These findings indicate that duration and intensity are the main prosodic cues associated with functional differentiation in hālā, especially in textual and interactive uses.
26. 【2607.28347】LLMs struggle to simulate human belief updates in controlled environments
链接:https://arxiv.org/abs/2607.28347
作者:Sebastian Pohl,Harsh Mehta,Pranav Mambayil,Abdul Ghafoor,Franziska Lesigang,Yufang Hou,Christian Hilbe
类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)
关键词:social science experiments, science experiments, tested directly, increasingly deployed, deployed as proxies
备注:
点击查看摘要
Abstract:LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.
27. 【2607.28319】Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
链接:https://arxiv.org/abs/2607.28319
作者:Pere Martra,Eugenio Martínez Cámara,Alfonso Ureña López
类目:Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
关键词:presents Fairness Pruning, work presents Fairness, Fairness Pruning, presents Fairness, large language models
备注: 15 pages, 3 figures, 9 tables. Code and datasets publicly available
点击查看摘要
Abstract:This work presents Fairness Pruning, a lightweight structural intervention method designed for the management and future mitigation of demographic bias in large language models (LLMs). As a foundational empirical validation of this method, this work focuses on causal bias localization. Using minimally contrastive prompt pairs and inference-time activation capture, the method identifies neurons that react differentially when processing demographic attributes in GLU architectures, evaluating the signal at the down_proj input. Empirical evaluation was conducted on models of up to 3 billion parameters (Llama-3.2 family and Salamandra-2B), combining standardized benchmark evaluation with qualitative text generation experiments. Results demonstrate that zeroing the identified neurons alters how the model responds to associated demographic variables. However, rather than producing flat mitigation, the intervention causes bidirectional bias destabilization: because BiasScore is unsigned, candidate sets mix neurons that push toward and against the stereotype, and the net effect on aggregate bias depends on which sign dominates. The intervention is extremely surgical: zeroing at most 40 neurons in Llama-3.2-1B (less than 0.031% of total MLP width) achieves a mean retention of 99.49% in reasoning and general knowledge capabilities. These findings empirically confirm that demographic bias processing and model capabilities operate on dissociable circuits, establishing the methodological foundations for transitioning from blind zeroing toward directional behavior modulation.
28. 【2607.28292】CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance
链接:https://arxiv.org/abs/2607.28292
作者:Anubhav Lakra,Yue Feng
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Machine Learning (cs.LG)
关键词:Large Language Models, Large Language, facts change continuously, corporate facts change, maintaining factual accuracy
备注: 10 pages, 12 figures
点击查看摘要
Abstract:Large Language Models (LLMs) deployed in dynamic financial environments face a critical challenge: maintaining factual accuracy as market conditions, regulations, and corporate facts change continuously. While 4-bit quantization enables efficient deployment, it severely limits the viability of sequential memory editing: existing methods undergo catastrophic performance degradation under this "quantization stability crisis." We introduce CACHE-UK (Contextual Adaptive Continual Hybrid Editor for UK Finance), a stability-aware memory editing framework specifically designed for domain-specific, quantized LLMs. CACHE-UK integrates three components: a rank-1 LoRA perturbation mechanism that confines edits to the low-rank adapter subspace, a financial domain prioritization module for content-adaptive edit strength, and a closed-loop Stability Controller that tracks "degradation debt" to prevent catastrophic forgetting across sequential updates. Evaluated on a 4-bit quantized OpenLLaMA-3B model with a curated UK financial corpus of 88,021 documents, CACHE-UK reduces knowledge degradation by 11-17% relative to adapted baselines under identical 4-bit constraints -- its most robust effect -- while attaining the highest test success (generalization) rate observed in our setting (28%, a 6 percentage point improvement over the strongest adapted baseline). These results indicate that stability-aware editing can improve factual maintenance in resource-constrained financial LLM deployments, though absolute generalization rates remain low.
29. 【2607.28282】(Towards) Scalable Reliable Automated Evaluation with Large Language Models
链接:https://arxiv.org/abs/2607.28282
作者:Bertil Braun,Martin Forell
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Large Language Models, Large Language, Language Models, remains challenging, challenging and resource-intensive
备注: 17 pages. Published in the Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM2) at ACL 2025
点击查看摘要
Abstract:Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated metrics often fail to capture the complexity and variability inherent in LLM-generated outputs. Moreover, these metrics typically rely on explicit reference standards, limiting their use mostly to domains with objective benchmarks. This work introduces a novel evaluation framework designed to approximate expert-level assessments of LLM-generated content. The proposed method employs pairwise comparisons of outputs by multiple LLMs, reducing biases from individual models. An Elo rating system is used to generate stable and interpretable rankings. Adjustable agreement thresholds, from full unanimity to majority voting, allow flexible control over evaluation confidence and coverage. The method's effectiveness is demonstrated through evaluating competency profiles extracted from scientific abstracts. Preliminary results show that automatically derived rankings correlate well with expert judgments, significantly reducing the need for extensive human intervention. By offering a scalable, consistent, and domain-agnostic evaluation layer, the framework supports more efficient and reliable quality assessments of LLM outputs across diverse applications.
30. 【2607.28274】MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek
链接:https://arxiv.org/abs/2607.28274
作者:Ioakeim Perros,Cleopatra Papadopoulou,Ayoub Kirouane,Christos Petrocheilos
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Modern Greek, Greek inflected forms, factual knowledge, richly inflected language, Morphological Open-class
备注: 12 pages, 8 tables
点击查看摘要
Abstract:Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicated to their inflectional competence. We introduce MORFES (Morphological Open-class Recognition-and-Formation Evaluation Suite), a benchmark of 500 expert-verified items that tests the recognition and production of Greek inflected forms, favoring lower-frequency lemmas so that a correct answer reflects the rule rather than a memorized form. We make it publicly available at this https URL. We evaluate a range of open language models on MORFES, situating them within the rapidly scaling open-weight ecosystem from LLaMA to Qwen3, DeepSeek-R1, Magistral, and Kimi K2, where multilingual coverage grows but grammatical competence in morphologically rich languages remains under-measured. Among them, Sophea-Genesis-1, a model we developed and release as open weights at this https URL, leads on inflectional morphology while matching similarly sized models in general capability.
31. 【2607.28263】Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
链接:https://arxiv.org/abs/2607.28263
作者:Hanzuo Liu,Xuan Qi,Chunyu Liu,Haotian Zhong,Yulong Wang,Rayying, Key,Alex Lamb,Mingyu Gao
类目:Computation and Language (cs.CL)
关键词:build semantic representations, middle layers build, layers build semantic, layers increasingly specialize, upper layers increasingly
备注: 19 pages, 4 figures, 27 tables. Submitted to ACL Rolling Review
点击查看摘要
Abstract:Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for prediction. We turn this division of labor into CoMem (Comprehension Memory), which writes each context chunk only through an intermediate layer, retrieves a fixed number of cached residual states, and recomputes the query-conditioned upper layers over the resulting pack. For a fixed retrieval budget, model-side read compute and memory are independent of stored-context length. We evaluate a continued-trained Qwen3-8B base LM under a unified chat-template-free protocol. The backbone is frozen; the flagship trains only a rank-32 self-distillation LoRA on plain PG19, and we report an adapter-free arm separately. CoMem reaches 97.05 on RULER and 38.27 on LoCoMo versus 34.59 for full-context KV-Direct; the dialogue-memory advantage survives conversation-cluster resampling and an independent judge. Results on additional long-context and long-document tasks expose both the benefits of bounded retrieval and its in-window compression tax. Controlled depth sweeps show that deeper caching lowers per-query recomputation but incurs a fidelity loss that self-distillation substantially repairs. In a separate adapter-free efficiency control on an NVIDIA H20 at 128k, CoMem uses 18.26 GB rather than 89.36 GB and achieves a 7.83x prefill speedup. These results show that long-context memory can be organized along the layer axis, not only the token axis.
32. 【2607.28236】CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising
链接:https://arxiv.org/abs/2607.28236
作者:Sina Heydari,Amirreza Abbasi,Mohsen Hooshmand,Majid Ramezani
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Pre-trained language models, Contrastive Denoising Autoencoder, significantly improved sentence, embedding remain sensitive, refines pre-trained BERT
备注: Submitted to 16th International Conference on Computer and Knowledge Engineering (ICCKE 2026)
点击查看摘要
Abstract:Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving textual perturbations such as synonym substitution, masking and word dropout. This work proposes a lightweight Contrastive Denoising Autoencoder (CDAE) that refines pre-trained BERT embedding by jointly optimizing contrastive and reconstruction objective to learn perturbation-invariant representation. We evaluate the proposed framework using multiple perturbation strategies with varying strengths and compare it against the original BERT embeddings and SimCSE. Experimental results show that CDAE consistently preserves higher embedding similarity under perturbations, with the improvements becoming more pronounced as framework effectively enhances representation stability while preserving semantic information, highlighting perturbation-invariant learning as a promising direction for improving sentence embeddings. The source code is publicly available at: this https URL
33. 【2607.28229】EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
链接:https://arxiv.org/abs/2607.28229
作者:Luigi Sigillo,Matteo Silvestri,Francesco Tabaro,Rajat Bhatnagar,Syed Irtaza Mubashar,Matt Jeffryes,Daljit Nijjer,Vittorio Perera,Ola Spjuth,Julio Saez-Rodriguez,Melissa Harrison,Fabio Petroni
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:Europe PMC, increasingly accessed, Europe PMC search, Europe PMC interface, live Europe PMC
备注:
点击查看摘要
Abstract:The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than $16$ points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about $8$ points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at this https URL
34. 【2607.28212】Causal Discovery with Inverted Self-attention for Multivariate Time Series
链接:https://arxiv.org/abs/2607.28212
作者:Yusen Liu,Yong Wang,Yifan Yin,Tianqing Zhu,Xiufeng Liu,Huan Huo
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Causal, high dimensionality, dependencies among variables, data is challenging, challenging due
备注:
点击查看摘要
Abstract:Causal discovery in multivariate time series data is challenging due to complex interactions, high dimensionality, and nonlinear dependencies among variables. Existing methods often struggle to capture these complexities, resulting in inaccurate causal structures. To address this issue, we propose a novel framework that leverages self-attention mechanisms within the transformer architecture for causal discovery. Our approach introduces a novel inverted causal self-attention mechanism (CSAM) that emphasizes latent and indirect causal relationships by inverting tokens and inducing sparsity in attention scores, focusing on significant causal interactions and reducing spurious correlations. Additionally, we develop a global causal algorithm to identify global causal links, providing a holistic metric for causal influence, along with a causal verification module to ensure robustness in the identified causal relationships, enhancing the reliability of our framework. Experiments on both linear and nonlinear datasets, along with ablation studies and sensitivity analyses, show that our framework outperforms existing methods, demonstrating its potential for causal discovery in complex multivariate time series.
35. 【2607.28196】Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
链接:https://arxiv.org/abs/2607.28196
作者:I. Kennedy,T.Kennedy
类目:Computation and Language (cs.CL)
关键词:original network internal, network internal representations, data-cheap quality guards, Practitioners accept, random probe inputs
备注:
点击查看摘要
Abstract:Practitioners accept a compressed language model once it clears a stack of data-cheap quality guards: perplexity within a small factor of the original, downstream accuracy (for example MMLU) inside a confidence interval, and data-free output-fidelity signals that compare the compressed and original network's internal representations under random probe inputs. This stack has a blind spot. Across three model families, gently-compressed models clear every guard and then invent procedure steps that were never in the instructions when they run a standard operating procedure (SOP) as an agent. The effect is operator-specific: coherent low-rank (SVD) truncation induces it, and magnitude pruning matched to the same perplexity does not. One dissociation isolates the cause. The same compressed weights that CI-win a paired output-fidelity test CI-fail the invented-step canary. The governing axis is the coherence of the compression error times its rate; the magnitude of the damage does not predict it. The data-free fidelity probe is a fidelity oracle by construction, so it cannot see this axis. We characterize the blindspot and dissociation with paired confidence intervals on a pre-registered, powered canary across three architectures. Operator-specificity replicates on all three, and the perplexity-guard evasion appears where the model admits in-guard low-rank headroom. We then give a data-free screen: a two-axis statistic of the compression error (coherent-fraction and error-rate) that flags the failing builds with fixed thresholds across architectures and matches the coherence-times-rate mechanism. Perplexity, MMLU, and fidelity acceptance do not certify agent safety. Screen gently-compressed low-rank builds before agentic deployment
36. 【2607.28190】he MADRS Pipeline: Supporting Depression Assessment in Clinical Trials
链接:https://arxiv.org/abs/2607.28190
作者:Mila Fodor,Katalin Ócsai,Francesco Periti,Rien Sonck,Alex Boudreau
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:major mental disorder, diagnosis relies primarily, major mental, diagnosis relies, clinical trials
备注:
点击查看摘要
Abstract:Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the psychiatric MADRS scale are getting more and more attention. While existing solutions primarily focus on detecting the disorder from different text sources (e.g., online text, social media), there is still limited support for clinical trials, where clinical assessments are conducted through structured interviews based on standard guidelines such as SIGMA. In this work, we develop a LLM pipeline specifically designed to support clinicians in supporting the assessment of depression in patients enrolled in clinical trials. Our pipeline converts audio interviews into transcripts, maps them into the ten MADRS symptom items, estimates their severity, and identify problematic clinical ratings associated with them. Evaluation on real clinical interviews shows a strong overall correlation of 0.867 with expert ratings, providing interpretable support for future assessments in clinical trials.
37. 【2607.28166】Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models
链接:https://arxiv.org/abs/2607.28166
作者:Chia-Ming Lee,Ming-Ching Chang,Xin Li,Yu-Lun Liu,Chih-Chung Hsu
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Diffusion language models, Diffusion language, generation-time early exit, language models, expose a provisional
备注: Code is available at [this https URL](https://github.com/ming053l/LATCH-dLLM)
点击查看摘要
Abstract:Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that stops decoding before the schedule is exhausted. Existing early-exit gates decide termination from fixed-region confidence statistics or schedule-dependent rules, evidence too coarse for a decision that freezes every remaining position at once, so they fire prematurely on long chain-of-thought outputs whose answers stabilize only near the end. Adaptive sampling, the other axis of training-free acceleration, paces how quickly positions commit while decoding continues but never verifies that the output itself has stabilized. We introduce a training-free, candidate-aware early-exit framework that keeps the two axes separate and matches each decision to evidence of its own scope. Confidence-Verified Commit (CVC) governs when the sequence may stop by verifying confidence and sustained argmax stability over the dynamically extracted candidate span using a deterministic parser specified from each task's output format. Block-Wise Early Commit (BWEC) governs where to accelerate by applying a cheaper local rule to non-final blocks, while leaving the final block and global termination under CVC. We refer to their combination as LATCH (Localized Acceleration with Tracked-Candidate Halting). Unlike prior methods, LATCH needs no suffix-prompt construction; it is prompt-anchor-free but format-aware. We evaluate LATCH end to end on 11 tasks under zero-shot settings using LLaDA and Dream. LATCH stays within 2.0 percentage points of full-decoding accuracy across all 22 evaluation settings, with one frozen hyperparameter set that transfers cross-backbone untuned, while achieving end-to-end TPS speedups of 9.3-17.8x on short-answer tasks and 2.0-3.3x on long-reasoning tasks.
38. 【2607.28156】RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
链接:https://arxiv.org/abs/2607.28156
作者:Jingxiang Fan,Junbao Zhuo,Bochao Zou
类目:Computation and Language (cs.CL)
关键词:Existing multimodal long-term, memory, overcome the limited, limited context, Reflective Retrieval Memory
备注:
点击查看摘要
Abstract:Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasize what to store rather than how stored memory should be retrieved. When retrieval becomes inaccurate or repeatedly fails to obtain useful evidence, existing agents lack mechanisms to diagnose failures from previous task trajectories and adapt future search this http URL introduce Reflective Retrieval Memory (RRM), a reflective memory framework for long-horizon multimodal reasoning. RRM augments an entity-centric multimodal memory graph with reflective experience memory, which distills transferable procedural retrieval knowledge from historical task trajectories. Unlike episodic and semantic memories that preserve factual evidence from the current video, reflective experience memory captures reusable search strategies across tasks. RRM converts retrieved experiences into query-level guidance, while answer generation remains conditioned only on factual evidence newly retrieved from the current video. A lifecycle management mechanism further regulates experience memory through usage frequency, reuse feedback, and temporal decay, thereby reducing redundancy and noise. RRM consistently outperforms previous state-of-the-art approaches on M3-Bench-Robot, M3-Bench-Web, and Video-MME-Long, demonstrating the effectiveness of reflective retrieval memory for long-horizon multimodal reasoning.
39. 【2607.28146】Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game
链接:https://arxiv.org/abs/2607.28146
作者:Niklas Bauer,Lars Benedikt Kaesberg,Akiko Aizawa,Jan Philip Wahle,Bela Gipp,Terry Ruas
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:high-stakes settings, legal systems, fundamental to safety, deployed as agents, agents in high-stakes
备注:
点击查看摘要
Abstract:As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.
40. 【2607.28128】Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
链接:https://arxiv.org/abs/2607.28128
作者:Shuyi Fan,Boyuan Deng,Mengyu Xu,Jiale Liu,Hongyang Zhang,Qiaoxin Yang,Chongyang Gao
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
关键词:LLM tutoring poses, distinguish direct answer-giving, LLM tutoring, rubric distinguish direct, measurement problem
备注: 24 pages, 4 figures, 6 tables
点击查看摘要
Abstract:LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|\delta|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.
41. 【2607.28127】FinSMART: Financial Sentiment Analysis for Algorithmic Trading through Market-Aligned Reinforcement Learning
链接:https://arxiv.org/abs/2607.28127
作者:Giorgos Iacovides,Wuyang Zhou,Danilo Mandic
类目:Computation and Language (cs.CL); Machine Learning (cs.LG); Statistical Finance (q-fin.ST); Trading and Market Microstructure (q-fin.TR)
关键词:Recent advances, advances in Generative, substantially improved financial, post-trained financial large, financial large language
备注:
点击查看摘要
Abstract:Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs). However, existing approaches remain confined to a market-agnostic, supervised learning paradigm that relies on limited, static and human-annotated datasets, and thus are incapable of adapting to evolving market conditions. To address this limitation, we introduce FinSMART, the first market-aligned reinforcement learning framework for financial sentiment analysis, which directly optimizes sentiment signals using realized market outcomes. To deal with the noisy, non-stationary, and multifactorial nature of financial markets, FinSMART incorporates a signal extraction pipeline that combines market-aware data filtering with a discrete asymmetric trading reward, enabling stable reinforcement learning from economically meaningful market feedback. Experimental results demonstrate that FinSMART significantly outperforms existing state-of-the-art methods in profitability, risk-adjusted performance, and sentiment signal quality, improving cumulative trading returns by 220% over the strongest baseline. Uniquely, the FinSMART framework naturally supports market-aware retraining, at any point in time, by replacing costly manual annotation with newly observed financial articles and their realized market outcomes. Such a retraining strategy enables the model to continuously adapt to changing market dynamics, resulting in consistent performance gains over its static counterpart. These findings demonstrate the practical applicability of market-aligned reinforcement learning and highlight its potential as a next-generation paradigm for developing adaptive financial LLMs.
42. 【2607.28119】Challenges in annotations by humans and LLMs: A case study of evaluative language
链接:https://arxiv.org/abs/2607.28119
作者:Mirela Imamovic,Aenne Cecilia Kristine Knierim,Khushi Pitroda,Ekaterina Lapshinova-Koltunski
类目:Computation and Language (cs.CL); Social and Information Networks (cs.SI)
关键词:complex linguistic phenomena, English TED talk, large language models, draw a comparison, generated by large
备注:
点击查看摘要
Abstract:In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
43. 【2607.28100】PCAP-LM: An LLM-Native Text Representation for TLS Bulk Traffic Analysis
链接:https://arxiv.org/abs/2607.28100
作者:Xavier Marjou,Lucas Tamic,Ilan Jaffeux-Cheniout
类目:Networking and Internet Architecture (cs.NI); Computation and Language (cs.CL)
关键词:Large language models, offer powerful reasoning, powerful reasoning capabilities, network traffic analysis, standard capture formats
备注: 6 pages
点击查看摘要
Abstract:Large language models (LLMs) offer powerful reasoning capabilities for network traffic analysis, but standard capture formats and their textual equivalents are prohibitively verbose, overflowing LLM context windows by two orders of magnitude. We present PCAP-LM, a flow-centric, LLM-native text representation that acts as a lossy knowledge extraction step rather than a standard compression tool: raw captures are transcoded into semantic summaries using PacketGlyphs - a novel ASCII alphabet coined in this paper that encodes packet direction, TCP/TLS state, log-scale size, and inter-packet delay. Combined with a constrained PMI-BPE tokenizer and motif run-length encoding, repetitive behavioural patterns are aggressively collapsed. A @REFS side-index preserves lossless drill-down into the original packets. Evaluated on a homogeneous corpus of 5G/4G TLS 1.3 bulk-download traffic, the BPE vocabulary fully saturates at 159 tokens, achieving an 812x size reduction over tshark -V and fitting entire captures within a single LLM context window. In a forensic question-answering evaluation over 30 held-out files, a frontier LLM achieves 99.3% accuracy from PCAP-LM documents versus 51.0% from a token-budget-matched tshark -V prefix. The lossy design introduces known blind spots - most notably a 24% false-negative rate for TCP retransmissions - and extending to heterogeneous mixed-protocol environments will require vocabulary retraining.
44. 【2607.28098】SciDataSailor: Deep Scientific Data Exploring
链接:https://arxiv.org/abs/2607.28098
作者:Jiyong Rao,Yicheng Qiu,Chi Zhang,Chunfeng Song,Runkai Zhao
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Deep Scientific Data, making their inspection, Scientific Data Exploration, domain expertise, real scientific data
备注: 63 pages, 10 figures
点击查看摘要
Abstract:Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments. We introduce Deep Scientific Data Exploration, an agentic task paradigm in which agents navigate repositories, interpret heterogeneous files and schemas, execute analyses, integrate cross-file evidence, and produce conclusions grounded in executed observations. To operationalize this paradigm, we present SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation. SciDataSailor instantiates trajectory synthesis as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms: difficulty-stratified exploration seeds, dual-feedback first-play urgency, hierarchical strategy-to-tool action generation, and entropy-guided branching. Using this framework, we construct SciDataSailor-SFT-2K for supervised fine-tuning and SciDataSailor-Bench for evaluation, with the latter comprising 627 meta-information summarization tasks and 586 scientific question-answering tasks across 27 datasets spanning the life, earth, and physical sciences.
45. 【2607.28082】GGC: Selective Query Correction for Reliable Text-to-SPARQL Generation
链接:https://arxiv.org/abs/2607.28082
作者:Ziyi Yang,Thanh-Son Nguyen,Tuan Anh Nguyen,Lihui Chen
类目:Computation and Language (cs.CL)
关键词:Large language models, translates natural language, demonstrated strong capabilities, Large language, natural language questions
备注: 18 pages, 1 figure
点击查看摘要
Abstract:Large language models (LLMs) have demonstrated strong capabilities in structured query generation, making them a natural choice for Text-to-SPARQL, which translates natural language questions into executable SPARQL queries over knowledge graphs. However, their initial outputs remain unreliable: generated queries may be executable yet semantically misaligned with input questions, leading to incorrect retrieval. To address this issue, we propose Generator-Gate-Corrector (GGC), a framework for reliable LLM-based Text-to-SPARQL generation. GGC first uses a Generator to produce an initial query, then applies a Gate to predict whether correction is needed, and finally invokes a Corrector only for selected high-risk queries. This selective correction mechanism avoids unnecessary modifications and reduces the risk of degrading originally correct queries. Experiments on MCQA show that GGC improves query-level accuracy from 90.23\% to 98.33\% while reducing inference overhead by 45\% compared with correcting all generated queries. Ablation studies show that the Gate is robust across thresholds and that Corrector training data composition affects correction effectiveness and stability. Overall, the results demonstrate that selective correction enhances the accuracy, reliability, and efficiency of LLM-based text-to-SPARQL generation.
46. 【2607.28077】LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
链接:https://arxiv.org/abs/2607.28077
作者:Shuang Liang,Haoyang Zhou,Yifan Gong,Guowei Wang,Xiting Wang
类目:Computation and Language (cs.CL)
关键词:effective learning signals, Reinforcement learning, large language models, learning signals, effective learning
备注: 15pages
点击查看摘要
Abstract:Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation. However, existing pre-rollout methods struggle to balance exploitation and exploration: repeatedly exploiting historically informative prompts can narrow training coverage, whereas broader exploration can lower the fraction of informative prompts. To address these limitations, we introduce LEEPS, a Latent-Guided Explore--Exploit Prompt Sampler that adaptively balances the reuse of previously observed informative prompts with continued exploration of uncertain ones. LEEPS partitions candidates into exploit and explore portfolios and adaptively allocates rollout budget according to their recent non-trivial ratios. It further uses representation-space neighbors and historical rollout outcomes to prioritize uncertain prompts likely to yield non-zero reward variance, thereby making exploration more targeted without additional rollouts. Across six mathematical reasoning benchmarks, LEEPS achieves the highest average score at both model scales, with relative gains of 2.6\% and 3.7\% over the strongest baseline for Qwen2.5-Math-1.5B and 7B, respectively, and generally improves faster during the training process. It also achieves the highest average score across the three evaluated OOD general-reasoning benchmarks at both model scales and adds only about 2 seconds of online sampling overhead per training step. Code is available at this https URL.
47. 【2607.28008】RepBench: Compiling Benchmarks into Capability Representations for Large Language Models
链接:https://arxiv.org/abs/2607.28008
作者:Yanshi Li,Xueru Bai,Shuman Liu,Long Zhang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Representation engineering reads, paper-specific synthetic data, steers capability directions, large language models, engineering reads
备注: 22 pages, 8 figures, with appendices. Yanshi Li and Xueru Bai contributed equally
点击查看摘要
Abstract:Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthetic data. The resulting measurements are difficult to compare or reproduce and may reflect surface patterns rather than capabilities. We present RepBench, a benchmark-grounded data layer for capability-aligned representation probing. Crawling 13,427 benchmark papers yields a taxonomy of 182 capability clusters in 13 families; harvesting 353 public benchmark datasets yields 46,149 audited probe texts covering 94 capabilities, each supported by at least two independent benchmarks. This multi-benchmark design reduces dependence on any single source: raw per-text vectors exhibit no natural cluster granularity, whereas benchmark-pooled capability vectors show an interior clustering optimum at a small number of clusters on all 12 evaluated models, with low agreement to the human taxonomy. Under cross-benchmark transfer evaluation across twelve models completed by all four readouts, difference-in-means attains the highest model-level mean on ten models, while logistic regression wins the most capability-model cells. This disagreement shows that the readout method and aggregation criterion are meaningful evaluation dimensions. The pipeline, corpus, and evaluation code are released as a reusable closed-loop workflow.
48. 【2607.27955】SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
链接:https://arxiv.org/abs/2607.27955
作者:Jennifer D'Souza,Sameer Sadruddin,Anisa Rula,Ana Bossler,Andrés Fullana,Enric Bas,Syed Ather,Defne Circi,Anlan Chen,L. Catherine Brinson,Alyssa Columbus,George Demetriou,Dongjun Jeong,Tarun Kumar,Frank Krüger,Sascha Genehr,Kai Budde-Sagert,Anamaria Leonescu,Francesco Lodola,Chiara Florindi,Gagana Balasubramanya Murthy,Samson Oluwapelumi Olagbile,Nazia Riasat,Yan Sha,Kevin Shen,Shaokai Yang
类目:Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:heterogeneous article discourse, spanning Biology Biotechnology, dispersed across prose, supplementary files, details needed
备注: 25 pages, 9 figures, Submitted for peer review to Nature Scientific Data
点击查看摘要
Abstract:Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of this http URL, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology Biotechnology, Materials Chemistry, Imaging Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.
49. 【2607.27940】riShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement
链接:https://arxiv.org/abs/2607.27940
作者:Cheng Wei(Honor Device Co., Ltd., Shenzhen, China)
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:exposing raw data, Federated fine-tuning, enables collaborative training, large language models, textbf
备注: 12 pages,3 figures
点击查看摘要
Abstract:Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint [1] (arXiv:2606.20553), demonstrates that a malicious parameter server can corrupt a PEFT adapter into a privacy backdoor: by assigning a dedicated memorization neuron to each training sample and ensuring each neuron updates at most once, the server can analytically reconstruct 59\%--79\% of client training data with high semantic fidelity. Existing defenses---including local differential privacy (LDP) [8] and gradient clipping---either fail against this attack or impose unacceptable utility degradation. We present \textbf{TriShield}, a three-layer deterministic defense that completely prevents NeuroImprint-style reconstruction with \textbf{zero model utility loss} and \textbf{no additional communication rounds}. TriShield consists of: (1) a \textbf{Parameter Artifact Detector} that identifies memory-neuron signatures in distributed model parameters before local training begins; (2) a \textbf{Stateful Virtual Iteration} mechanism that forces Adam/AdamW's momentum state to irreversibly entangle gradients across virtual steps, invalidating NeuroImprint's closed-form inversion; and (3) a \textbf{Zero-Utility Orthogonal Projection} operator that projects all local gradient updates onto the main-task semantic subspace computed via SVD, physically eliminating any gradient components that carry private memorization. We prove theoretically that after Layers 2 and 3, the mutual information between the uploaded gradient and any individual training sample is zero. Experiments on GPT-2 (117M) and Llama-Guard-3-1B verify that TriShield reduces NeuroImprint reconstruction rate to \textbf{0\%} across all tested attack variants, while maintaining or improving training accuracy, with less than 5\% additional GPU computation overhead.
50. 【2607.27919】Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
链接:https://arxiv.org/abs/2607.27919
作者:Rubin Wei,Jiaqi Cao,Jiarui Wang,Junming Zhang,Qipeng Guo,Bowen Zhou,Zhouhan Lin
类目:Computation and Language (cs.CL)
关键词:Decoder-only language models, single parameter set, Decoder-only language, entangle long-term memory, making it difficult
备注:
点击查看摘要
Abstract:Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory Decoder at Scale, scaling memory models up to 6.9B parameters and pretraining them on 300B tokens. At this data scale, the combined cost of indexing and search makes a standard Faiss pipeline infeasible. We address this bottleneck with a distributed pipeline for Faiss indexing and retrieval, together with sparse, batch-wise loading of kNN distributions. Across model scales, we find that allocating more parameters to memory yields a better parameter-performance tradeoff than scaling the base model alone. On 17 benchmarks, pairing a 6.9B general memory with Pythia-410M raises its average score from 29.86 to 37.34, surpassing Pythia-12B (37.24) with 39% fewer total parameters. For Qwen3 Base models ranging from 0.6B to 14B, 1.7B domain memories improve the average score across the three domains by more than 9 points at every scale. Overall, our results demonstrate that independently scaling pretrained memory offers a more parameter efficient path to improving language model performance.
51. 【2607.27912】IFHierBench: Hierarchical Instruction Following for Large Language Models
链接:https://arxiv.org/abs/2607.27912
作者:Yuetian Mao,Chunyang Chen
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:deploying large language, downstream components depend, output satisfying specific, large language models, satisfying specific constraints
备注:
点击查看摘要
Abstract:Instruction-following ability is critical for deploying large language models in real-world applications, where downstream components depend on the output satisfying specific constraints. Modern deployments increasingly handle the full task in a single LLM call, with one prompt specifying a layered output whose overall artifact, structural sections, and nested fields must each satisfy concrete constraints. Existing instruction-following benchmarks treat the constraint set as a flat list applied uniformly to the response, so they cannot scope a check to a particular section of the output. We introduce IFHierBench, a hierarchical instruction-following benchmark of 600 prompts stratified across four constraint-tree depths and 35 distinct constraints, each prompt paired with a deterministic checker that verifies satisfaction at every scope. Evaluating seven leading proprietary and open-weight models, we find that even the strongest model only marginally exceeds 50% prompt-level accuracy and that accuracy degrades sharply as constraint depth grows. Reliably following nested constraints remains a substantial gap for current LLMs, motivating future training methods that consider constraint adherence at finer granularity to achieve better instruction-following ability.
52. 【2607.27853】FinanceHarness: Autonomous Financial Deep Research Framework
链接:https://arxiv.org/abs/2607.27853
作者:Yijia Xiao,Rujun Han,Yanfei Chen,Zifeng Wang,Ke Jiang,Zhongying CuiZhu,Vishy Tirumalashetty,Wei Wang,Burak Gokturk,Tomas Pfister,Chen-Yu Lee
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Finance (q-fin.CP)
关键词:adopted agentic products, widely adopted agentic, Powered by advances, financial deep research, deep research
备注:
点击查看摘要
Abstract:Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at this https URL.
53. 【2607.27851】Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm
链接:https://arxiv.org/abs/2607.27851
作者:Ming Wang,Jiaqi Wu Young,Wenfang Wu,Daling Wang,Shi Feng
类目:Computation and Language (cs.CL); Human-Computer Interaction (cs.HC); Social and Information Networks (cs.SI)
关键词:influential strategy traditions, dialogue research includes, includes two influential, Emotional dialogue, Emotional dialogue research
备注:
点击查看摘要
Abstract:Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users' capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. We propose capability-sustaining emotional dialogue (CSED) as a longitudinal research paradigm that aligns supportive strategy with this goal and organizes data, models, system design, evaluation, and governance around repeated use, non-use, transition, and termination. A targeted literature-and-corpus audit motivates this position. In a PRISMA-ScR-guided sample, 95% of 60 system-building papers pursue relief-oriented goals. None evaluates capability or longitudinal outcomes, and only 1 considers dependency, autonomy, or termination risk. In 300 ESConv supporter turns, capability-relevant functions appear in 43.0%, while generic suggestions account for 22.0%, compared with 4.0% reappraisal, 6.7% self-efficacy support, and 0.3% boundary behavior. We release a protocol for extending the audit to model behavior. An illustrative process model connects latent user capability to six design commitments, four evaluation timescales, and lifecycle constraints. The resulting agenda makes CSED testable across data, policy design, training, evaluation, and governance.
54. 【2607.27845】AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
链接:https://arxiv.org/abs/2607.27845
作者:Haobo Li,Eunseo Jung,Wenxiao Zhao,Feng Liu,Jiong Wang,Kaiyi Xu,Zijie Guo,Zixin Chen,Ben Fei,Fenghua Ling,Lei Bai
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Recent advances, assist scientific research, large language models, advances in large, large language
备注:
点击查看摘要
Abstract:Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted scientific workflows remains underexplored: verifying whether reviewer feedback leads to meaningful and evidence-supported manuscript improvements. We introduce AutoSupervision, which evaluates whether scientific manuscript revisions genuinely address reviewer concerns through grounded evidence. AutoSupervision leverages transparent peer-review records as a natural source of supervision, where reviewer comments specify scientific concerns, author responses describe claimed resolutions, and revised manuscripts provide evidence of changes. Given reviewer comments, author responses, and revised manuscripts, models must characterize reviewer concerns, determine whether concerns have been addressed, and identify supporting manuscript evidence. We construct AutoSupervision from 56,000 Nature Communications articles and corresponding review records. Then we conducted experiments on LLMs, the ablation study, and the case study. Our results show that while LLMs perform well in characterizing reviewer concerns, with GPT-5.5 achieving a score of 0.754, evidence-based verification remains the primary bottleneck, with the best-performing model reaching only 0.501.
55. 【2607.27834】MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
链接:https://arxiv.org/abs/2607.27834
作者:Hanshuai Cui,Zhiqing Tang,Zhi Yao,Fanshuai Meng,Qianli Ma,Weijia Jia
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:long-running large language, agents reuse information, large language model, language model agents, model agents reuse
备注:
点击查看摘要
Abstract:Persistent memory lets long-running large language model agents reuse information across sessions and tasks. Yet errors in writable memory can persist and corrupt future behavior. Existing systems improve storage and retrieval, but they do not provide a transaction boundary for reliable updates and recovery. We therefore propose MemTxn, a governance layer outside the answer model. MemTxn verifies whether an update is supported by its source. It also selects the visible version when facts conflict and restores the application-visible state after a fault. The system uses Ordered PatchTest to validate writes, a Temporal Resolver to select versions, and a durable snapshot journal to recover state. On an item-disjoint audit, MemTxn accepts all 60 supported originals and rejects all 179 hard negatives. Under persistent multi-key faults on LongMemEval-S and LoCoMo states, it restores the complete declared active map without knowing the actual physical write set. On MemoryAgentBench FactConsolidation, MemTxn achieves the highest average F1 across all twelve answer-model configurations. It outperforms Dense by 17.06--24.07 points in five representative settings.
56. 【2607.27816】Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
链接:https://arxiv.org/abs/2607.27816
作者:Yuhang Zhu(1),Mingxuan Du(1),Benfeng Xu(1 and 2),Jie Gao(2),Lingyun Yu(1),Hongtao Xie(1) ((1) University of Science and Technology of China, (2) MetaStone Technology, Beijing, China)
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large language models, important consumer applications, language models, important consumer, consumer applications
备注: 29 pages, 3 figures, including supplementary material. Resources: [this https URL](https://github.com/Zhuyh1139/PALATE)
点击查看摘要
Abstract:Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.
57. 【2607.27790】Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis
链接:https://arxiv.org/abs/2607.27790
作者:Wei Chen,Junkai Li,Tongguan Wang,Hui Liu,Feiyue Xue,Chuanxiang Ma,Ying Sha
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Multimodal Sentiment Analysis, integrating natural language, complex human emotions, Multimodal Sentiment, Sentiment Analysis
备注: Accepted by MM 2026
点击查看摘要
Abstract:Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal modalities share a structural isomorphism with natural language, as both can be viewed as feature sequences evolving over time. This isomorphism enables the transformation of non-verbal modalities into text-like tokens for unified semantic reasoning. Large Language Models (LLMs), designed to understand and generate sequential data, can thus be utilized to interpret complex affective sequences. However, existing LLM-based methods primarily capture low-level superficial features, failing to model affective semantics arising from structural variations and contextual interactions. To address this limitation, we propose \textbf{SentiLLM}, a unified framework that leverages \textit{Semantic-Aligned Structural Abstraction} to distill continuous raw signals into compact, semantically meaningful tokens. Specifically, we introduce a \textit{Dual-Stream Salience-Context Calibration Mechanism}, which disentangles non-verbal feature sequences into a focus stream and an ambient stream. The focus stream captures salient sentiment shifts (e.g., facial expressions) guided by textual priors, while the ambient stream characterizes stable background states. Through calibrating these dynamic sentiment shifts against background states, SentiLLM effectively projects non-verbal modalities into a unified semantic space, making them naturally understandable for LLMs. Serving as a plug-and-play module, SentiLLM significantly improves discriminative performance with only a small number of trainable parameters. Our method achieves superior performance on four datasets, MOSI, MOSEI, CH-SIMS, and CH-SIMS v2, demonstrating the effectiveness of the structural abstraction paradigm in MSA. Our code is available at: \href{this https URL}.
58. 【2607.27783】Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
链接:https://arxiv.org/abs/2607.27783
作者:Amruta Parulekar,Jinu Lee,Dilek Hakkani-Tür,Hari Sundaram
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:Large Language Models, Large Language, Language Models, unstructured prose, Directed Acyclic Graphs
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return "Consensus Reasoning". Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman $\rho = 0.30$-$0.51$, and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4-65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem.
59. 【2607.27773】ChronoMem: Version Control and Semantic Rollback for Large Language Model Agent Memory
链接:https://arxiv.org/abs/2607.27773
作者:Yongye Su,Wujiang Xu,Chaoji Zuo,Elisa Bertino
类目:Computation and Language (cs.CL)
关键词:agents increasingly rely, support multi-session interaction, LLM agents increasingly, interaction and personalization, increasingly rely
备注:
点击查看摘要
Abstract:LLM agents increasingly rely on long-term memory to support multi-session interaction and personalization. However, existing agent memory systems are designed around forward-only evolution, continuously accumulating, consolidating, and overwriting knowledge, with no principled mechanism to inspect, version, or revert prior states. This makes agents brittle under corrections, concept drift, and memory corruption, particularly after they have already been exposed to subsequent information. We present ChronoMem, a semantic version-control layer for agentic memory integrated into the production-ready, open-source Agent Development Kit by Google. ChronoMem commits whole-memory snapshots at each memory write, maintains structured version histories, and supports natural-language rollback requests by mapping undo intents to concrete historical versions through hybrid lexical and semantic retrieval, rank fusion, and reranking. We further introduce a post-exposure evaluation protocol that tests whether an agent can behave counterfactually after rollback by answering queries and summarizing history as if future updates had never occurred. On long-horizon conversational benchmarks augmented with evolving memory states and rollback tasks, ChronoMem substantially improves rollback-consistent question answering and history summarization relative to prompt-only and retrieval-only baselines, while achieving strong performance in semantic version selection. To our knowledge, ChronoMem is the first open-source system and benchmark for systematic semantic global memory rollback in LLM agents.
60. 【2607.27766】Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
链接:https://arxiv.org/abs/2607.27766
作者:Xinyu Luo,Hui Liu,Yihua Shao,Junyi Yang,Arindam Basu,Haoliang Li
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:On-device in-context learning, downstream model inference, On-device in-context, in-context learning, relies on pre-inference
备注: Under review
点击查看摘要
Abstract:On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local memories under limited computation, memory, and data-exposure budgets. We propose Conditional Retrieval Alignment (CoRA), a gradient-free framework that converts a frozen encoder into a task-conditioned retriever using paired candidate inputs and outputs. CoRA selects complementary encoder layers, constructs an output-derived conditioning space from candidate memory, and aligns candidate input representations to this space through closed-form ridge regression. Low-rank factorization then produces a compact retrieval basis where candidate outputs are used only during offline index construction, whereas query-time retrieval requires only the query input and precomputed index. We show that CoRA's rank-constrained basis is the optimal low-rank compression of the output-conditioned fitted representation, and derive an exact two-pass streaming construction that avoids materializing the full fitted matrix. We further extend the framework to multimodal exemplar retrieval by incorporating visual representations into the conditioning and retrieval spaces. Experiments across ten textual datasets and four multimodal benchmarks with Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B, and Qwen3.5-2B, as well as end-to-end Raspberry Pi~5 deployment demonstrate that CoRA supports effective task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls.
61. 【2607.27756】Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
链接:https://arxiv.org/abs/2607.27756
作者:Xilin Jiang,Riki Shimizu,Sukru Samet Dindar,Junkai Wu,Zhongweiyang Xu,Nima Mesgarani
类目:ound (cs.SD); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
关键词:designed for clean, dyadic interactions, turns speaking, typically designed, single user
备注:
点击查看摘要
Abstract:Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant's behavior with three action tokens: |respond|, |listen|, and |ignore|, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in |respond| mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.
62. 【2607.27747】Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
链接:https://arxiv.org/abs/2607.27747
作者:Liangjie Zhao,Jiaqing Lyu,Kexin Tang,Zecheng Fang,Rong Yin,Yulan Hu,Da Li,Jianing Li
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Large Vision Language, Vision Language Models, Large Vision, Vision Language, Language Models
备注:
点击查看摘要
Abstract:Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation.
63. 【2607.27739】Measuring Alignment With Reader Highlights Net of Position and Length
链接:https://arxiv.org/abs/2607.27739
作者:Kazuki Nakayashiki,Keisuke Watanabe
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:Context compression discards, language model reads, downstream task accuracy, Context compression, language model
备注: 15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper's numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A
点击查看摘要
Abstract:Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting offers a non-circular reference: many people independently marking passages on the same page. But the obvious metric, the fraction of crowd-marked sentences a compressor keeps, is confounded twice: crowd marks are front-loaded and crowd-marked sentences are longer, so any method favouring early or long sentences scores well regardless of readers. We remove both by matching each marked sentence against unmarked sentences of the same document at equal relative depth and equal within-document length rank, and we calibrate every estimator on synthetic nulls built from position and length alone - a step that matters, since depth-only stratification returns a false positive on 20-36% of nulls containing no effect. On 120 web documents (at least 12 independent readers each), a language-model importance ranking keeps 38.4% of crowd-marked sentences against 19.9% of their matched neighbours: an enrichment of +0.196 [+0.148, +0.239], at p = 0.0005 under an exact randomization test that assumes nothing about clustering, and replicated cross-vendor. Naive truncation, whose keep rule is position, correctly falls to +0.003. To give the number a scale: scored identically, on the same budget, against a crowd label recomputed to exclude them, a single human reader reaches +0.182 - indistinguishable from GPT-5.4 (+0.002 [-0.081, +0.088]) and below Claude Opus 5. Classical methods are not null - Luhn's 1958 heuristic reaches +0.088 - so reader selection is partly recoverable by counting words; conditioning additionally on lexical centrality removes only 0.010, so the agreement is not centrality. We also report that a claim in our own prior work does not reproduce on this corpus.
64. 【2607.27735】A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
链接:https://arxiv.org/abs/2607.27735
作者:Yuesong Liu,Yuan Zeng,Min Lyu,Ruilin Liu,Yu Guo,Yinlong Xu
类目:Computation and Language (cs.CL)
关键词:Speculative decoding alleviates, large language model, language model inference, Speculative decoding, alleviates the memory-bandwidth
备注: 9 pages, 4 figures, subbmited to AAAI 2027
点击查看摘要
Abstract:Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafting overhead, token acceptance, and speculation length. We present a unified efficiency analysis showing that extending the speculation horizon can reduce rather than improve speedup when the marginal acceptance probability falls below the relative drafting cost. Guided by this analysis, we introduce SparseSpec-L, a training-free self-speculative decoding framework for long-context inference. SparseSpec-L generates lightweight drafts directly from the target model using a dynamically sparsified and recallable KV cache. It recycles per-head attention statistics produced during full-context verification as a no-extra-forward importance signal, allowing critical historical tokens to be recalled without permanently discarding the dense KV cache. An online entropy-based controller further selects the speculation length according to expected step-wise efficiency. Experiments across multiple long-context tasks and model scales show consistent end-to-end acceleration, with up to speedup over autoregressive decoding while preserving the target model's output distribution.
65. 【2607.27726】Baikal: Structured Search for Deep Research over Data Lakes
链接:https://arxiv.org/abs/2607.27726
作者:Dhruv Agarwal,Rishitha Guttapalle Mohan,Aarti Kumari,Ashi Sinha,Athulya Anil,Kavitha Srinivas,Horst Samulowitz,Andrew McCallum
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:requires an LLM, data lakes requires, data lakes, LLM agent, TAT-QA data lakes
备注:
点击查看摘要
Abstract:Deep research over data lakes requires an LLM agent to investigate evidence across thousands of heterogeneous tables and passages to synthesize a report. Existing methods perform iterative retrieval and generation, letting accumulated context determine what to investigate next, which can overexploit locally promising evidence and fail to cover distinct semantic regions under a fixed budget. To address this, we cast deep research over data lakes as a budgeted search problem and present Baikal - a framework that clusters heterogeneous evidence into semantic regions, then searches over them adaptively to balance exploration and exploitation. Within each selected region, Baikal generates and investigates region-grounded subquestions, using finding quality as rewards to update region-level value estimates and guide search under policies ranging from random and LLM-guided selection to Bayesian $\epsilon$-greedy and UCB. We evaluate Baikal on 15 queries each over HybridQA and TAT-QA data lakes containing 10,993 and 2,757 tables, respectively, together with 227K Wikipedia passages and 13K financial report passages. We assess research quality with a new rubric covering groundedness, relevance, diversity, and utility, and use GPT-5-mini to score Baikal and strong baselines, including DeepSearcher and an OpenCode research agent with retrieval and clustering variants. Across both data lakes, Baikal performs strongly under several region-selection policies; its best configuration improves report scores over the strongest baselines by 28% on HybridQA and 36% on TAT-QA. Our analyses attribute these gains to organizing and exploring semantic evidence regions, which improves groundedness and diversity and yields more useful findings under the same subquestion budget. These results demonstrate the value of structured semantic exploration for systematic research and discovery over heterogeneous data lakes.
66. 【2607.27692】Recall Before You Rank: Similarity-Guided Top-$K$ Reuse for Efficient Long-Context Attention
链接:https://arxiv.org/abs/2607.27692
作者:Wenshuai Yao,Wenyong Zhou,Hanyong Shao,Yizhe Chen,Zhiyuan Ning,Yuannuo Feng,Ru Huang,Kechao Tang
类目:Computation and Language (cs.CL); Machine Learning (cs.LG)
关键词:Exact Top, Top, sparse attention reduces, full-history Exact Top, performing global Top
备注: 9 pages, 9 figures, and 5 tables
点击查看摘要
Abstract:Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identifying this subset still requires scoring the current query against the full KV cache and performing global Top-$K$ selection, leaving selector cost linear in context length and limiting the practical efficiency of sparse attention for long-context decoding. In this paper, we introduce ReTopK, a training-free method that accelerates dynamic Top-$K$ attention by reusing historical retrieval decisions. ReTopK builds on the observation that similar queries often attend to overlapping supports and that partially overlapping supports can still preserve most of the Exact Top-$K$ attention mass. For each attention head, it maintains a bounded cache of historical query--support pairs, retrieves the most similar cached queries for each new query, unions their stored supports with a recent window, and reranks only the resulting compact candidate set using exact current-query scores. A similarity-based fallback invokes full-history Exact Top-$K$ when reuse is unreliable, while periodic exact refreshes limit cache drift. ReTopK retains the complete KV cache and reuses only selected indices, rather than historical scores, attention weights, or outputs. Across 16K--128K contexts, ReTopK achieves the lowest PG19 perplexity and the highest NIAH and LongBench scores among the evaluated approximate methods. At 128K with $K=512$, ReTopK incurs only a 0.50\% perplexity increase over Exact Top-$K$ while accelerating attention computation by $3.07\times$.
67. 【2607.27680】ght Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection
链接:https://arxiv.org/abs/2607.27680
作者:Arunan J
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:statistical properties remain, partially understood, standard mechanism, properties remain, remain only partially
备注: Springer Nature Submission
点击查看摘要
Abstract:Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only partially understood. Existing generalization results provide upper bounds of the form O~(sqrt(rd/n)) or O~(rd/n), but a matching lower bound is missing, and the question of how to choose the LoRA rank r has no formal answer. Both gaps are closed here. A local Rademacher argument establishes an upper bound of O~(rd/n) on the excess risk of the empirical risk minimizer over rank-r LoRA, whenever the target adaptation has rank at most r. A matching minimax lower bound of Omega(rd/n) is then proved via a Fano-type packing of the rank-r subspace of R^{d x d}; the bound applies to any estimator whose output lies in the rank-r LoRA class. Combining the two yields a rank-selection dichotomy. For the constrained empirical risk minimizer, the optimal rank equals the intrinsic rank r*, and over-ranking strictly hurts. For adaptive estimators of the nuclear-norm-then-truncate type, over-ranking is harmless and the rate saturates at Theta~(r* d / n) regardless of r. Taken together, the three results characterize the statistical complexity of LoRA fine-tuning within the well-specified locally quadratic regime, and identify the empirically observed over-parameterization penalty as a property of unregularized empirical risk minimization rather than of the LoRA class itself. Predictions of the theory are verified on a synthetic trace-regression benchmark and on real LoRA fine-tuning across three (model, task) configurations covering DistilBERT and RoBERTa on SST-2 and MRPC. All configurations exhibit the predicted U-shape in validation loss, with two showing statistically significant loss inflation at large ranks (paired permutation p = 0.016).
68. 【2607.27671】ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
链接:https://arxiv.org/abs/2607.27671
作者:Shengjie Li,Vincent Ng
类目:Computation and Language (cs.CL)
关键词:ASAP, automated essay scoring, AES, ICLE, evaluated solely
备注: Accepted as a long paper to NAACL 2024
点击查看摘要
Abstract:The majority of the recently-developed models for automated essay scoring (AES) are evaluated solely on the ASAP corpus. However, ASAP is not without its limitations. For instance, it is not clear whether models trained on ASAP can generalize well when evaluated on other corpora. In light of these limitations, we introduce ICLE++, a corpus of persuasive student essays annotated with both holistic scores and trait-specific scores. Not only can ICLE++ be used to test the generalizability of AES models trained on ASAP, but it can also facilitate the evaluation of models developed for newer AES problems such as multi-trait scoring and cross-prompt scoring. We believe that ICLE++, which represents a culmination of our long-term effort in annotating the essays in the ICLE corpus, contributes to the set of much-needed annotated corpora for AES research.
69. 【2607.27656】Looped Transformers with Source-Centered State Evolution
链接:https://arxiv.org/abs/2607.27656
作者:Bum Jun Kim,Kohei Hayashi,Shunsuke Kamiya,Masanori Koyama,Yusuke Iwasawa,Yutaka Matsuo
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:Looped Transformers create, test-time compute axis, increasing effective depth, Looped Transformers, fixed parameter count
备注: 24 pages, 5 figures
点击查看摘要
Abstract:Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective depth at a fixed parameter count. However, that shared block must then govern an entire trajectory of varying hidden states over trained and extrapolated depths. Furthermore, in additive-injection looped Transformers, an input-conditioned signal is reintroduced at every recurrent step, so applying the shared transition at an input-conditioned reference can still move the hidden state. In this paper, we propose Source-Centered State Evolution (SCSE), which is designed to reconcile input conditioning with reference-preserving shared recurrence. Specifically, SCSE retains input dependence through its learned anchor and initial deviation, allows nonzero deviations to drive recurrent computation while mapping zero deviation to zero, and guarantees exact anchor invariance through its zero-deviation mask. The designated anchor is thereby a one-step fixed point by construction. The zero-deviation forcing bias is the next deviation produced from the anchor itself and vanishes in SCSE, while nonzero deviations remain active and support state-dependent recurrent computation. Our theory shows that the zero-deviation forcing bias is a design degree of freedom whose task effect can be harmful, neutral, or beneficial; SCSE resolves this choice in favor of exact anchor invariance by setting the bias to zero. Across WikiText-2, WikiText-103, direct web-corpus pretraining, held-out web-text transfer, and LAMBADA completion, SCSE improves the controlled recurrent quality frontier. Ablation studies identify the learned anchor and the anchor-coordinate deviation recurrence as the primary contributors to the gain, and a trained-model case study grounds the anchor-response diagnostic in observed recurrent motion.
70. 【2607.27654】From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
链接:https://arxiv.org/abs/2607.27654
作者:Tao Wen,Shuai Shao,Pei Ke,Xu Han,Jie Zou,Guannan Li,Tao Tian,Jinjie Qiu,Lan Wang,Ke Qin
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:information extraction, involving various event-centric, Event analysis, essential and fundamental, fundamental direction
备注: 9 pages. Published in the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2026)
点击查看摘要
Abstract:Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.
71. 【2607.27652】Harness-G: A Graph-Structured Harness for Search Agents
链接:https://arxiv.org/abs/2607.27652
作者:Yanning Hou,Haoyuan Chen,Sihang Zhou,Xiaoshu Chen,Xirui Liu,Duanyang Yuan,Lingyuan Meng,Quan Liu,Jian Huang
类目:Computation and Language (cs.CL)
关键词:search agents commonly, optimize multi-turn interactions, Reinforcement learning, agents commonly model, search agents
备注: Code: [this https URL](https://github.com/7HHHHH/Harness-G)
点击查看摘要
Abstract:Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainly improve training with denser or more structured credit signals, but rarely examine whether retrieval is properly formulated at the policy-environment interface. We observe pronounced retrieval aliasing during Search-R1 training: rollouts for the same question continue to generate distinct query strings, yet their accumulated evidence sets increasingly overlap. We call this phenomenon retrieval-equivalence collapse; in this regime, trajectories approach utility equivalence with respect to retrieval decisions, leaving within-group returns with little effective retrieval contrast. To address this problem, we propose Harness-G, a graph-structured retrieval framework that redesigns this interface. It reformulates free-form query generation as finite action selection: the policy selects an evidence sentence or entity, or chooses to answer, while the environment constructs the menu, tracks retrieval state, and validates and executes each choice. This interface reduces linguistic aliasing and makes same-state alternatives directly comparable. Building on this interface, we introduce Structured Non-myopic Credit (SNC), which uses a frozen answer scorer to compare the selected action with its alternatives and assigns downstream gains to the earlier actions that enabled them. Across six QA benchmarks, Harness-G achieves the highest average F1 at both evaluated model scales, outperforming the strongest baseline, Graph-R1, by 10.74 points at 1.5B and 3.98 points at 3B.
72. 【2607.27631】ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning
链接:https://arxiv.org/abs/2607.27631
作者:Zhenrong Zhang,Fei Wu,Jun Du,Jianshu Zhang,Si Wei
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:large language models, Reinforcement learning, language models, learning has emerged, effective paradigm
备注:
点击查看摘要
Abstract:Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among existing policy optimization methods, Proximal Policy Optimization (PPO) remains particularly appealing because its learned critic can, in principle, provide token-level credit assignment. However, in mathematical reasoning tasks characterized by long reasoning horizons and sparse outcome rewards, reliable token-level credit assignment remains challenging. The standard critic often fails to accurately evaluate intermediate reasoning states, resulting in noisy advantage estimates and suboptimal policy updates. In this paper, we propose ReDiPPO, a Reference-guided and Discrepancy-aware PPO framework for mathematical reasoning. ReDiPPO introduces a reference-guided critic that uses reference answers as training-time privileged signals to provide more accurate value estimation. Meanwhile, it retains a standard critic and quantifies the token-level reference-standard discrepancy between the standard value estimate and the reference-guided value estimate. This discrepancy serves as an indicator of difficult reasoning states and is used to reweight the corresponding token-level advantages during PPO optimization. Extensive experiments on diverse mathematical reasoning benchmarks demonstrate that ReDiPPO improves value-estimation accuracy and consistently outperforms strong policy optimization baselines, including PPO, DAPO, and GSPO, in final reasoning performance. Our code is available on this https URL.
73. 【2607.27614】DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
链接:https://arxiv.org/abs/2607.27614
作者:Hongbin Zhang,Junhao Liu,Xuefeng Bai,Youcheng Pan,Yang Xiang,Kehai Chen
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:large language models, Recent advances, increasingly adopt LLMs, led sign language, converting sign-language videos
备注:
点击查看摘要
Abstract:Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM's language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.
74. 【2607.27611】AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
链接:https://arxiv.org/abs/2607.27611
作者:Qi Wang
类目:Computation and Language (cs.CL); Risk Management (q-fin.RM)
关键词:Corporate annual reports, foreign-exchange risk management, Corporate annual, weakly structured evidence, risk management
备注: 40 pages, 4 figures, 12 tables. Preprint; not peer reviewed
点击查看摘要
Abstract:Corporate annual reports contain weakly structured evidence about foreign-exchange risk management, derivative use, natural hedging, and explicit non-use. This study develops AWARE-FX, an auditable AI/NLP decision-support system that converts report text into traceable firm-year hedging-disclosure measures. The system combines a professional-source lexicon, negation and accounting-status logic, channel-specific financial encoders, exact evidence gates, conservative aggregation, and an audit ledger. Across 24,909 Hong Kong firm-years from 2008-2025, it retrieves and scores 543,527 snippets. Reliability is evaluated through ablations, a stratified 300-snippet human audit, three-seed FinBERT-ModernBERT comparisons, strict 2023-2025 temporal tests, probability calibration, selective prediction, and fixed-prompt generative-model benchmarks. FinBERT has the higher mean F1 in seven of eight encoder task-split comparisons; its temporal F1 ranges from 0.702 to 0.872. Abstaining on the 20% least-confident temporal observations raises retained-sample F1 by 0.050-0.077. Deterministic Qwen3-8B performs strongly on commodity and negation evidence but poorly on foreign-debt and accounting-context labels, showing that a general-purpose LLM does not uniformly replace domain constraints. The strict FX score is negatively associated with linked baseline and stress-period FX exposure, whereas the generic broad score is not. These associations provide external construct validation, not causal estimates of hedging effectiveness. AWARE-FX contributes a tested decision-support architecture in which retrieval, status logic, classification, uncertainty handling, aggregation, and external validation remain separately auditable.
75. 【2607.27595】Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories
链接:https://arxiv.org/abs/2607.27595
作者:Zhaoji Wang,Wanyu Si,Jun Wang
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL)
关键词:Computational approaches, neural retrieval, parallel-passage lists, advanced from string, string matching
备注: 9 pages, 4 figures, 3 tables
点击查看摘要
Abstract:Computational approaches to intertextuality have advanced from string matching to neural retrieval, yet their outputs, similarity scores and parallel-passage lists, identify where texts reuse one another without characterizing how or why. We recast fine-grained intertextuality extraction as an agentic task in which a large language model (LLM) reads two text units in full and, through a constrained tool interface, must ground each proposed reuse in exact character spans on both sides and label it under a five-dimension typology of reuse (form, aspect, source-marking, function, stance). We validate the approach on an exhaustive comparison of the Analects with the Book of Han, where three domain experts adjudicate a pooled multi-model candidate set into a benchmark of 2,533 intertextual pairs. Against this standard we study twelve LLMs, reporting precision (56%-93%), a 51$\times$ cost spread at comparable quality, and how well their confidence is calibrated. Expert agreement traces a reliability gradient: dimensions legible on the textual surface are annotated consistently, while those requiring inference of intent are contested, delimiting the claims such annotation supports. Scaling the validated extractor to the full Twenty-Four Histories (65,380 comparisons, 5,766 pairs) recovers corpus-level structure a similarity score cannot express. The interpretive composition of citation shows no systematic change across eighteen centuries, yet the same passage is quoted less and less literally. Stability in the aggregate with drift in the individual case is what a cultural-attraction account expects. We release the extraction protocol and the expert-adjudicated benchmark.
76. 【2607.27591】Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
链接:https://arxiv.org/abs/2607.27591
作者:Jinyi Liu,Wei Chen,Pengyu Chen,Xinyi Yuan,Minghe Bai,Guoquan Wu,Jun Wei
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:dominate memory traffic, Feed-forward networks, large language model, dominate memory, activation sparsification
备注:
点击查看摘要
Abstract:Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activation sparsification. However, existing training-free methods suffer substantial model-quality degradation at high sparsity due to limitations in their channel-selection strategies. We observe that the SwiGLU intermediate state provides a highly effective channel-selection signal, but obtaining it requires costly dense computation. To address this, we present \emph{Prox}, a two-stage training-free framework for sparse SwiGLU FFNs. Prox hinges on the key insight: sparse execution requires only the channel mask induced by the intermediate state, which can be constructed from the magnitude ranking of its entries rather than their exact values. Specifically, Stage 1 uses input sparsity and quantized proxy weights to construct a shared mask; Stage 2 computes the selected channels exactly, enabling sparse execution of all three projections. Across ten LLMs from six model families, Prox outperforms training-free baselines at all sparsity levels, achieves up to a $1.99\times$ end-to-end decoding speedup at 70\% FFN sparsity, and is compatible with quantization and sparse attention.
77. 【2607.27557】raining Skills Like Parameters via Self-Supervised Semantic Diffusion
链接:https://arxiv.org/abs/2607.27557
作者:Mo Li,Zixin Yin,Ting Cao,Yunxin Liu
类目:Computation and Language (cs.CL)
关键词:Large Language Models, Large Language, remarkable general instruction-following, general instruction-following capabilities, demonstrate remarkable general
备注: Preprint, work in progress
点击查看摘要
Abstract:While Large Language Models (LLMs) demonstrate remarkable general instruction-following capabilities, they often fall short of human experts in highly specialized, open-ended domains such as creative screenwriting. Prior approaches typically adopt post-training, yet both supervised fine-tuning and reinforcement learning require weight access that closed-source frontier models do not offer, and demand heavy compute. Moreover, what is learned is tied to a single checkpoint and cannot be inspected by humans. Recent advancements in agentic continual learning instead attempt to bridge this gap by accumulating external textual skills. However, these methods heavily rely on costly human expert annotations or unreliable LLM-as-a-judge feedback for reflection. To overcome this bottleneck, we propose a novel, unsupervised self-evolving agent framework inspired by the corruption-and-reconstruction paradigm of diffusion models. Instead of relying on explicit external scoring, we leverage existing high-quality human artifacts to construct self-supervised signals. Training then follows the familiar loop of neural network training, forward, loss, and backward, with the loss coming from contrasting the agent's reconstruction against the human original. What is updated is not model weights but an external library of textual skills. We evaluate our framework on the challenging task of short drama screenwriting. Experimental results demonstrate that our method enables the agent to autonomously extract and internalize highly generalizable skills, significantly enhancing its domain-specific generation capabilities. Furthermore, this self-contrastive reflection paradigm offers a scalable pathway for agents to teach themselves the production of complex, high-quality human artifacts, without requiring external supervision.
78. 【2607.27553】Using Large Language Models for Idea Generation in Innovation
链接:https://arxiv.org/abs/2607.27553
作者:Lennart Meincke,Karan Girotra,Gideon Nave,Christian Terwiesch,Karl T. Ulrich
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); General Economics (econ.GN)
关键词:large language models, ideas, language models, efficacy of large, large language
备注:
点击查看摘要
Abstract:This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for new products targeted toward college students and priced at 50 dollars or less. The first pool of ideas was created by university students in a product design course before the availability of LLMs. The second and third pools of ideas were generated by GPT-4 from OpenAI using zero-shot and few-shot prompting, respectively. We evaluated idea quality using standard market research techniques to predict average purchase intent probability. We used text mining to assess idea similarity and human raters to evaluate idea novelty. We find that AI-generated ideas outperform human-generated ideas in terms of average purchase intent, with few-shot prompting yielding slightly higher intent than zero-shot prompting. However, AI-generated ideas are perceived as less novel and exhibit higher pairwise similarity, particularly with few-shot prompting, indicating a less diverse solution landscape. When focusing on the quality of the best ideas rather than the average ideas, we find that AI-generated ideas are seven times more likely to rank among the top 10 percent of ideas, demonstrating a significant advantage over human-generated ideas. We propose that this seven-to-one advantage is a conservative estimate because it does not account for the greater productivity of AI. Our findings suggest that despite some drawbacks, AI creativity presents a substantial benefit in generating high-quality ideas for new product development.
79. 【2607.27539】Subtract or Replay? Exact Deletion from Language-Model Memory
链接:https://arxiv.org/abs/2607.27539
作者:Vishwajith Ramesh
类目:Machine Learning (cs.LG); Computation and Language (cs.CL)
关键词:persistent language-model memory, language-model memory depends, persistent language-model, language-model memory, memory depends
备注: 22 pages, 8 figures
点击查看摘要
Abstract:Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic decrement; influence transformed by later writes inside shared recurrent state requires rebuilding from before the write. We test this distinction in two pretrained models against explicit record-omitted references. First, we replace Gemma 3's global-attention layers with support-vector memory. After low-rank recovery at 1B, decrement and retained-key refit agree at the next-token output to median KL $5.4\times10^{-15}$ over 31 support-token deletions, with $+2.0\%$ perplexity relative to a matched fine-tune. A masked-refit proxy is indistinguishable from the never-ingested floor under elicitation, relearning, sampling, and LiRA attacks. At 4B and 12B, certificate ordering persists but utility cost rises to $11.2\%$ and $44.3\%$. Second, in a 48B Kimi Linear hybrid, additive writes admit a fixed decrement and diagonal decay a corrected one, whereas the delta rule makes $12$--$49\%$ of a record's contribution suffix-dependent. Checkpointed rewind-and-replay deletes real clinical records at contexts up to 18,842 tokens, matching never-ingested logits and all recurrent states bit for bit within a deterministic MLX implementation; replaying a correction provides exact amendment. Exact deletion is therefore a property of memory representation: subtract addressable records and replay entangled writes.
80. 【2607.27528】hreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping
链接:https://arxiv.org/abs/2607.27528
作者:Cristian Leo,Anton Dykyi,Danny Cortegaca,Daniel Begimher,Prakash Jha
类目:Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Software Engineering (cs.SE)
关键词:secure software development, scarce security expertise, demands scarce security, software development, security expertise
备注: 20 pages, 12 tables, 1 figure
点击查看摘要
Abstract:Threat modeling is essential for secure software development, yet manual analysis of cloud-native architectures is slow and demands scarce security expertise. We present ThreatForest, a multi-agent system that generates structured attack trees from source code repositories, maps attack steps to adversary tactics, techniques, and procedures (TTPs) from a pluggable set of frameworks (MITRE ATTCK, CAPEC, and cloud-specific threat matrices), and synthesizes actionable mitigations. ThreatForest decomposes threat modeling into a multi-stage agent pipeline -- repository analysis, context refinement, threat generation, parallel attack-tree construction with TTP mapping and mitigation synthesis, and report generation -- orchestrated as a directed graph with deterministic verification gates, bounded retries, and three human-in-the-loop validation points. A domain-specific sentence-transformer maps each attack step to candidate techniques by cosine similarity; we show empirically that this embedding stage, not the surrounding pipeline, is the dominant accuracy bottleneck. We evaluate ThreatForest across seven application domains on a sixteen-dimension rubric, scored by a panel of independent LLM raters with an adversarial verification pass and expert review. Panel-measured quality reaches 0.63-0.68 (on a 0-1 scale) for threat statements, attack trees, and mitigations, but only 0.29 for embedding-only TTP mapping -- a gap stable across all seven domains that isolates the binding constraint. A controlled single-call baseline on the same model more than doubles mapping defensibility, pinning the limitation on the embedding encoder rather than the multi-agent design. To our knowledge, ThreatForest is the first end-to-end system that turns a code repository into TTP-mapped attack trees with evidence-based mitigations across adversary frameworks, with a reusable framework for benchmarking such systems.
81. 【2607.27512】Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models
链接:https://arxiv.org/abs/2607.27512
作者:Germans Savcisens,Samantha Dies,Courtney Maynard,Tina Eliassi-Rad
类目:Computation and Language (cs.CL); Multiagent Systems (cs.MA); Social and Information Networks (cs.SI)
关键词:Large language models, Large language, increasingly deployed, LLM, Large
备注: 33 pages (14 pages of main text), 7 figures, 14 tables
点击查看摘要
Abstract:Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework for studying belief diffusion within networked LLM populations. CoevolveSim allows us to isolate and study three factors: domain specialization, social-role assignment, and social network structure. Within this framework, generalist and specialist LLM agents exchange and revise beliefs. In each round, an LLM agent observes a summary of its neighbors' beliefs before updating its own. We run 1,280 controlled simulations spanning four scenarios, two network structures, and 20 medical-indication statements. We find that persona-style role assignment and network structure reshape individual belief revision but have minimal effect on population-level consensus. In contrast, introducing (finetuned) specialist LLMs more than doubles the shift in consensus and gives rise to consistent asymmetries in exerted influence. We further show that simple persistence-based opinion-dynamics models reproduce collective outcomes in all-generalist LLM populations, whereas heterogeneous LLM populations require population-level belief composition to reproduce consensus and agent identity to predict individual belief transitions. Our results indicate that realistic simulation of belief diffusion in multi-agent LLM systems requires a diverse set of underlying LLMs, not persona prompting alone.
82. 【2607.27506】Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
链接:https://arxiv.org/abs/2607.27506
作者:Shreyas Subramanian,Mecit Gungor,Vikram Elango
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Relative Policy Optimization, Group Relative Policy, Language and embedding, systems are conventionally, conventionally assumed
备注: 28 pages, 3 figures. Submitted to COLM 2026
点击查看摘要
Abstract:Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built components: a compact embedding model and a purpose-built SLM. B1ade-embed, a 335M parameter retrieval model constructed via parameter-free fusion of five pretrained encoders achieves top MTEB scores among sub-500M models with zero additional training, and B1ade-1B, an SLM trained on low-cost GPUs using Group Relative Policy Optimization (GRPO) on 723M tokens (2.2M examples) of curated context-question pairs with rewards that optimize only answer similarity. Our central finding is emergent attribution: despite receiving no explicit supervision for source citation, B1ade-1B cites retrieved passages in 42.4% of responses, exceeding the attribution rate of its training distribution by 5.5 percentage points. This demonstrates that grounding behavior can emerge as an accuracy-maximizing strategy under RL training, without explicit reward engineering. On standard QA benchmarks, B1ade-1B achieves 81.82% on PopQA, 65.8% on PubMedQA, and 51.09% on FEVER. In end-to-end RAG evaluation, B1ade-1B achieves an average score of 0.654 across correctness, completeness, coherence, and faithfulness, a 10.8% improvement over the SFT, while closing the gap with models 1.5x its size. These results show that strategic model composition and reward design suffice for resource-efficient RAG, without large-scale pretraining.
83. 【2607.27497】SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge
链接:https://arxiv.org/abs/2607.27497
作者:Lucio M. Dery,Benedict Aaron Tjandra,Siavash Samiei,Adhiguna Kuncoro,Zohar Yahav,Jiajun Shen,Arthur Szlam
类目:Computation and Language (cs.CL)
关键词:Agentic systems driven, solve complex problems, autonomously solve complex, synthesizing text-based knowledge, Agentic systems
备注:
点击查看摘要
Abstract:Agentic systems driven by large language models (LLMs) regularly feature two key mechanisms to autonomously solve complex problems: synthesizing text-based knowledge and procedures from past experiences and building parametric (weight-space) skill libraries for recurring sub-goals. To date, research has largely treated these as orthogonal pursuits: either organizing textual knowledge through composition and reflection, or consolidating parametric skills via weight-space merging. Consequently, the seamless integration of text and model weights for targeted performance improvements remains largely unexplored. This work bridges this modality gap by treating model weights as an additional modality that an LLM can natively reason over. We instantiate parametric learning via prefix-tuning and augment an LLM to ingest both prefix weights and rich textual data which capture relationships to a target capability. Our augmented LLM, which we call SkillSmith, synthesizes these inputs to perform instruction-steered parametric synthesis, directly outputting new prefix weights that manifest the target skill. We demonstrate that our approach significantly outperforms both text-only and weight-space-only baselines, unlocking performance gains that are out of reach for uni-modal (text-only or weight-only) adaptations.
84. 【2607.27482】Latent States in Neural Networks: Recovering the Temporal Structure of Drifting Data from Model Weights
链接:https://arxiv.org/abs/2607.27482
作者:Kevin Guan
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:drifting data stream, changing continuously, pass through discrete, discrete regimes, temporally drifting data
备注:
点击查看摘要
Abstract:A temporally drifting data stream may pass through discrete regimes rather than changing continuously. We ask whether such regimes are recoverable from the weights of models trained on the stream, using a hidden Markov model (HMM) fit to the chronologically ordered trajectory of those weights. We study this question in two domains known to drift over time: multimodal misinformation detection, using the Fakeddit dataset; and sentiment analysis, using the Yelp dataset. We train classifiers on consecutive temporal windows and fit an HMM to the trajectory of their aligned weights, recovering latent states that partition each timeline into coherent phases. On both datasets, classifiers generalize better to data from windows sharing the state of their training window than to windows across state boundaries. This within-state transfer advantage survives a control for temporal proximity and modestly exceeds the advantage recovered by a naive partition into contiguous states of equal size. Although the states are estimated solely from model weights, they correlate more strongly with shifts in the data's class distribution than with the weight-space geometry used to estimate them. After class divergence and lag are residualized out, the within-state advantage exceeds its permutation null on both tasks, indicating that the states recover structure relevant to transfer beyond the data distribution. Every effect replicates on both tasks but is attenuated on Yelp, whose label distribution is more temporally stable.
85. 【2607.27421】Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models
链接:https://arxiv.org/abs/2607.27421
作者:Parishruthi Ganesh,Gerry Dozier,Cheryl Seals
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:task-oriented dialogue systems, deployable open-weight language, open-weight language models, dialogue systems, open-weight language
备注:
点击查看摘要
Abstract:Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable open-weight language models under compute, latency, and robustness constraints. We present a systematic zero-shot evaluation of 41 open-weight language models spanning 15 families and the 135M--9B parameter range across eight English single-label intent-classification datasets. A ninth dataset, ATIS, uses five labeled demonstrations and is reported as an auxiliary five-shot result. The evaluation includes standard benchmarks, a large-scale voice-assistant corpus, and production-derived e-commerce datasets. Beyond exact-match accuracy, we analyze confidence calibration, robustness to realistic input perturbations, statistical reliability of model rankings, deployment efficiency, and benchmark saturation. Our results show that instruction-tuned 3B models can outperform several evaluated 7B base models, that differences among leading models on MASSIVE are statistically indistinguishable under pairwise McNemar tests, and that widely used benchmarks such as SNIPS have become saturated and no longer meaningfully discriminate among current open-weight models. Instruction tuning's effect on confidence calibration is inconsistent rather than uniformly harmful. These findings provide practical guidance for selecting and evaluating open-weight language models for intent classification.
86. 【2607.27420】Dimensionality and Measurement Precision in HLE's Multiple-Choice Subset
链接:https://arxiv.org/abs/2607.27420
作者:Mayank Sharma,Savira Nadela,Tyler Matteson
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Humanity Last Exam, evaluate frontier language, HLE, frontier language models, Humanity
备注:
点击查看摘要
Abstract:Humanity's Last Exam (HLE) is widely used to evaluate frontier language models. HLE organizes its questions into eight subject-domain categories, whose subscores are often interpreted as evidence of distinct capabilities. However, no study has assessed whether these labels correspond to empirically separable latent constructs, nor whether the benchmark effectively differentiates between models of similar ability. We evaluate 29 LLMs on the text-only multiple-choice subset of HLE ($J = 428$ items) and apply psychometric methods to assess both the dimensionality of the benchmark and the distribution of its measurement precision. Fitting a two-parameter logistic IRT model, we find convergent evidence that HLE measures a single general reasoning factor: McDonald's $\omega_h = 0.998$, domain labels explain only 3.5\% of item response variance, within- and between-domain residual correlations are nearly identical (Cohen's $d = 0.016$), and domain-specific ability estimates are near-redundant with the total score ($r \geq 0.81$). A separate analysis of the test information function reveals that measurement precision concentrates at moderate ability levels and drops sharply above $\theta = 0$, where frontier models sit. These findings suggest that HLE's domain subscores do not warrant distinct capability interpretations and that the benchmark's ability to discriminate among the strongest models is limited.
87. 【2607.27405】Benchmarking LLM Competence on Logical Inference over Probability Operators
链接:https://arxiv.org/abs/2607.27405
作者:Nayera Hasan,Jack Greff,Alvin Grissom II
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:expressions of uncertainty, natural-language expressions, medicine and law, ubiquitous in natural, everyday conversations
备注: Under review
点击查看摘要
Abstract:Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.
88. 【2607.27393】AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
链接:https://arxiv.org/abs/2607.27393
作者:Mohamed Bayan Kmainasi,Ali Ezzat Shahroor,Abul Hasnat,Md. Rafiul Biswas,Wajdi Zaghouani,Firoj Alam
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Arabic hateful meme, Arabic HAteful, Arabic HAteful Memes, multimodal online harm, hateful meme detection
备注: 26 pages, 14 figures, 15 tables
点击查看摘要
Abstract:Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels. We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label annotations. The dataset includes 5K manually annotated memes using a taxonomy that captures hate types, i.e., attack strategies. We further provide ~66K silver-labeled memes to support future studies. We benchmark text-only, image-only, and late-fusion multimodal models, as well as few-shot in-context learning (ICL) and open- and closed-weight Vision-Language Models (VLMs) under zero-shot and fine-tuning settings. Our results establish strong baselines and highlight key challenges in culturally grounded Arabic hateful meme detection. We release the dataset, annotation guidelines, and evaluation scripts to support future research. WARNING: This paper contains examples that may be disturbing to readers.
89. 【2607.27384】Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
链接:https://arxiv.org/abs/2607.27384
作者:Prabhjot Singh,Pritam Deka,Vijay Chennareddy
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
关键词:Narrative Anchoring Gap, Narrative Anchoring, Large language models, Large language, mode Narrative Anchoring
备注:
点击查看摘要
Abstract:Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.
90. 【2607.27379】HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs
链接:https://arxiv.org/abs/2607.27379
作者:Ru Peng,Tianyu Zhao,Xijun Gu,Zhiting Fan,Haokai Xu,Jinyang Zhang,Yawen Zeng,Yihong Zhuang,Kexin Yang,Junyang Lin,Dayiheng Liu,Junbo Zhao
类目:Computation and Language (cs.CL)
关键词:large language models, language models, scarce and costly, vital for large, large language
备注: ACL Findings 2026 Paper
点击查看摘要
Abstract:High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying "requirements + persona" to backtranslate seed documents into diverse yet faithful instructions with a strict QA alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth's robustness and transferability. Our code is publicly available at this https URL.
91. 【2607.27372】Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
链接:https://arxiv.org/abs/2607.27372
作者:Alexi Gladstone,Heng Ji,Yilun Du
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:deep learning revolution, training beats decomposing, learning revolution, hand-designed stages, Generative modeling
备注:
点击查看摘要
Abstract:The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generative modeling, however, has remained the exception-despite generative models being remarkably capable, they are still not trained end-to-end. This is because, at its core, generative modeling is about handling distributions with many modes, and existing scalable approaches handle this the same way, by factoring the generation procedure, which prevents end-to-end generation. In this work, we introduce Explorative Modeling, a new paradigm that instead factors the training loop, exploring K candidate matches between model generations and data, and training on the best, so predictions commit to modes rather than blurring them. We find Explorative Models (XMs) useful in two settings. First, increasing exploration adds a third pretraining axis beyond parameters and data for existing generative models-where scaling exploration monotonically improves performance across both continuous and discrete domains (images, video, and language). Notably, gains from exploration increase with scale, climbing from 7% to 36% as data scales and from 13% to 23% as models grow, with efficiency gains more than doubling at 3x the compute. Concretely, exploration improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, parameter efficiency by 47%, lifts the strongest of image-generation recipes to a near-state-of-the-art 1.43 FID on ImageNet without guidance, enables scaling how end-to-end existing models are, and unlocks scaling generalization. Second, XMs enable end-to-end reconstructive generative modeling, matching diffusion on control tasks with 16-256x fewer inference steps. Together, these results establish XMs as both a new pretraining axis for existing generative models and a standalone end-to-end generative modeling paradigm.
92. 【2607.27366】BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences
链接:https://arxiv.org/abs/2607.27366
作者:Ru Peng,Haokai Xu,Xijun Gu,Tianyu Zhao,Zhiting Fan,Yawen Zeng,Yihong Zhuang,Jinyang Zhang,Kexin Yang,Jian Wu,Hao Chen,Junyang Lin,Dayiheng Liu,Junbo Zhao
类目:Computation and Language (cs.CL)
关键词:overlooking open-ended humanities, large language models, primarily targets domains, broad HSS disciplines, nuanced quality judgments
备注:
点击查看摘要
Abstract:While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than objective correctness. This makes preference alignment a natural paradigm for broad HSS tasks. Yet existing methods are either costly or not tailored to broad HSS disciplines. We thus propose BridgeAlign, among the first preference-alignment pipelines for broad HSS disciplines, with three phases: i) Seed Curation: curating HSS seed documents from web corpora via heuristic/LLM-based filtering and text refinement; ii) Preference Data Synthesis: generating preference triplets via persona-based instruction inversion with QA consistency checks; iii) Preference Optimization: moving beyond naive human-vs-model heuristics by first grounding preferences in HSS quality rubric, then generating transitional responses via controlled quality degradation to form near-boundary preference pairs for finer-grained quality discrimination. Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines; importantly, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them, as supported by extensive experiments and contextualized by existing theories.
93. 【2607.27353】LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
链接:https://arxiv.org/abs/2607.27353
作者:Musa Shams(Independent Researcher)
类目:Computation and Language (cs.CL)
关键词:Agentic retrieval-augmented generation, retrieval-augmented generation systems, Agentic retrieval-augmented, retrieval-augmented generation, generation systems
备注: 10 pages, 9 tables. Code and data: [this https URL](https://github.com/MusaShams/layerrag-bench)
点击查看摘要
Abstract:Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. Schema normalization raises schema-drift success from 0.000 to 0.913, but stale evidence, missing tool output, denied permissions, and wrong-session context are not recovered by schema normalization. Groundedness-only evaluation also produces substantial false positives under stale and wrong-session evidence. These results support a layer-specific evaluation principle: a reliability intervention should be credited for repairing its target layer without being mistaken for a universal fix.
94. 【2607.27303】HGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model
链接:https://arxiv.org/abs/2607.27303
作者:Yixin Peng,Diego Collarana,Er Jin,Stefan Decker
类目:Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML)
关键词:relation types co-exist, dynamic relational systems, Temporal Attention, heterogeneous graphs offer, Rotary Temporal Attention
备注:
点击查看摘要
Abstract:Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve over time. Learning on such graphs requires jointly modeling cross-type structural heterogeneity and the temporal dynamics of interactions, yet existing methods still struggle to reconcile parameter-efficient cross-type transfer with relation-aware specialization, and typically inject time only as additive features outside the attention kernel. We propose \textbf{THGFM}, a web-scale temporal heterogeneous graph fusion model that addresses both limitations within a unified dual-path architecture. THGFM couples a \textit{Shared-Space Temporal Attention} branch for parameter-efficient cross-type transfer with a \textit{Relational Type-Partitioned Temporal Attention} branch for relation-aware specialization, and integrates them through \textit{Dual-Path Relational--Shared Fusion}, instantiated with \textit{Type-Conditioned Non-Competitive Gated Sum Fusion}: a adaptive mechanism that assigns independent, type-conditioned feature-wise gates to the shared and specialized branches, allowing both to be amplified or suppressed without zero-sum competition. To directly incorporate relative time into the attention score, THGFM further introduces \textit{Rotary Temporal Attention}, which rotates queries and keys by half-phases of relative time before matching. THGFM consistently outperforms baseline graph transformer models on academic graphs benchmarks, delivering a $+3.25\%$ six-task mean gain, with peak relative gains of $+12.37\%$ on OAG-CS PV, $+4.87\%$ on PF-$L_2$, and $+1.18\%$ on PF-$L_1$, and $+4.24\%$, $+3.73\%$, and $+4.61\%$ on OGBN-MAG, HTAG-ArXiv, and HTAG-DBLP, respectively.
95. 【2607.27232】Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups
链接:https://arxiv.org/abs/2607.27232
作者:Haran Shani-Narkiss,Michael Fire,Oren Tsur
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY); Machine Learning (cs.LG)
关键词:Large Language Models, Large Language, form our worldview, Language Models, increasingly shaping
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.
96. 【2607.27228】AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026
链接:https://arxiv.org/abs/2607.27228
作者:Tazro Ohta,Nomi L. Harris,Seth Carbon
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Digital Libraries (cs.DL)
关键词:prepare submission materials, submission materials, prepare submission, Open Source Conference, Bioinformatics Open Source
备注: 18 pages, 4 figures
点击查看摘要
Abstract:Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences are seeing an overwhelming surge of submissions. We wanted to see if generative AI could help our conference's volunteer reviewers by pre-reviewing abstracts for certain criteria. The Bioinformatics Open Source Conference (BOSC) was well-positioned to experiment with this, as we already had a detailed rubric used by reviewers to evaluate submitted abstracts on multiple criteria, including openness (public availability of the code or other content associated with the project), valid open source license, and "runnability" (how easy it is to download, build, and run the project - an important measure of reusability). For BOSC 2026, we built bosc-pre-review, an agentic skill that assessed six review criteria, and Runabilly, which builds and tests each project in a disposable Docker container for safety. The AI only gathered evidence to present to the reviewers; humans made every decision regarding the acceptance of the abstracts. After the review period, we surveyed the reviewers to determine how useful they found the pre-review. Most of those who responded said they found it useful, but they preferred to check the AI's conclusions against their own, rather than accepting the AI results unquestioningly.
97. 【2607.27212】Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy
链接:https://arxiv.org/abs/2607.27212
作者:Asif Azad,Mohammad Sadat Hossain,MD Sadik Hossain Shanto,Sabri Boughorbel,Abdulrhman Aljouie,Bdour Alwuqaysi,Yahya Bokhari,Ayah Othman Sindi,Ehsan Hoque
类目:Human-Computer Interaction (cs.HC); Computation and Language (cs.CL)
关键词:Autism Spectrum Disorder, Arabic-speaking countries face, Children with Autism, limited service reach, Autism Spectrum
备注:
点击查看摘要
Abstract:Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of qualified specialists, limited service reach beyond urban centers, and a near-total absence of culturally grounded digital therapy materials. We present Digital Harf, a pervasive, multimodal AI platform that extends clinician-led speech and language therapy into the home. The system integrates three therapeutic modules - language therapy, speech intelligibility, and picture description - within a unified workflow that adapts to each child's performance over time. To address the critical shortage of Arabic therapy content, we introduce an Agentic Synthetic Data Engine (ASDE) that automatically generates culturally relevant images, prompts, and language tasks guided by explicit therapeutic and cultural criteria. Expert evaluation with 13 licensed Speech-Language Pathologists yielded a 90.1% clinical acceptance rate for ASDE-generated content without any manual curation or selection, and strong ratings for cultural and linguistic alignment across the full platform. Digital Harf demonstrates that AI-driven therapeutic systems can be built from the ground up for underrepresented linguistic settings, treating cultural grounding as core infrastructure rather than an adaptation afterthought.
98. 【2607.27210】Prompt Chaining in Practice: A Case Study in Automated Scholarly Report Generation
链接:https://arxiv.org/abs/2607.27210
作者:Andrei Lazarev
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computational Engineering, Finance, and Science (cs.CE); Digital Libraries (cs.DL); Software Engineering (cs.SE)
关键词:publications requires automated, requires automated tools, scholarly publications requires, effective information synthesis, exponential growth
备注: This is the version of the article accepted for publication in SUMMA 2025 after peer review. The final, published version is available at IEEE Xplore: [this https URL](https://doi.org/10.1109/SUMMA68668.2025.11302303)
点击查看摘要
Abstract:The exponential growth of scholarly publications requires automated tools for effective information synthesis. However, simple, single-shot prompting methods often lack the reliability and quality required for complex synthesis tasks. This paper introduces and empirically evaluates a multi-stage prompt chaining methodology as a more reliable architectural pattern for such tasks. This approach is implemented in our system, AI SciBrief, which automatically generates scholarly digests. We conducted a comparative experiment, measuring the performance of our prompt chaining method against a carefully optimized single-shot baseline. Both systems were evaluated against a human-authored "gold standard" report for the "Education" domain. The results demonstrate a significant difference in reliability: our prompt chaining method achieved a 100% success rate, whereas the optimized baseline failed in 50% of its runs. In terms of quality, the proposed method also demonstrated a clear advantage, achieving a superior ROUGE-L F1-score (0.507 vs. 0.486), driven primarily by higher precision. We conclude that prompt chaining is a more dependable and effective engineering approach for complex, multi-step generative tasks, significantly mitigating the risks of failure and inconsistency inherent in monolithic prompts.
99. 【2607.17751】MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
链接:https://arxiv.org/abs/2607.17751
作者:HONOR Agentic Search Team:Zhengzong Chen,Lei Tang,Lijun Liu,Chuandi Jiang,Fan Yang,Keyun Chu,Chu Zhao,Shihao Liu,Minghang Li,Bo Liang,Can Wen,Hailong Wu,Jingnan Ju,Mian Liu,Nengbin Zhang,Peiqiang Wang,Penghe Nie,Qinhui Gu,Sijia Lv,Siqi Chen,Wei Zhang,Yang Xu,Yuhao Qian,Yuxiang Zhang,Zeng Cheng,Zhen Wang,Zuan Chen,Yuanyuan Zhao,Fei Huang
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:joint optimization framework, optimization framework integrating, integrating Counterfactual task, framework integrating Counterfactual, Counterfactual task decomposition
备注:
点击查看摘要
Abstract:We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios. We empower MagicSelector with these capabilities through three key contributions: (1) a preference-guided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.
信息检索
1. 【2607.28618】AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
链接:https://arxiv.org/abs/2607.28618
作者:Bing Yan,Gregory Wolfe,Stefano Martiniani,Kyunghyun Cho
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:ranked document lists, requires assembling specific, assembling specific findings, specific findings scattered, primarily return ranked
备注:
点击查看摘要
Abstract:Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to locate relevant information, verify their provenance, and assemble cross-paper answers manually. We present AskChem, a claim-centered infrastructure for cross-paper chemistry search. AskChem changes the unit of retrieval from the paper to the provenance-carrying claim: each paper is converted into atomic, typed claims, each grounded by a source DOI and a verbatim quote or an explicit evidence locator. Over this shared claim store, AskChem exposes complementary structures for search and synthesis: a stabilized faceted taxonomy for hierarchical retrieval and browsing, an evidence graph linking claims through relations, and an exploratory living taxonomy that situates indexed papers under scientific principles. AskChem currently indexes 2.4M claims from 147K papers and provides a web interface, as well as REST, SDK, and MCP access for AI agents. On AskChem-Bench, grounding a GPT-5.5 reader in AskChem yields 100% resolvable DOIs, compared with 88.3% without retrieval, and the highest citation density among five tested systems. AskChem is live at this https URL.
2. 【2607.28571】Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently
链接:https://arxiv.org/abs/2607.28571
作者:Simon Roy,Mark Bong,Giovanni Beltrame
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:Operational Earth observation, Earth observation increasingly, Operational Earth, observation increasingly calls, satellite image pairs
备注: 10 pages, 3 figures
点击查看摘要
Abstract:Operational Earth observation increasingly calls for answering queries such as ``find the image pairs where a new building appeared.'' This means searching an archive of before-and-after (bi-temporal) satellite image pairs and ranking each pair by how well it matches a natural-language description of the change. The component that performs this match, the fusion module that combines the ``before'' and ``after'' views, must be run at query time across many candidate pairs, so its speed largely sets the cost of every search. We present a controlled comparison of how to build that module. Using one fixed image encoder (a frozen CLIP model) and one training recipe for all variants, we evaluate eight designs drawn from three families: attention, state-space models (Mamba), and learned compression (our Temporal Bottleneck Fusion, TBF). Each design is tested on two benchmarks (LEVIR-CC and Dubai-CC) with ten random seeds, so the reported differences are statistically grounded. We outline three findings: first, a training-free two-stage search (a cheap difference model that shortlists candidates, followed by attention fusion that re-ranks them) matches or exceeds full-fusion recall on LEVIR-CC while cutting query cost $10$-$15\times$, with comparable R@1/R@5 on Dubai-CC; second, the linear-time scan of Mamba, attractive on paper, gives no speed benefit at the patch counts typical of vision transformers ($L{=}196$): the scan is limited by memory bandwidth, whereas attention maps cleanly onto parallel hardware; and third, compressing the fused representation (TBF) reduces parameters by $2.3\times$ and latency by $1.6\times$ for a change-only BLEU-1 cost of $0.007$, although more aggressive compression quietly discards change-relevant detail that aggregate metrics fail to reveal.
3. 【2607.28498】CA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval
链接:https://arxiv.org/abs/2607.28498
作者:Yuto Suzuki,Farnoush Banaei-Kashani
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:Science typically involves, Science typically, Scientific hypothesis generation, typically involves Scientific, hypothesis composition
备注:
点击查看摘要
Abstract:Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Existing SIR methods rank papers by topical similarity and do not explicitly represent how a candidate inspiration transfers to a target problem. This is especially limiting for remote inspirations, whose value often lies in reusable problem-solving principles rather than topical overlap. Motivated by how humans abstract transferable aspects of a source and remap them to a new target, we reformulate SIR as target-conditioned abstraction (TCA). The retrieval object is a transferable abstract principle extracted from a candidate specifically for the target. We present TCA-SIR, which learns to generate target-conditioned abstractions and uses their representations to predict transferability. On ResearchBench, TCA-SIR outperforms prior SIR methods and direct LLM retrieval, improving HitRate@top4% over MOOSE-Chem by more than 10 percentage points. Learned abstractions also recover target-relevant mechanisms more clearly than an untrained TCA prompt, yielding both stronger retrieval and an interpretable rationale for scientific inspiration.
4. 【2607.28397】GLM-RAG: Graph Language Models for Graph-Based Retrieval-Augmented Generation
链接:https://arxiv.org/abs/2607.28397
作者:Maya Arseven,Anette Frank,Beni Egressy,Johann Higl,Moritz Plenz
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:Retrieval-augmented generation, knowledge graphs requires, effectively capture, Retrieval-augmented, semantic information
备注: 10 pages, 19 figures
点击查看摘要
Abstract:Retrieval-augmented generation (RAG) over knowledge graphs requires retrievers that can effectively capture both graph structure and semantic information. Recent approaches have explored graph neural network (GNN)-based retrievers to model graph topology in multi-hop reasoning tasks. In parallel, graph language models (GLMs) have emerged as a promising paradigm that integrates graph reasoning and the semantic capabilities of language models. In this work, we introduce a GLM-based retriever and investigate the comparative strengths of GLM-based, GNN-based, and traditional vector-search-based retrievers in single- and multi-hop RAG settings, and with a particular focus on transferability to unseen domains. Our findings suggest that finetuned GLM retrievers generalize better out of domain, achieving SOTA on two multi-hop benchmarks. On in-domain multi-hop QA datasets they remain comparable to prior work, with promising scaling as parameters and subgraph coverage increase. GNN-based retrievers achieve higher graph coverage with an efficient training setup, whereas the vector-search baseline excels at single-hop datasets.
5. 【2607.28229】EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents
链接:https://arxiv.org/abs/2607.28229
作者:Luigi Sigillo,Matteo Silvestri,Francesco Tabaro,Rajat Bhatnagar,Syed Irtaza Mubashar,Matt Jeffryes,Daljit Nijjer,Vittorio Perera,Ola Spjuth,Julio Saez-Rodriguez,Melissa Harrison,Fabio Petroni
类目:Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:Europe PMC, increasingly accessed, Europe PMC search, Europe PMC interface, live Europe PMC
备注:
点击查看摘要
Abstract:The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a crucial part of that need, and resources such as Europe PMC, with over 40M indexed records, are widely used to meet it. Yet these resources were not built for AI agents: they take keywords and complex syntax and return whole papers, so every agent must learn the syntax, issue several searches, and read full papers to find the evidence it needs. We introduce EMBL AI Librarian, a knowledge layer that upgrades the Europe PMC interface for AI agents: an agent asks in natural language and receives evidence that answers it. A single LLM orchestrates the whole knowledge retrieval process: it plans complementary subqueries executed by the live Europe PMC search engine, then reads the selected papers and locates the relevant evidence. We evaluate Librarian across four benchmarks: literature synthesis, claim verification, open-domain question answering, and downstream biology tasks such as protocol questions and sequence manipulation. On ScholarQABench, Librarian improves Citation F1 by more than $16$ points over strong recently published baselines. Used as the retrieval layer of an existing claim-verification pipeline, it increases agreement with expert consensus; and on the open-form LitQA2 benchmark, a GPT-5.4 agent scores about $8$ points higher when grounded in Librarian than with web search. Overall, our results show that equipping life-science agents with the Librarian knowledge layer improves performance across a range of tasks. We release our code publicly at this https URL
6. 【2607.28136】Extended Depth-First Representations of $k^2$-trees
链接:https://arxiv.org/abs/2607.28136
作者:Gabriel Carmona,Paolo Ferragina,Giovanni Manzini,Francesco Tosoni
类目:Data Structures and Algorithms (cs.DS); Information Retrieval (cs.IR); Performance (cs.PF)
关键词:lossless compression formats, study static, operational efficiency, lossless compression, layouts
备注: 44 pages, 7 figures, 18 tables
点击查看摘要
Abstract:In this paper, we study static, computation-friendly, lossless compression formats for graphs, focusing on memory locality and operational efficiency of $k^2$-trees. We observe that their traditional level-wise layouts suffer from poor cache performance due to weak locality, especially in operations such as matrix-vector and matrix-matrix operations. To address this limitation, we propose four depth-first representations of $k^2$-trees: a plain depth-first layout (EDF-1), a balanced-parenthesis representation (BP), and their compressed variants (CEDF and CBP). We further introduce a linear-time compression method based on suffix and LCP arrays to identify and compress identical subtrees. We experimentally evaluate the execution time, the disk space, and the peak-memory usage of our approaches against classical level-wise $k^2$-trees and DFUDS-based representations across two real and one synthetic dataset (i.e., Web Graphs, Wikidata, and random adjacency matrices) over the above linear-algebra operations. Results show that our depth-first layouts are competitive and often superior than known approaches: CEDF achieves the best compression in most settings, EDF-1 and CEDF reduce the peak memory usage consistently, and performance varies by workload, with different layouts excelling in different operations and data regimes. Overall, this work demonstrates that depth-first layouts of $k^2$-trees provide a practical and efficient alternative to traditional layouts, improving both compression and computational performance in matrix operations.
Comments:
44 pages, 7 figures, 18 tables
Subjects:
Data Structures and Algorithms (cs.DS); Information Retrieval (cs.IR); Performance (cs.PF)
MSC classes:
68P05, 68P30, 68R10, 65F50
ACMclasses:
E.1; E.2; E.4; F.2.2
Cite as:
arXiv:2607.28136 [cs.DS]
(or
arXiv:2607.28136v1 [cs.DS] for this version)
https://doi.org/10.48550/arXiv.2607.28136
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
7. 【2607.28129】Face and Voice Cross-modal Association with Learning Convex Feature Embedding
链接:https://arxiv.org/abs/2607.28129
作者:Taewan Kim,Jiwoo Kang
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:deep learning, association learning, cross-modal, cross-modal association tasks, learning
备注:
点击查看摘要
Abstract:Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull. Moreover, the incorporation of cross-modal attention mechanisms with convex embedding techniques represents a highly effective strategy for the attenuation of false positives and false negatives, accomplished via the minimization of inter-class discrepancies. We exhaustively evaluated our method for cross-modal verification, matching, and retrieval tasks on the large-scale VoxCeleb dataset. Extensive experimental results demonstrate that the proposed method achieves notable improvements over existing state-of-the-art methods.
8. 【2607.28070】CCFormer: Efficient Cross-Field Interaction and Hierarchical Sequence Compression for Industrial Recommendation at Tencent
链接:https://arxiv.org/abs/2607.28070
作者:Yunlong Wang,Huizhe Zhang,Haonan Hu,Yudong Li,Bing Wen,Jianchao Tu,Chengxiang Zhuo,Zang Li
类目:Information Retrieval (cs.IR)
关键词:predictable scaling laws, Recent studies, increasing sequence length, demonstrated that sequential, built upon self-attention
备注:
点击查看摘要
Abstract:Recent studies in industrial recommendation systems have demonstrated that sequential recommendation models built upon self-attention can benefit from predictable scaling laws by increasing sequence length and model capacity. However, practical recommender systems impose strict latency and resource constraints, making it challenging to balance computational overhead with fine-grained feature interaction. In this paper, we propose CCFormer, an efficient Transformer backbone that unifies cross-field feature interaction and compressed long-sequence modeling for industrial recommendation. Specifically, CCFormer combines feature-field separated cross attention with long-sequence subspace token mixing to exploit long-term preference signals across heterogeneous feature domains. A hierarchical sequence compression strategy with progressively expanded receptive fields enables efficient long-sequence modeling with reduced information loss. Extensive experiments on two public benchmarks and a large-scale industrial dataset demonstrate that CCFormer consistently outperforms state-of-the-art baselines. Online A/B tests in a video recommendation scenario and an advertising ranking scenario at Tencent further validate its industrial practicality, yielding a 3.57% CTR gain and a 1.71% advertising revenue lift, respectively, while accelerating model training by 2.21x over the strong HSTU baseline. CCFormer has been fully deployed in Tencent's production recommendation system, serving the main traffic of both scenarios.
9. 【2607.28055】VIG-RL: Learning to Search and Insert for Verified Image Grounding
链接:https://arxiv.org/abs/2607.28055
作者:Qinhan Yu,Jun Guang,Chong Chen,Wentao Zhang
类目:Information Retrieval (cs.IR)
关键词:Verified Image Grounding, requires Verified Image, responses requires Verified, Image Grounding, providing reliable interleaved
备注:
点击查看摘要
Abstract:In knowledge-intensive scenarios, providing reliable interleaved text-image responses requires Verified Image Grounding (VIG): the precise integration of retrieved authentic visual evidence. Existing retrieval-augmented frameworks predominantly rely on decoupled, static pipelines, inherently failing to dynamically reason about when external knowledge is required and where visual assets should be contextually inserted. To bridge this gap, we propose VIG-RL, an autonomous agentic framework that formulates the search-selection-insertion workflow as an active decision-making process. Operating within a dynamic ReAct-style loop, VIG-RL is optimized via reinforcement learning, guided by a composite reward system that holistically evaluates the agent's step-by-step tool execution and final multimodal alignment. Extensive evaluations demonstrate that VIG-RL establishes a new state-of-the-art, significantly outperforming existing static baselines.
10. 【2607.27959】FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
链接:https://arxiv.org/abs/2607.27959
作者:Bohan Hou,Haoqiang Lin,Xuemeng Song,Haokun Wen,Meng Liu,Yupeng Hu,Xiangyu Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:Multimodal Large Language, Large Language Models, Large Language, strong generalizable multimodal, generalizable multimodal processing
备注:
点击查看摘要
Abstract:Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR). Therefore, in this work, we propose an automated fine-grained multimodal quintuple dataset construction pipeline and a novel two-stage fine-grained multimodal fine-tuning strategy. The dataset generation pipeline produces a comprehensive CIR dataset with fine-grained image captions and modification text, facilitating fine-grained context modeling. Beyond the previously entangled fine-tuning paradigm, our approach separates the fine-tuning process into two distinct stages: (1) fine-grained context reasoning-oriented fine-tuning and (2) fine-grained retrieval-oriented fine-tuning. These stages aim to sequentially enhance the model's context understanding and query-target alignment capabilities, thereby improving retrieval performance. Extensive experiments across five datasets encompassing diverse and complex image retrieval tasks demonstrate the remarkable superiority of our method over existing approaches in zero-shot retrieval settings, even with a more lightweight MLLM backbone compared to those methods.
11. 【2607.27955】SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions
链接:https://arxiv.org/abs/2607.27955
作者:Jennifer D'Souza,Sameer Sadruddin,Anisa Rula,Ana Bossler,Andrés Fullana,Enric Bas,Syed Ather,Defne Circi,Anlan Chen,L. Catherine Brinson,Alyssa Columbus,George Demetriou,Dongjun Jeong,Tarun Kumar,Frank Krüger,Sascha Genehr,Kai Budde-Sagert,Anamaria Leonescu,Francesco Lodola,Chiara Florindi,Gagana Balasubramanya Murthy,Samson Oluwapelumi Olagbile,Nazia Riasat,Yan Sha,Kevin Shen,Shaokai Yang
类目:Digital Libraries (cs.DL); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR)
关键词:heterogeneous article discourse, spanning Biology Biotechnology, dispersed across prose, supplementary files, details needed
备注: 25 pages, 9 figures, Submitted for peer review to Nature Scientific Data
点击查看摘要
Abstract:Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automation dispersed across prose, tables, figures, protocols, and supplementary files. We present the first release of this http URL, a multidisciplinary collection of 16 expert-annotated schemas spanning Biology Biotechnology, Materials Chemistry, Imaging Measurement, Physics, and Psychology. Each schema defines reusable fields for describing process instances, including inputs, outputs, materials, instruments or software, parameters, conditions, procedural steps, measurements, and provenance-related information. The schemas were created through a human-in-the-loop schema-mining workflow in which large language models generated candidate structures from process specifications, scientific articles, and expert feedback, followed by domain-expert construction of final master schemas. The dataset contains final schemas in JSON Schema and SHACL formats, intermediate model-generated schemas, expert-feedback records, source-paper metadata, community-development materials, and analysis scripts. Technical validation assessed schema structure, development provenance, expert review, and syntactic conformance. The collection supports structured annotation, metadata enrichment, scientific knowledge graphs, information extraction, semantic publishing, and cross-study comparison.
12. 【2607.27944】Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation
链接:https://arxiv.org/abs/2607.27944
作者:Long Zhang,Hao Jiang,Sheng Yu,Fei Pan,Peng Jiang,Kun Gai
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:large language models, frameworks largely follow, advanced ID-based recommendation, generation frameworks largely, existing SID generation
备注:
点击查看摘要
Abstract:While large language models (LLMs) have advanced ID-based recommendation through Semantic ID (SID) modeling, existing SID generation frameworks largely follow a single-representation-then-quantization paradigm. This design faces two bottlenecks: semantic entanglement mixes heterogeneous attributes, such as geography, brand, and category, causing information loss during quantization, low-quality SIDs, and severe collisions; moreover, black-box representation learning provides neither explicit attribute semantics nor clear geographic or semantic meanings for SID positions. These limitations weaken both retrieval reliability and the ability to diagnose or control SID generation. We propose Interpretable Representation via LLM-Driven Generative Disentanglement for Local-Life Service Recommendation (LGRID). LGRID introduces a generative disentanglement paradigm through an Encode - Disentangle - Align - Quantize pipeline. It first uses joint LLM encoding to preserve cross-attribute geographic-semantic dependencies, rather than encoding fields independently. A Structured Disentangled Block then routes hidden states into attribute-aligned slots for geographic and semantic factors. Synergistic Alignment Learning makes these slots both generatively decodable and discriminative for retrieval, while Dual-Stream Residual Quantization separately discretizes the two streams into compact SIDs with explicit attribute correspondence. This design yields interpretable SIDs with positions grounded in item attributes and local-service semantics. Experiments on Kuaishou and Foursquare show that LGRID consistently outperforms strong SID baselines, achieving up to a 5.44 percent relative AUC gain. It also achieves over 99 percent attribute-decoding accuracy for coarse geographic fields and reduces the full-SID collision rate to 39.9 percent, compared with 97.0 percent for LGSID.
13. 【2607.27789】From Understanding to Action: Feedback-Grounded Policy Discovery for Generative Recommendation
链接:https://arxiv.org/abs/2607.27789
作者:Zhi Chen,Minmao Wang,Xingchen Liu,Haoqiang Liang,Huihuang Lin,Likang Wu,Hongke Zhao,Yulong Wang,Shijie Yi,Fei Pan,Peng Jiang
类目:Information Retrieval (cs.IR)
关键词:generative recommenders enable, efficient next-item generation, recommenders enable efficient, enable efficient next-item, captures behavioral co-occurrence
备注:
点击查看摘要
Abstract:Semantic-ID-based generative recommenders enable efficient next-item generation, but their item-level supervision mainly captures behavioral co-occurrence and local transitions. Large language models (LLMs) can complement these models by reasoning over heterogeneous interaction histories to understand the user's current demand. However, LLMs are not inherently trained with recommendation-specific outcome feedback, and linguistically plausible reasoning therefore does not necessarily lead to effective recommendation decisions. We term this mismatch the Understanding-Action Gap. Accordingly, we distinguish intent knowledge, which captures the user's current demand, from policy knowledge, which specifies the recommendation direction and rejection boundary under that demand. To bridge this gap, we propose a feedback-driven agent framework that first induces task-oriented intent and then discovers recommendation policies according to their incremental utility over an intent-only baseline. Candidate policies are evaluated and refined using outcome-derived feedback rather than linguistic plausibility. We further transfer the resulting intent and policy knowledge into two latent tokens of a lightweight Semantic-ID generator through dual-space relational distillation, enabling LLM-free online inference. Experiments on public benchmarks show consistent improvements over baselines, while large-scale online A/B tests achieve gains of 4.506% in Revenue and 4.621% in ADVV.
14. 【2607.27766】Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning
链接:https://arxiv.org/abs/2607.27766
作者:Xinyu Luo,Hui Liu,Yihua Shao,Junyi Yang,Arindam Basu,Haoliang Li
类目:Computation and Language (cs.CL); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:On-device in-context learning, downstream model inference, On-device in-context, in-context learning, relies on pre-inference
备注: Under review
点击查看摘要
Abstract:On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local memories under limited computation, memory, and data-exposure budgets. We propose Conditional Retrieval Alignment (CoRA), a gradient-free framework that converts a frozen encoder into a task-conditioned retriever using paired candidate inputs and outputs. CoRA selects complementary encoder layers, constructs an output-derived conditioning space from candidate memory, and aligns candidate input representations to this space through closed-form ridge regression. Low-rank factorization then produces a compact retrieval basis where candidate outputs are used only during offline index construction, whereas query-time retrieval requires only the query input and precomputed index. We show that CoRA's rank-constrained basis is the optimal low-rank compression of the output-conditioned fitted representation, and derive an exact two-pass streaming construction that avoids materializing the full fitted matrix. We further extend the framework to multimodal exemplar retrieval by incorporating visual representations into the conditioning and retrieval spaces. Experiments across ten textual datasets and four multimodal benchmarks with Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B, and Qwen3.5-2B, as well as end-to-end Raspberry Pi~5 deployment demonstrate that CoRA supports effective task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls.
15. 【2607.27763】DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
链接:https://arxiv.org/abs/2607.27763
作者:Bowen Wang,Youwen Zhang,Ritesh Mehta
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:Concept Unique Identifiers, UMLS Concept Unique, assigning UMLS Concept, generating natural-language captions, Caption Prediction
备注: 21 pages, 9 figures
点击查看摘要
Abstract:We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.5790$ and a secondary $F_1$ of $0.9657$. In parallel, we submitted a training-free KNN retrieval pipeline over frozen BiomedCLIP embeddings, which reached a primary $F_1$ of $0.5780$ and a secondary $F_1$ of $0.9599$-essentially matching the fine-tuned ensemble on the primary track at a fraction of the cost. For Task 2, our submissions included a fine-tuned Gemma-3 27B model (overall $0.3571$, ranking third in the official submission), a fully fine-tuned BLIP pipeline with custom Vizwins merging ($0.3564$), and a zero-shot MedGemma-4B run with a PubMed-style prompt ($0.3186$), spanning a wide range of model scales and training costs. Code: this https URL.
16. 【2607.27760】Hierarchical Latent Reasoning for LLM-based Recommendation
链接:https://arxiv.org/abs/2607.27760
作者:Peiyu Hu,Siying Gu,Weihai Lu,Zhuodong Liu,Yuntian Tang,Jiahao Liang,Yiying Xie,Jiang Rong,Zhaokai Luo,Zhiyong Wang,Jia Wang
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:Large Language Models, Large Language, Language Models, contextual modeling capabilities, shown strong potential
备注:
点击查看摘要
Abstract:Large Language Models (LLMs) have shown strong potential for recommendation by leveraging their semantic understanding and contextual modeling capabilities. Recent studies further introduce reasoning mechanisms to improve user preference modeling. However, explicit natural-language reasoning incurs substantial inference overhead, whereas existing latent reasoning methods mainly focus on generating or verifying intermediate states, leaving their layer-wise preference roles and contributions insufficiently characterized. We propose HiLaR, a Hierarchical Latent Reasoning framework with layer-aware reinforcement optimization for LLM-based recommendation. HiLaR constructs temporal-guided hierarchical user preference representations, aligns them with multiple LLM latent reasoning states, and organizes the reasoning process from broad preferences to fine-grained current intents. To further optimize the reasoning trajectory, HiLaR combines final recommendation feedback with layer-aware process rewards derived from the marginal target-likelihood gain of each state. Experiments on four Amazon benchmark datasets show that HiLaR generally outperforms strong sequential, generative, and LLM-based recommendation baselines. Ablation and sensitivity analyses further verify the contribution of hierarchical representation learning, latent alignment, and process-level optimization. Our code is available in this https URL.
17. 【2607.27748】A Structured Knowledge Infrastructure for Domain-Specific Data Asset Discovery
链接:https://arxiv.org/abs/2607.27748
作者:Mengdi Chen,Yuanxin Huang,Yulin Jiang,Wei Sun
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Databases (cs.DB)
关键词:generic RAG retrieves, Enterprise data analytics, prevent metric misinterpretation, analytics agents face, data analytics agents
备注: 6 pages, 2 figures, 2 tables. Submitted to DAI 2026 Industry Track
点击查看摘要
Abstract:Enterprise data analytics agents face two structural failures: generic RAG retrieves the wrong asset (Hit@10=19.1%) and delivers no usage knowledge to prevent metric misinterpretation---stemming from four root causes (C1--C4) ranging from semantic gap and entity ambiguity to schema drift and asset-usage gap. We present a two-layer solution deployed in the commercial advertising data warehouse at Xiaohongshu (5,300+ Hive tables, 14 domains). A three-tier dual-purpose knowledge base (179 documents, eight-section annotation template) serves both retrieval and generation, with a closed-loop refresh pipeline maintaining day-level freshness (one yes/no approval, 30s hot-reload). The Graph-Guided Retriever (GGR) uses a 2,859-node knowledge graph as a candidate gate with intent routing to deliver 71.6x token reduction. The Scene-Aware Ranker (SAR) applies 19-class entity recognition and explicit scenario annotations; negative knowledge alone contributes 25 percentage points of Hit@10 gain. On two 100-question benchmarks, Hit@10 rises from 19.1% to 96.6% (+77.5pp) and knowledge coverage from 56% to 77%, at 4.84--5.33s end-to-end latency.
18. 【2607.27744】ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation
链接:https://arxiv.org/abs/2607.27744
作者:Yuxin Chen,Liang Luo,Buyun Zhang,Jian Jiao,Boda Li,Haoyu Wang,Tongyi Tang,Ao Cai,Zijian Shen,Zhengkai Zhang,Wenyi Xie,Ryan Dick,Han Liu,Neng Shi,Bin Yu,Jianbo Xiao,Shuyao Bi,Hongtao Yu,Yuanwei Fang,Zhuoran Zhao,Sijia Chen,Yang Chen,Shuqi Yang,Qianru Li,Zikun Liu,Wei Ling,Sihan Zeng,Longhao Jin,Jiaxin Lu,Yinbin Ma,Jiawei Li,Yichen Ruan,Yong Ler Lee,Birmingham Guan,Zijian Li,Jianbo Sun,Zhengyu Zhang,Zeliang Chen,Xiaohan Wei,Yuchen Hao,GP Musumeci,Venkatesh Ranganathan,Yantao Yao,Chunqiang Tang,Wenlin Chen,Santanu Kolay,Ellie Dingqiao Wen
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
关键词:Modern recommendation models, production cost constraints, cost constraints cap, Modern recommendation, constraints cap
备注:
点击查看摘要
Abstract:Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far systems can scale. In this work, we propose Request-Oriented Compute Sharing (ROCS), a modeling and inference paradigm that exploits a unique property of recommendation inference: each user request is evaluated against many candidates, while request-side features are shared across candidates. ROCS defers request-candidate interactions as late as possible, isolates candidate-dependent representations, and evaluates substantial portions of the model once per request rather than once per candidate, significantly improving inference efficiency while maintaining or improving prediction quality. To realize this paradigm, we develop Generalized Layer Masking (GLM) to enforce candidate isolation in feature-interaction architectures, and Deep Cross Attention (DCA) to extend request-oriented sharing to sequence architectures. To support efficient GPU deployment, we co-design In-Kernel Broadcast Optimization (IKBO) that significantly accelerates ROCS model execution. Experiments on public benchmarks show that ROCS consistently improves the quality-efficiency tradeoff across recommendation backbones. On production-scale workloads, ROCS achieves up to a 3x QPS improvement on retrieval models without quality degradation and a 0.5% relative LogLoss improvement with a 50% QPS gain on a short-form video ranking model. ROCS has been deployed across large-scale recommendation systems spanning ads and organic surfaces, retrieval and ranking stages, and more than two orders of magnitude in inference complexity, delivering significant online gains at reduced infrastructure cost.
Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR)
Cite as:
arXiv:2607.27744 [cs.LG]
(or
arXiv:2607.27744v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2607.27744
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
19. 【2607.27739】Measuring Alignment With Reader Highlights Net of Position and Length
链接:https://arxiv.org/abs/2607.27739
作者:Kazuki Nakayashiki,Keisuke Watanabe
类目:Information Retrieval (cs.IR); Computation and Language (cs.CL); Human-Computer Interaction (cs.HC)
关键词:Context compression discards, language model reads, downstream task accuracy, Context compression, language model
备注: 15 pages, 7 tables. Analysis code and de-identified artifacts included as ancillary files; five of six scripts reproduce the paper's numbers from the shipped artifacts alone. Reports claims from our own prior work that this corpus does not reproduce, and lists twelve claims withdrawn during internal adversarial review in Appendix A
点击查看摘要
Abstract:Context compression discards most of a document before a language model reads it, and is normally evaluated by downstream task accuracy - which makes another model the judge of what mattered. Naturalistic social highlighting offers a non-circular reference: many people independently marking passages on the same page. But the obvious metric, the fraction of crowd-marked sentences a compressor keeps, is confounded twice: crowd marks are front-loaded and crowd-marked sentences are longer, so any method favouring early or long sentences scores well regardless of readers. We remove both by matching each marked sentence against unmarked sentences of the same document at equal relative depth and equal within-document length rank, and we calibrate every estimator on synthetic nulls built from position and length alone - a step that matters, since depth-only stratification returns a false positive on 20-36% of nulls containing no effect. On 120 web documents (at least 12 independent readers each), a language-model importance ranking keeps 38.4% of crowd-marked sentences against 19.9% of their matched neighbours: an enrichment of +0.196 [+0.148, +0.239], at p = 0.0005 under an exact randomization test that assumes nothing about clustering, and replicated cross-vendor. Naive truncation, whose keep rule is position, correctly falls to +0.003. To give the number a scale: scored identically, on the same budget, against a crowd label recomputed to exclude them, a single human reader reaches +0.182 - indistinguishable from GPT-5.4 (+0.002 [-0.081, +0.088]) and below Claude Opus 5. Classical methods are not null - Luhn's 1958 heuristic reaches +0.088 - so reader selection is partly recoverable by counting words; conditioning additionally on lexical centrality removes only 0.010, so the agreement is not centrality. We also report that a claim in our own prior work does not reproduce on this corpus.
20. 【2607.27682】Restoring Collaborative Signals in Semantic-ID Generative Recommendation via Personalized Natural Language
链接:https://arxiv.org/abs/2607.27682
作者:Changjiang Han,Qingyang Li,Yaqiang Zang,Jikun Kang,Pinghua Gong,Xue Liu,Bowei He
类目:Information Retrieval (cs.IR)
关键词:Making LLM-based generative, Making LLM-based, LLM-based generative recommendation, unsolved goal, SID
备注: 8 pages, 4 figures
点击查看摘要
Abstract:Making LLM-based generative recommendation models stronger and more personalized through natural language and explicit reasoning is a widely anticipated yet still unsolved goal. Such models cast recommendation as autoregressively generating an item's semantic-ID (SID), a short tuple of discrete codes, so that recommending well reduces to emitting the right SID. In this setting the model verbalizes its knowledge poorly, and text and SID tokens live in misaligned embedding spaces. Deep reasoning therefore rarely turns into a correct SID, and enabling explicit "thinking" often gives no gain or even hurts. The deeper cause is that a compact SID cannot hold content and collaborative signal at once: the two compete, and collaboration loses. Because a mis-predicted SID is a wrong recommendation, this caps accuracy directly. Costly multi-round training barely helps, and few methods try to supply the missing signal at inference time. What is missing is a reliable channel that carries collaborative signal into SID generation. We therefore propose a framework, guided by personalized natural language, that adds hierarchical collaborative cues as the model generates, without altering the backbone or retraining the SIDs. Rather than mapping language onto SIDs directly, it uses language to attach analyzable links between collaborative patterns and their audiences, restoring the collaborative signal that SIDs miss. The result is consistent gains in recommendation accuracy, grounding generation in collaborative structure at inference time rather than relying on explicit reasoning or retraining.
21. 【2607.27647】LoopMemGR: From Behavior Logs to Evolving Memory for Generative Recommendation
链接:https://arxiv.org/abs/2607.27647
作者:Hui Qian,Changfa Wu,Chang Liu,Binbin Cao,Jian Wu,Yuliang Yan,Han Zhu,Bo Zheng
类目:Information Retrieval (cs.IR)
关键词:discrete Semantic IDs, large-scale item spaces, formulates next-item prediction, conditional autoregressive generation, Semantic IDs
备注:
点击查看摘要
Abstract:Generative recommendation formulates next-item prediction as conditional autoregressive generation over discrete Semantic IDs, enabling end-to-end recommendation over large-scale item spaces. However, most existing methods follow a history-as-context paradigm that repeatedly reconstructs user preference from behavior history while discarding system-side recommendation decisions after each request. This creates an asymmetric memory: the system remembers what the user has done, but not what it has previously recommended or learned from the resulting feedback. Consequently, useful preference-validation signals, potential negative evidence, and historical exploration information cannot be directly reused across requests. To address these limitations, we propose LoopMemGR, a closed-loop recommendation experience memory framework for generative recommendation. In addition to the conventional behavior log, LoopMemGR maintains a recommendation experience log that records past recommendation--feedback trajectories. It extracts request-relevant evidence through three complementary views: the recency view captures short-term interaction dynamics, the frequency view summarizes recurring recommendation patterns, and the global view distills transferable regularities shared across users. These signals are compressed into a fixed number of experience tokens to condition the generative backbone under a bounded input budget. Extensive experiments on an industrial Taobao dataset demonstrate the effectiveness of closed-loop experience accumulation and multi-view experience extraction.
22. 【2607.27640】Dynamic Exploration Graph: A Novel Approach for Efficient Nearest Neighbor Search in Evolving Multimedia Datasets
链接:https://arxiv.org/abs/2607.27640
作者:Nico Hezel,Kai Uwe Barthel,Bruno Schilling,Konstantin Schall,Klaus Jung
类目:Information Retrieval (cs.IR)
关键词:Approximate Nearest Neighbor, Nearest Neighbor Search, Approximate Nearest, Nearest Neighbor, Neighbor Search
备注:
点击查看摘要
Abstract:Approximate Nearest Neighbor Search (ANNS) represents a fundamental problem in various applications (image-search, recommendation systems). While graph-based algorithms have demonstrated a good balance between search accuracy and time, handling dynamic datasets, where data points are continuously added or removed, remains a challenge. This paper introduces the Dynamic Exploration Graph (DEG), an extension of the continuous refining Exploration Graph, which retains high search efficiency for static dataset while adding essential support for dynamic data. At the core of the DEG design are two key innovations: a novel vertex deletion algorithm which guarantees graph connectivity and a data distribution-agnostic method for graph expansion. Through these mechanisms, the DEG maintains a balanced and well-connected structure, even under continuous data alterations. Empirical experiments in both streaming and online scenarios demonstrate the superior performance of the DEG, surpassing existing dynamic graph algorithms in terms of construction time and search efficiency. Although optimized for dynamic datasets, the DEG delivers results as good as current state-of-the-art approaches for static dataset, underscoring its broad applicability.
23. 【2607.27623】An Exploration Graph with Continuous Refinement for Efficient Multimedia Retrieval
链接:https://arxiv.org/abs/2607.27623
作者:Nico Hezel,Kai Uwe Barthel,Konstantin Schall,Klaus Jung
类目:Information Retrieval (cs.IR)
关键词:Approximate Nearest Neighbor, Nearest Neighbor Search, Approximate Nearest, Nearest Neighbor, feature vectors continue
备注:
点击查看摘要
Abstract:As datasets and the dimensionality of feature vectors continue to grow, Approximate Nearest Neighbor Search (ANNS) in large multimedia databases becomes increasingly relevant. Graph-based approaches have demonstrated to offer the best trade-off between retrieval precision and search time. Despite their ability to deliver search times several orders of magnitude faster than exact search techniques, existing methods suffer from slow constructions speeds or high memory requirements. This paper presents a "continuous refining Exploration Graph" (crEG), a novel approach for rapidly constructing a compact exploration graph with state-of-the-art search performance. Additionally, it provides the ability to enhance its effectiveness even further through an optional edge optimization algorithm. Both algorithms are specifically designed to produce and operate on undirected graphs with even degrees and guarantee graph connectivity at any time, a property particularly valuable for "exploratory search", where the query is part of the database elements. Although such queries provide an advantageous starting point for graph search algorithms, they have been rarely considered in the context of ANNS, yet are crucial for recommendation and exploration systems. Our experiments demonstrate high efficiency in ANNS does not necessarily translate to a good performance in "exploratory search".
24. 【2607.27577】Heterogeneous Ranking in Industrial-Scale Recommender Systems: A Case Study
链接:https://arxiv.org/abs/2607.27577
作者:Di Bai,Jintao Liu,Zhenwei Tang,Peifan Wu,Nada Al-Thawr,Luoshu Wang
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:highly homogeneous environments, video-only closed-ecosystem platforms, homogeneous environments, music-only or video-only, closed-ecosystem platforms
备注: Accepted to ACM RecSys Industry Track 2026
点击查看摘要
Abstract:Heterogeneous recommendation feeds present complex challenges that extend beyond those found in highly homogeneous environments (e.g., music-only or video-only closed-ecosystem platforms). In Google Discover, a unified feed integrates diverse content sourced from the decentralized open web, including web articles, long-form and short-form videos, user-generated content (UGC), and beyond. Different content types exhibit distinct feature densities and user interaction patterns. Building a unified ranking model that sustains high performance across such heterogeneity, while avoiding negative transfer or majority bias, remains a significant industrial challenge. This paper presents an end-to-end case study on the industrial-scale multi-task ranking of heterogeneous feeds, grounded in real-world deployment. We introduce HA-MoE, a heterogeneity-adaptive multi-gated mixture-of-experts architecture that incorporates explicit heterogeneity context into both gating networks and expert representations. This approach enables effective specialization without significantly increasing operational overhead. To support reliable deployment, we introduce LENS, a lightweight observability framework that provides interpretable diagnostics of expert specialization and tracks this functional heterogeneity across continuous retraining. We evaluate our method using Dual-Level AUC (DL-AUC), a heterogeneity-aware evaluation metric that combines global ranking performance with cross-segment ranking correctness. Offline evaluations on a large-scale industrial dataset demonstrate consistent improvements over baseline models. Furthermore, online A/B testing confirms gains in feed activity and exploration metrics. Together, offline and online results validate the effectiveness of our approach for managing heterogeneity in industrial-scale recommender systems.
Comments:
Accepted to ACM RecSys Industry Track 2026
Subjects:
Information Retrieval (cs.IR); Machine Learning (cs.LG)
Cite as:
arXiv:2607.27577 [cs.IR]
(or
arXiv:2607.27577v1 [cs.IR] for this version)
https://doi.org/10.48550/arXiv.2607.27577
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)
Journalreference:
Proceedings of ACM RecSys 2026
Related DOI:
https://doi.org/10.1145/3773078.3831848
Focus to learn more
DOI(s) linking to related resources</p>
25. 【2607.27523】Hierarchical Reranking for Scalable Financial RAG System
链接:https://arxiv.org/abs/2607.27523
作者:Joohyun Lee,Sungwoo Hong
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
关键词:Analyzing financial documents, macroeconomic reports demands, reports demands expert, demands expert reasoning, Analyzing financial
备注: 8 pages, 1 figure, 3 tables. Accepted at FinLLM @ IJCAI-ECAI 2026
点击查看摘要
Abstract:Analyzing financial documents such as 10-K filings, tabular disclosures, and macroeconomic reports demands expert reasoning and extensive time. However, existing Retrieval-Augmented Generation systems often struggle to process hybrid text-table structures or the massive scale of financial documents. To address these challenges, we propose Hierarchical Reranker, a RAG framework designed to improve retrieval performance and generative reliability across large-scale financial datasets. The system integrates three key innovations: Pre-Retrieval Optimization, enhancing query clarity and search efficiency through normalization, keyword expansion, and table transformation; Hierarchical Reranker Architecture, improving retrieval precision through a two-stage ranking mechanism; and Long-Context Management, preserving reasoning accuracy through adaptive input partitioning and fusion under extensive contexts. Across multiple benchmarks, including FinQA, FinanceBench, and ConvFinQA, the proposed system achieved an NDCG@20 score of 0.7918 and demonstrated superior factual consistency. Its robustness was further validated by achieving second place in the ACM-ICAIF '24 FinanceRAG Challenge. This work presents a deployable, domain-optimized RAG pipeline that enhances both the accuracy and scalability of financial reasoning, paving the way for automated audit reporting and quantitative investment analysis. The source code will be made publicly available on GitHub upon acceptance.
26. 【2607.27475】OneShot: Index-in-Ranking with Neural Scoring for Large-Scale Retrieval
链接:https://arxiv.org/abs/2607.27475
作者:Ziwei Li,Shuyao Li,Xufeng Cai,Xue Zou,Yiming Ma,Huiting Lu,Wujie Yan,Zhichen Zhao,Yang Lu,Zhe Wang,Rui Luo,Zhengyu Su,Dan Zhang,Ji Liu
类目:Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:primary stage responsible, modern recommendation systems, primary stage, stage responsible, responsible for filtering
备注:
点击查看摘要
Abstract:In modern recommendation systems, retrieval serves as a primary stage responsible for filtering billions of candidate items down to thousands prior to refined ranking. To make this massive search effective and efficient, the system relies on ranking accuracy and indexing efficiency. However, these two objectives are traditionally misaligned: while the former optimizes for the alignment between ranking predictions and user behavior, the latter optimizes for a structural grouping of item representations which enables fast search among billions of candidates. Thus, despite extensive efforts to scale up interaction modeling for retrieval, they remain fundamentally limited by the structural misalignment between the ranking objectives and the proximity-learned index. In this work, we address this long-standing dichotomy by proposing a new holistic retrieval framework, OneShot. It is an end-to-end, in-model index learning framework that natively aligns index learning with ranking objectives. Using this joint learning as a structural foundation, OneShot pushes the boundaries of retrieval expressiveness by scaling interaction modeling with neural scoring beyond the persistent dot-product bottleneck. OneShot is fully deployed in Instagram's industrial short-video recommendation system, driving significant wins in user daily sessions, engagement, and time-spent. Additionally, OneShot achieves a $20\%$ recall gain at the operational ranking volume and a 10x efficiency improvement at an equivalent recall level.
27. 【2607.17751】MagicSelector: Joint Optimization for Agent Tool Selection via Counterfactual Decomposition and Progressive Reranking
链接:https://arxiv.org/abs/2607.17751
作者:HONOR Agentic Search Team:Zhengzong Chen,Lei Tang,Lijun Liu,Chuandi Jiang,Fan Yang,Keyun Chu,Chu Zhao,Shihao Liu,Minghang Li,Bo Liang,Can Wen,Hailong Wu,Jingnan Ju,Mian Liu,Nengbin Zhang,Peiqiang Wang,Penghe Nie,Qinhui Gu,Sijia Lv,Siqi Chen,Wei Zhang,Yang Xu,Yuhao Qian,Yuxiang Zhang,Zeng Cheng,Zhen Wang,Zuan Chen,Yuanyuan Zhao,Fei Huang
类目:Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
关键词:joint optimization framework, optimization framework integrating, integrating Counterfactual task, framework integrating Counterfactual, Counterfactual task decomposition
备注:
点击查看摘要
Abstract:We present MagicSelector, a joint optimization framework integrating Counterfactual task decomposition, Progressive reranking, and Dynamic Top-K, designed to address the fundamental challenges of tool retrieval in agents. MagicSelector is a specialized framework capable of translating ambiguous user instructions into executable atomic subtasks and guiding high-precision tool retrieval, effectively mitigating redundant noise and severe context distraction in out-of-domain (OOD) scenarios. We empower MagicSelector with these capabilities through three key contributions: (1) a preference-guided counterfactual task decomposition mechanism that utilizes a counterfactual reward to quantify the marginal causal gain of decomposition on retrieval ranking, effectively imposing fine-grained structural supervision on logical coherence; (2) a progressive tool reranking method driven by self-distillation hard negative mining, which optimizes both point-wise and list-wise relevance to enhance fine-grained discrimination among highly similar tools; and (3) a dual semantic boundary-aware dynamic Top-K strategy that adaptively monitors reranking score cliffs and inter-tool semantic shifts to dynamically truncate the candidate list, maximizing relevant tool recall while filtering long-tail noise. Evaluated on MTDTool, the first task decomposition benchmark we constructed tailored for mobile multi-turn interactions with process-level annotations, MagicSelector yields promising performance. Extensive experiments demonstrate that MagicSelector significantly outperforms state-of-the-art methods in terms of tool retrieval accuracy, OOD generalization capability, and overall token efficiency, thereby demonstrating the effectiveness of our proposed framework.
计算机视觉
1. 【2607.28627】ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
链接:https://arxiv.org/abs/2607.28627
作者:Yao Xiao,Reuben Tan,Zhen Zhu,Yuqun Wu,Jianfeng Gao,Derek Hoiem
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:GPU memory constraints, visual context poses, infeasible under GPU, GPU memory, Long visual context
备注: Code: [this https URL](https://github.com/avaxiao/ReToken)
点击查看摘要
Abstract:Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints. We present ReToken, a single learnable embedding trained as an explicit retrieval target that selects a sparse set of query-relevant visual tokens from a pre-filled visual KV cache. Trained on only a small image-QA dataset, ReToken yields consistent gains across image and video benchmarks: on Visual Haystacks it improves Qwen3VL-8B by 13.4 points and InternVL3.5 by 12.4 points (20% relative), and on LVBench it transfers zero-shot to long video for an 8.0-point gain with Qwen3VL-8B. Thanks to its lightweight design, both training and long-video inference fit on a single H100. Code is available at: this https URL
2. 【2607.28625】ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
链接:https://arxiv.org/abs/2607.28625
作者:Yukang Cao,Haozhe Xie,Beichen Wen,Runmao Yao,Yinghao Liu,Yue Huang,Zhichao Liao,Yunxiang Wang,Haiheng Liu,Xingshun Tian,Dawei Su,Long Zhuo,Dacheng Tao,Xiaogang Wang,Liang Pan,Ziwei Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:fundamental data bottleneck, Embodied intelligence faces, intelligence faces, faces a fundamental, Ambient Capture Engine
备注: Project Page: [this https URL](https://ace-data-engine.github.io/ACE-Data-0/)
点击查看摘要
Abstract:Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.
3. 【2607.28624】PhiZero: A World Model Built Around Physical Language
链接:https://arxiv.org/abs/2607.28624
作者:Shuyao Shang,Yuqi Wang,Ruopeng Gao,Xu Chen,Tieniu Tan,Lue Fan,Zhaoxiang Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:compact discrete representation, world model built, compact discrete, discrete representation, representation of world-state
备注: Project page: [this https URL](https://phi-zero.github.io/)
点击查看摘要
Abstract:We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
4. 【2607.28611】Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
链接:https://arxiv.org/abs/2607.28611
作者:Chongjian Ge,Hanwen Jiang,Tianyu Wang,Jiuxiang Gu,Yiran Xu,Ziwen Chen,Shaoteng Liu,Jing Shi,Yicong Hong,Zefan Cai,Hailin Jin,Hao Tan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Visual generation increasingly, generation increasingly requires, increasingly requires high-resolution, full attention prohibitive, requires high-resolution images
备注: 40 pages
点击查看摘要
Abstract:Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera, a hybrid visual diffusion backbone with a principled scaling recipe. Chimera processes text, image, and video tokens in one raster-ordered stream without positional embeddings. It combines Kimi Delta Attention (KDA) for long-context state tracking with O(N) complexity, interleaved Multi-head Latent Attention (MLA) for direct global interaction, and modality-aware short convolutions for local spatiotemporal context. Sparse Mixture-of-Experts (MoE) layers expand capacity while controlling activated compute. To scale this heterogeneous architecture, we introduce HeteroP, a module-wise scheme that transfers hyperparameters across width and depth according to each tensor's functional fan-in and model depth. HeteroP yields a consistently tuned family used to fit Chinchilla-style compute-optimal laws for activated model size, training-token count, and image-video data ratio. Guided by these laws, we train an 11B-parameter Chimera with 2B activated parameters. Experiments show three results. First, measured by pretraining diffusion loss, the dense backbone is 1.7x as compute-efficient as a matched full-attention Wan-2.1 2B baseline, while the complete system reaches 7.3x. Second, without length-specific fine-tuning, Chimera extrapolates zero-shot from 5-second training clips to 30-second videos, with only 6.5% FID degradation in the last five seconds. Third, the fitted laws show that compute-optimal image pretraining divides compute nearly evenly between activated model size and training-token count, whereas video pretraining modestly favors model size at higher budgets. These results establish a foundation for designing and scaling efficient long-context diffusion architectures.
5. 【2607.28609】OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
链接:https://arxiv.org/abs/2607.28609
作者:Qiushi Sun,Kanzhi Cheng,Yian Wang,Bowen Yang,Hang Yan,Liheng Chen,Fangzhi Xu,Zichen Ding,Nuo Chen,Jialin Cao,Xingdong Gong,Zehao Li,Kaiming Jin,Xinfeng Yuan,Zhoumianze Liu,Jingyang Gong,Zhangyue Yin,Jiahui Gao,Zhiyong Wu,Tianbao Xie,Jianbing Zhang,Ben Kao,Lingpeng Kong
类目:Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:Computer-using agents, digital world, advancing rapidly, CUA, VLM judges
备注: Work in progress
点击查看摘要
Abstract:Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at this https URL.
6. 【2607.28595】Beacon: Knowing When and How to Perform Agentic Visual Reasoning
链接:https://arxiv.org/abs/2607.28595
作者:Qixun Wang,Yang Shi,Letian Cheng,Zhuoran Zhang,Yan He,Yuqi Tang,Qi Zhang,Xinlei Yu,Ruizhe Chen,Tianrun Xu,Yuanxing Zhang,Pengfei Wan,Haotian Wang,Xianghua Ying
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Mode Adaptiveness, multimodal large language, agentic visual reasoning, inefficient reasoning paradigm, large language models
备注: 33 pages
点击查看摘要
Abstract:The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: Mode Adaptiveness (MA) and Tool Effect (TE). Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through text-only reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose Beacon, a novel agentic visual reasoning model that achieves stronger overall performance, improved Mode Adaptiveness, and genuine tool-induced performance gains. At the core of Beacon are the Necessity-Aware Adaptive Reward and the Hint-Guided Capability Expansion mechanism in the reinforcement learning stage, which respectively encourage adaptive tool invocation based on task necessity and strengthen the model's tool-use capability on the most challenging problems. Extensive experiments across diverse benchmarks demonstrate the strong overall performance of Beacon and its substantial improvements in both Mode Adaptiveness and Tool Effect.
7. 【2607.28590】VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
链接:https://arxiv.org/abs/2607.28590
作者:Kangning Zhang,Yixing Li,Shuai Shao,Qingyao Li,Zhengxi Lu,Zhiyuan Yao,Jianghao Lin,Wenxiang Jiao,Yuan Lu,Weiwen Liu,Weinan Zhang,Yong Yu
类目:Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
关键词:Multimodal on-policy distillation, Multimodal on-policy, supervising student-generated trajectories, knowledge by supervising, transfers fine-grained visual
备注: The project is accessible at [this https URL](https://github.com/DeepExperience/VAD_Multimodal_OPD)
点击查看摘要
Abstract:Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.
8. 【2607.28589】MixFrag: Fragility-Guided Mixed-Precision Post-Training Quantization for Vision Transformers
链接:https://arxiv.org/abs/2607.28589
作者:Md. Mehrab Hossain Opi,Robiul Islam Ryad,Md. Umar Faruk
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:deploying Vision Transformers, Vision Transformers, resource-constrained devices, PTQ, solution for deploying
备注:
点击查看摘要
Abstract:Post-training quantization (PTQ) has emerged as an effective solution for deploying Vision Transformers (ViTs) on resource-constrained devices. However, existing PTQ methods typically employ uniform bit-widths across transformer components, overlooking their heterogeneous sensitivity to quantization and leading to inefficient precision allocation. In this paper, we propose {MixFrag, a fragility-guided mixed-precision PTQ framework for Vision Transformers. MixFrag first estimates component-level quantization fragility by measuring the Kullback--Leibler (KL) divergence between full-precision and isolated quantized output distributions using a small calibration set. It then formulates bit allocation as a Multiple-Choice Knapsack Problem (MCKP), enabling adaptive layer-wise precision assignment under a target bit budget. Extensive experiments on ImageNet-1K across multiple Vision Transformer architectures demonstrate that MixFrag achieves competitive classification performance under practical mixed-precision settings. Furthermore, evaluations on COCO object detection and instance segmentation show that MixFrag achieves state-of-the-art performance among existing mixed-precision PTQ methods, improving the previous best method by up to 9.6 AP under the challenging MP3/MP3 setting. Additional analyses validate the proposed fragility metric and demonstrate its strong correlation with the learned bit allocation. These results establish MixFrag as an effective framework for mixed-precision post-training quantization of Vision Transformers.
9. 【2607.28581】ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation
链接:https://arxiv.org/abs/2607.28581
作者:Xiao Luo,Mingyang Du,Xin Zhou,Tianrui Feng,Xiwu Chen,Xiaofan Li,Jiangning Zhang,Dingkang Liang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:scaling model capacity, generation predominantly relies, incurs prohibitive computational, predominantly relies, relies on scaling
备注:
点击查看摘要
Abstract:High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typically requires learning geometry from scratch and overlooks the rich semantic and structural priors already encapsulated in discriminative 3D foundation models. We contend that leveraging the profound understanding of the 3D world possessed by these discriminative models can significantly reduce generative cost. To this end, we propose ROAD, a framework that reduces the training cost of 3D generation by transferring these rich discriminative priors into diffusion transformers. To address the inherent semantic-structural heterogeneity between generative and discriminative latents, we introduce a reciprocal-objective alignment strategy. This method synergizes Holistic Semantic Condensing to enforce global semantic coherence and Structural Optimal Alignment, which is formulated as a bipartite matching problem to rigorously align microscopic geometric details between disparate latent spaces. The 3D foundation model is only used for training-time supervision of alignment and is not used at inference, incurring no additional inference cost. Compared with the industrial baseline Step1X-3D, the proposed ROAD achieves highly competitive generation performance with only 1.5% of the training data and significantly reduces training costs, effectively reducing the computational overhead of high-fidelity 3D generation. Code is available at this https URL.
10. 【2607.28571】Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently
链接:https://arxiv.org/abs/2607.28571
作者:Simon Roy,Mark Bong,Giovanni Beltrame
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:Operational Earth observation, Earth observation increasingly, Operational Earth, observation increasingly calls, satellite image pairs
备注: 10 pages, 3 figures
点击查看摘要
Abstract:Operational Earth observation increasingly calls for answering queries such as ``find the image pairs where a new building appeared.'' This means searching an archive of before-and-after (bi-temporal) satellite image pairs and ranking each pair by how well it matches a natural-language description of the change. The component that performs this match, the fusion module that combines the ``before'' and ``after'' views, must be run at query time across many candidate pairs, so its speed largely sets the cost of every search. We present a controlled comparison of how to build that module. Using one fixed image encoder (a frozen CLIP model) and one training recipe for all variants, we evaluate eight designs drawn from three families: attention, state-space models (Mamba), and learned compression (our Temporal Bottleneck Fusion, TBF). Each design is tested on two benchmarks (LEVIR-CC and Dubai-CC) with ten random seeds, so the reported differences are statistically grounded. We outline three findings: first, a training-free two-stage search (a cheap difference model that shortlists candidates, followed by attention fusion that re-ranks them) matches or exceeds full-fusion recall on LEVIR-CC while cutting query cost $10$-$15\times$, with comparable R@1/R@5 on Dubai-CC; second, the linear-time scan of Mamba, attractive on paper, gives no speed benefit at the patch counts typical of vision transformers ($L{=}196$): the scan is limited by memory bandwidth, whereas attention maps cleanly onto parallel hardware; and third, compressing the fused representation (TBF) reduces parameters by $2.3\times$ and latency by $1.6\times$ for a change-only BLEU-1 cost of $0.007$, although more aggressive compression quietly discards change-relevant detail that aggregate metrics fail to reveal.
11. 【2607.28565】MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion
链接:https://arxiv.org/abs/2607.28565
作者:Yunzhan Fu,Xiangyu Shen,Yifei Sun,Yuhan Chen,Jian Wu,Hongxia Xu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:integrate complementary information, diverse imaging modalities, aims to integrate, integrate complementary, complementary information
备注: 14pages, 14 figures, accepted by ACM MM2026
点击查看摘要
Abstract:Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter. This module explicitly extracts source image features before serialization, injecting them into the network via strict dimensional alignment to effectively supplement image features. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss. This loss ensures deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion, holding significant promise for intent-driven intelligent clinical decision support systems.
12. 【2607.28538】ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs
链接:https://arxiv.org/abs/2607.28538
作者:Ruman Wang,Hangting Ye
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:Classifying pathological scars, requires distinguishing keloids, substantial acquisition variation, Classifying pathological, photographs requires distinguishing
备注:
点击查看摘要
Abstract:Classifying pathological scars from clinical photographs requires distinguishing keloids from hypertrophic scars despite limited expert-labeled data and substantial acquisition variation across hospitals. End-to-end image models remain data-dependent, whereas sending photographs to a hosted vision-language model (VLM) may conflict with local data-governance requirements and yields decisions that are difficult to reproduce and audit. We introduce ScaFE (Scar Feature Engineering), which transfers clinical knowledge from a large language model (LLM) into deterministic, executable feature programs instead of asking the model to diagnose images. A web-enabled LLM retrieves clinical evidence and synthesizes programs that measure visually assessable scar attributes. Candidate programs execute in a restricted local environment, and only aggregate validation statistics and feature-level SHAP summaries are returned for iterative repair and refinement; raw images and patient-level outputs remain local. A lightweight Random Forest then operates on the resulting structured representation. On 600 photographs from three hospitals under leave-one-site-out evaluation, ScaFE achieves 81.0% site-macro balanced accuracy, exceeding the strongest baseline, BiomedCLIP, by 10.0 percentage points. With only 10% of the development data, ScaFE retains 72.0% balanced accuracy and an 11.8-point lead. Iterative refinement also raises the executable-program rate from 66.7% to 95.0%, with verified evidence for 91.7% of the final features. These results show that LLM knowledge can support data-efficient, cross-site medical image classification through local and auditable feature programs rather than direct VLM decisions.
13. 【2607.28532】MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition
链接:https://arxiv.org/abs/2607.28532
作者:Alex Andonian,Samuel G Rodriques,Andrew D White,Siddharth M Narayanan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Markush structure, scientific literature, Markush, structure, optical chemical structure
备注:
点击查看摘要
Abstract:Chemical structures appear in patents and the scientific literature as images. For programmatic usage, such as indexing in databases or constructing machine learning model training sets, they must be transformed into line notations. The two common forms of this task are translating an image of a single molecule (optical chemical structure recognition - OCSR) and translating a Markush structure that represents a family of molecules. While prior work in the former case is quite mature, Markush structure parsing remains a challenging task. In this work, we treat both tasks as an image-to-text translation problem. We then propose OCSRGlyph, a state-of-the-art OCSR model, improving performance over prior methods by carefully considering stereochemistry. For the Markush task, we introduce MarkushGlyph, a vision-language model that reads the entire Markush structure as an image. This contrasts with prior systems, which often use multiple stages to separately process visual and text input content. Finally, we introduce a new metric for determining the accuracy of Markush structure translations, handling failure modes present in prior metrics.
14. 【2607.28526】What to Remove, What to Preserve: Dual-Ambiguity Rectification for All-in-One Image Restoration
链接:https://arxiv.org/abs/2607.28526
作者:Cencen Liu(1),Wen Yin(1),Dongyang Zhang(1),Dongmin Li(1),Shan Zhao(2),Bing Su(2),Tao He(1),Jielei Wang(1),Guoming Lu(1) ((1) University of Electronic Science and Technology of China, (2) Jiigan Technology)
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:handle diverse degradations, image restoration aims, unified framework, aims to handle, handle diverse
备注:
点击查看摘要
Abstract:All-in-one image restoration aims to handle diverse degradations within a unified framework. Existing methods commonly encode heterogeneous degradation conditions in a shared latent space, where degradation-related cues and scene content can remain entangled. We characterize the resulting challenge as dual ambiguity: semantic ambiguity in channel-wise modulation and spatial ambiguity in restoration responses, which can lead to content corruption and residual artifacts. To mitigate this issue, we propose DAR-Net, a Dual-Ambiguity Rectification Network for all-in-one image restoration. DAR-Net first introduces a Degradation Archetype Representation (DAR) module to construct a structured degradation state through simplex-constrained archetype mixture modeling. Based on this state, a Semantic Ambiguity Rectification (SeAR) module generates degradation-aware prompts to improve channel-wise conditioning in the decoder. A Spatial Ambiguity Rectification (SpAR) module further regularizes degradation-aware and complementary features toward orthogonal response subspaces, reducing spatial interference between removal and preservation cues. Extensive experiments on standard all-in-one restoration benchmarks show that DAR-Net achieves the best overall performance under both three-degradation and five-degradation settings, improving the average PSNR over the strongest competitor by 0.14 dB and 0.34 dB, respectively; it additionally shows superior performance on CDD-11 and WeatherBench.
15. 【2607.28516】Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
链接:https://arxiv.org/abs/2607.28516
作者:Bowen Liu,Shuning Wang,Xinpeng Ding,Zhiheng Wu,Bodong Du,Xiaomeng Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Long-video understanding commonly, understanding commonly compresses, commonly compresses videos, Long-video understanding, understanding commonly
备注:
点击查看摘要
Abstract:Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet making evidence available does not ensure that complementary cues across moments are integrated for answering. Our key idea is to organize selected frames into query-relevant cross-frame evidence before generation. We formulate this post-selection stage as a latent evidence interface and instantiate it with GenEvA ($\textbf{Gen}erative$ $Latent$ $\textbf{Ev}idence$ $\textbf{A}ggregation$), a distribution-guided latent evidence aggregation framework. Specifically, GenEvA uses a query-conditioned evidence distribution to focus aggregation on relevant frames, forming compact cross-frame latent evidence from their frame-specific information. Since cross-frame integration is not always needed, the same distribution determines whether to insert this latent complement. Across four benchmarks and two Video-MLLM backbones, GenEvA consistently improves matched-frame baselines. At 8 frames, it raises the four-benchmark LLaVA-Video average by $+5.2$ points and Qwen2.5-VL accuracy on LVBench by $+10.1$ points. These gains require only $0.11\%$--$0.40\%$ average video-token overhead; analyses further show task-aware allocation and benefits from Adaptive Evidence Invocation.
16. 【2607.28509】RefCaptioner: Multi-Reference Image-Grounded Video Captioning
链接:https://arxiv.org/abs/2607.28509
作者:Tengfei Liu,Yang Shi,Yuran Wang,Xiaohan Zhang,Yuqing Wen,Yuqi Tang,Qixun Wang,Zhuoran Zhang,Xuanyu Zhu,Weihong Lin,Xinlei Yu,Yujie Wei,Xinwei Long,Fengxiang Wang,Xinlong Chen,Yue Ding,Jialu Chen,Haotian Wang,Yuanxing Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:explicitly ground local, ground local visual, local visual elements, Existing video captioning, generate natural descriptions
备注: [this https URL](https://github.com/pkucs-Ltf/RefCaptioner)
点击查看摘要
Abstract:Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing $20,000$ videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.
17. 【2607.28487】AuricularWorld: Hierarchical Action-Guided World Modeling for Fine-Grained Auricular Structure Segmentation from CT Scans
链接:https://arxiv.org/abs/2607.28487
作者:Jingwen Yang,Senmao Wang,Luoyao Kang,Runmeng Cui,Keying Zhang,Yunjia Bao,Haifan Gong,Lin Lin,Haiyue Jiang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:surrounding soft tissues, Fine-grained segmentation, ear occupies, boundaries are highly, surrounding soft
备注:
点击查看摘要
Abstract:Fine-grained segmentation of auricular structures in CT is challenging because the ear occupies a small image region, cartilage boundaries are highly irregular, and interfaces between cartilage and surrounding soft tissues are often ambiguous. Clinical annotations may also include both composite structures containing cartilage and adjacent skin and their corresponding cartilage-only regions, producing nested and overlapping labels. We propose a world-model-based segmentation framework that enables iterative anatomical reasoning beyond conventional feed-forward prediction. Built on an encoder-decoder architecture, the framework introduces a deterministic recurrent state-space model into the intermediate latent space. Multi-scale encoder features and partially decoded representations are fused to form a structural observation that initializes the latent dynamics. During inference, the model performs a three-step latent rollout without ground-truth guidance. Hierarchical anatomical actions update the recurrent state and progressively refine the latent representation. The resulting latent trajectory is projected back into the decoder and combined with high-resolution features to produce the final segmentation. To learn reliable latent transitions, we introduce a balanced hierarchical action objective that addresses foreground sparsity, missing anatomical groups, and imbalance between add and remove operations. Extensive experiments show that the proposed framework consistently improves segmentation accuracy and reduces HD95 by more than 43% for small, irregular, and overlapping auricular structures in CT. These results demonstrate the effectiveness of latent world-model reasoning for challenging medical image segmentation.
18. 【2607.28483】owards Real-Time PixOOD: Efficient Anomaly Segmentation for Autonomous Vehicles
链接:https://arxiv.org/abs/2607.28483
作者:Luca de Martino,Federico Aromolo,Federico Nesti,Giorgio Buttazzo
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Real-time anomaly segmentation, Real-time anomaly, NVIDIA Jetson AGX, Jetson AGX Orin, autonomous systems
备注: 12 pages, 2 figures, 3 tables. Accepted at the Efficient Deep Learning: Methods and Applications workshop, 35th International Conference on Artificial Neural Networks (ICANN 2026)
点击查看摘要
Abstract:Real-time anomaly segmentation is essential for the safety of autonomous systems. Although recent approaches offer high accuracy, their computational cost limits their deployment on embedded hardware. This work presents an efficient and accelerated pipeline designed for both embedded and desktop platforms, targeting the autonomous driving and railway domains. The proposed approach reformulates the Neyman-Pearson scoring stage of PixOOD, a state-of-the-art out-of-distribution detection method, and deploys the full pipeline through hardware-optimized TensorRT compilation, reaching up to 182 FPS on a desktop NVIDIA RTX 4060 GPU and 75 FPS on the NVIDIA Jetson AGX Orin embedded platform, respectively 20x and 18x faster than the original baseline. The achieved results demonstrate that advanced anomaly segmentation can be efficiently deployed for onboard processing in autonomous driving and railway applications.
19. 【2607.28470】owards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation
链接:https://arxiv.org/abs/2607.28470
作者:Antonio Delgado-Rosa,David Muñoz-Valero,Enrique Adrian Villarrubia-Martin,Juan Moreno-Garcia
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:limited downlink budget, low Earth orbit, low Earth, interconnected bottlenecks, downlink budget
备注: 43 pages, 14 figures
点击查看摘要
Abstract:Airborne surveillance from low Earth orbit is hindered by two interconnected bottlenecks: nanosatellites have a limited downlink budget, yet the conventional approach still transmits terabytes of raw imagery to the ground for processing, and open satellite datasets for aircraft are scarce and severely class-imbalanced. These limitations either delay timely decision-making or prevent standard detectors from learning robust representations of rare aircraft classes. In this paper, a workflow that combines on-board inference with generative data augmentation is proposed to address both limitations jointly. Inference is executed on a 6U CubeSat equipped with a low-power edge tensor accelerator, while a diffusion model fine-tuned through low-rank adaptation generates synthetic minority-class imagery. This synthetic output is automatically annotated, pseudo-labelled, by an intermediate detector and merged with classically augmented samples. The results show that the balanced dataset increases global mean average precision from 77.9% to 82.2%, with the minority class rising from F1=0.683 to F1=0.811, and that the quantised detector fits the on-chip memory and projects 25-30 frames per second on orbit. This approach contrasts with the conventional bent-pipe architecture, in which the satellite acts as a passive data collector. Therefore, the computational tests support the proposed workflow as a decision-support tool for real-time, autonomous airborne surveillance from nanosatellites.
20. 【2607.28464】Can Vision-Language Models Reason about AI Edits in Images?
链接:https://arxiv.org/abs/2607.28464
作者:Darsha Udayanga,Pin-Yu Chen,Payel Das,Qiang Ji
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:modern generative models, manipulations increasingly difficult, difficult to identify, critical for trustworthy, modern generative
备注:
点击查看摘要
Abstract:Detection and localization of AI-tampered images are critical for trustworthy AI, yet modern generative models have made such manipulations increasingly difficult to identify. While traditional binary classifiers can detect image tampering, they lack interpretability and generalization. Vision-Language Models (VLMs) offer a promising alternative due to their strong visual understanding and reasoning capabilities; however, existing approaches typically rely on supervised finetuning with curated explanations rather than exploiting their inherent reasoning capabilities. In this work, we investigate whether VLMs can be trained to reason about AI-generated image edits using reinforcement learning (RL) rather than explicit reasoning supervision. Motivated by the success in Group Relative Policy Optimization (GRPO), an RL technique that incentivizes the model to reason by generating thinking traces prior to giving the final answer, we propose a GRPO-based training framework that utilizes simple accuracy and format rewards. Given an input image, the model produces a structured reasoning trace and predicts whether the image has been tampered with. A lightweight segmentation model is then guided by the reasoning output to generate pixel-level localization masks. Experiments across multiple image manipulation datasets demonstrate that our approach achieves competitive detection and localization performance compared to state-of-the-art image forgery detectors, despite requiring substantially weaker supervision. We introduce effective intersection over union (eff-IoU), a unified metric to jointly evaluate detection and localization. These results suggest that reinforcement learning provides an effective and scalable mechanism for teaching VLMs to reason about AI-generated content.
21. 【2607.28463】VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
链接:https://arxiv.org/abs/2607.28463
作者:Haiyue Zhang,Yi Bin,Xun Jiang,Zeyu Ma,Duo Peng,Guoqing Wang,Yang Yang,Heng Tao Shen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large vision-language models, understanding long videos, long videos remains, videos remains challenging, achieved significant progress
备注:
点击查看摘要
Abstract:Large vision-language models (LVLMs) have achieved significant progress in video understanding, yet understanding long videos remains challenging due to the large number of visual tokens and limited context windows. Visual sampling provides a practical solution by selecting an informative subset of frames. However, existing methods typically either rely on relevance-aware sampling, leading to redundant frame selection and insufficient temporal coverage, or adopt a fixed sampling strategy regardless of query type. In this paper, we propose VisualRouter, a training-free and plug-and-play framework for query-grounded visual sampling. VisualRouter first classifies each query as either global or local and then applies the corresponding sampling strategy. For global queries, it employs a relevance-coverage hybrid strategy that preserves temporal coverage while retaining query-relevant visual evidence. For local queries, it adopts an event-aware frame selection strategy that performs event partitioning, segment-level frame allocation, and intra-event frame selection, jointly balancing relevance, coverage, and diversity with a limited number of input frames. Experiments show that VisualRouter consistently improves multiple LVLMs over uniform sampling, achieving gains of 5.2%, 7.7%, and 11.6% on Video-MME, LongVideoBench, and MLVU with Qwen2.5-VL-7B, and outperforming existing training-free visual sampling methods under the same setting.
22. 【2607.28442】ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
链接:https://arxiv.org/abs/2607.28442
作者:Ping-Kun Chiang,Kun-Ru Wu,Po-han Li,Sandeep Chinchali,Ufuk Topcu,Yu-Chee Tseng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent advances, large language models, advances in large, large language, enabled new possibilities
备注:
点击查看摘要
Abstract:Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbf{ViewMind3D}, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird's-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What'' questions in SQA3D, while maintaining strong overall accuracy (50.8\%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.
23. 【2607.28428】Kohn-Sham Spectral Embedding on Sparse Graphs at the Nishimori Temperature for Image Classification
链接:https://arxiv.org/abs/2607.28428
作者:V.S. Usatyuk,D. A. Sapozhnikov,S. I. Egorov
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Information Theory (cs.IT)
关键词:Sham Spectral Embedding, replacing dense CNN, dense CNN classifiers, sparse-graph spectral embedding, spectral embedding evaluated
备注: 42 pages, 10 figures, 5 tables, was presented at the 10th International Conference 'Deep Learning on Computational Physics (DLCP2026)', under review for the Moscow University Physics Bulletin, Physics series
点击查看摘要
Abstract:We introduce Kohn--Sham Spectral Embedding (KSSE), a physics-inspired energy-based model replacing dense CNN classifiers with a sparse-graph spectral embedding evaluated at the Nishimori temperature of an associated Random-Bond Ising Model. By mapping pre-trained features onto quasi-cyclic low-density parity-check graphs and constructing a regularized Laplacian acting as a Kohn--Sham Hamiltonian, we solve $D$ independent channel spectral problems in $\mathcal{O}(N\log N + k^2_{\text{mode}} N)$ time via FFT on circulant blocks (leveraging Pontryagin self-duality of $\mathbb{Z}/p\mathbb{Z}$) and low-order Rayleigh refinement. Graph topology is optimized using \emph{star-domain surgery}: rather than destroying information-carrying codewords by removing frustrated cycles, we construct edge shifts creating local convexity around codewords while bounding residual frustration to $\rho(B_\gamma)\leq 1+\delta$. Multi-scale fractal analysis ($D_2$ spectrum) and fractal learning-rate landscape certifies a landscape transition from rough regimes ($D_23$) to star-domain basins ($D_21$), enabling Rayleigh refinement with $k_{\text{mode}}=5$ modes. We prove six theoretical results: a generalized Ihara--Bass identity linking belief propagation to the Laplacian; trapping-set eigenvalue correspondence; additive channel separability with an explicit exchange-correlation bound; a surgery theorem bounding frustration with attractor width $\Omega(1/\sqrt{d_{\min}})$; a quasi-stationarity perturbation bound; and a fixed-point convergence theorem. In a transductive protocol on ImageNet-1000 with frozen EfficientNet-B4 features ($D=1792$), KSSE achieves \textbf{88.93\%} Top-1 accuracy using $\approx 21.24$M parameters, outperforming Swin-L (197M, 86.4--87.3\%) and matching ViT-H/14 (632M, 88.0--89.5\%) under standard inductive setups, while reducing model footprint by $10\times$ and $30\times$, respectively.
24. 【2607.28423】Negative controls reveal volume-driven confounding in radiomics and imaging foundation model features
链接:https://arxiv.org/abs/2607.28423
作者:Katy L. Scott,Sejin Kim,Joshua Siraj,Caryn Geady,Matthew Boccalon,Mattea Welch,Mogtaba Alim,Andrew J. Hope,Benjamin Haibe-Kains
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:foundation models promise, models promise non-invasive, promise non-invasive biomarkers, reflect tumour volume, promise non-invasive
备注: 22 pages (including supplementary), 6 figures, 2 supplementary tables, 5 supplementary figures
点击查看摘要
Abstract:Radiomics and imaging foundation models promise non-invasive biomarkers of tumour biology, yet predictive signatures may reflect tumour volume or acquisition artifacts rather than meaningful image structure. We introduce READII-2-ROQC, an open-source framework that uses volume-preserving negative controls to assess whether radiomic and deep imaging features capture independent spatial signals. READII-2-ROQC generates voxel-perturbed images across tumour, background and whole-image regions using configurable randomization strategies, then compares feature behaviour and model performance between original and control images. Applied to three public cancer imaging cohorts, the framework processed 3,552 tumour volumes and extracted PyRadiomics and foundation-model features from original images and nine matched controls. Reproducing published survival and HPV-status signatures, we show that multiple models retain performance after spatial structure is destroyed, revealing volume-driven or contextual confounding, whereas others show perturbation-sensitive signal. READII-2-ROQC provides a scalable quality-control strategy for developing interpretable, biologically grounded imaging biomarkers and reproducible radiomics workflows.
25. 【2607.28415】QQWorld: Quantile-Quantile Matching for World Model Regularization
链接:https://arxiv.org/abs/2607.28415
作者:Zhoushun Yu,Xiaoyu Hu,Xiangyu Xu
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Robotics (cs.RO)
关键词:compact representation space, world models enable, models enable efficient, predicting future states, performance depends critically
备注:
点击查看摘要
Abstract:Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.
26. 【2607.28401】Large scale cross-regional remote sensing flood monitoring framework for operative mapping and impact analysis
链接:https://arxiv.org/abs/2607.28401
作者:Ilya Novikov,Svetlana Illarionova,Ruslan Dzharkinov,Maria Smirnova,Ayrat Abdullin,Anna Korotkova,Mariia Ulianova,Dmitrii Shadrin,Evgeny Burnaev
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:populations and infrastructure, critical for minimizing, disasters on populations, Effective flood monitoring, flood monitoring
备注: 37 pages, 11 figures, 9 tables. Preprint submitted to Earth Systems and Environment. This version has not been peer reviewed
点击查看摘要
Abstract:Effective flood monitoring is critical for minimizing the impacts of flood disasters on populations and infrastructure. Yet reliable remote sensing across extensive and environmentally diverse regions remains challenging, as most segmentation algorithms lack the generalisation capacity required for large-scale application, while annotated flood data are scarce and unevenly distributed. This study presents an end-to-end multimodal framework for Russian Federation territories sustainable flood monitoring and damage assessment based on synthetic aperture radar data, multispectral imagery, and digital elevation models with their derivatives, forming a 21-channel input. Using a self-collected multimodal dataset covering seven Russian regions, two strategies for water surface detection under limited data conditions were compared: a supervised U-Net++ model and the self-supervised AnySat architecture pre-trained and fine-tuned for the segmentation task. Under the data conditions of this study, supervised learning proved more effective, while the AnySat-based approach offered greater stability and retains advantages for settings where larger unlabelled data or missing modalities at inference are expected. The best flood area predictions were used to estimate flood impact in urban areas in terms of the area affected, material damage, casualties, and ecological and agricultural impact. The estimations were conducted following the official methodology of the Russian Ministry of Emergency Situations. Applied to the 2019 Tulun flood, the obtained results closely matched official assessments, except for material damage, due to the open-source databases usage. The results demonstrate the potential of deep learning and multimodal satellite data integration for scalable, reliable flood monitoring across diverse environmental and data-limited conditions.
Comments:
37 pages, 11 figures, 9 tables. Preprint submitted to Earth Systems and Environment. This version has not been peer reviewed
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2607.28401 [cs.CV]
(or
arXiv:2607.28401v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2607.28401
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
27. 【2607.28394】Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
链接:https://arxiv.org/abs/2607.28394
作者:Weiquan Lin,Yu Deng,Shiyang Liu,Luping Xiao,Xu Tang,Junzhi Yu,Jiaolong Yang,Lei Zhang,Xingyu Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Hand-object interaction, modeling remains challenging, requires joint reasoning, HOI, object geometry
备注:
点击查看摘要
Abstract:Hand-object interaction (HOI) modeling remains challenging because it requires joint reasoning about hand articulation, object geometry, contact, semantics, and dynamics under severe visual uncertainty. Foundation models introduce transferable prior knowledge learned from large-scale cross-domain data, offering new ways to address these challenges beyond task-specific data and models. However, the rapidly growing literature remains fragmented, and existing studies typically describe these methods simply as ``using large models'' without systematically characterizing what knowledge is introduced, where it enters the HOI pipeline, or which HOI uncertainty it helps reduce. This survey presents the first systematic review of foundation-model priors for HOI. We organize the literature into six HOI tasks spanning reconstruction and generation. More importantly, we establish a taxonomy of eight foundation-model sub-priors grouped into geometric, semantic, and visual families. Geometric priors encompass shape retrieval, shape reconstruction, and spatial reconstruction; semantic priors include semantic grounding and language reasoning; and visual priors cover visual representation, image generation, and video generation. Based on this taxonomy, we systematically analyze how different priors are represented, injected, and adapted across HOI pipelines and tasks. Beyond how foundation models empower HOI, we further examine how HOI-derived knowledge is used in robot learning, including human-data pretraining, human-to-robot skill transfer, and HOI-to-robot data generation. Finally, we summarize datasets and evaluation protocols, and discuss limitations and future directions toward more generalizable HOI systems. To support long-term progress, we curate a live repository that continuously aggregates emerging methods and benchmarks.
28. 【2607.28386】Explaining Image Similarity with Automatically Extracted Concept Activation Vectors
链接:https://arxiv.org/abs/2607.28386
作者:Isaac Roberts,Petra Bevandic,Alexander Schulz,Barbara Hammer
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:computer vision applications, vision applications, underlies many computer, computer vision, receive a high
备注:
点击查看摘要
Abstract:Image similarity underlies many computer vision applications, yet it is often unclear why two images receive a high or low similarity score. Existing explainability methods often rely on gradient-based attribution maps to provide local justifications for similarity. These approaches struggle to provide global insights into what specifically drives similarity in regions of an embedding space, such as texture, shape, or color. We introduce a model- and metric-agnostic framework that explains image similarity using Concept Activation Vectors (CAVs) extracted automatically via Sparse Autoencoders (SAEs). Given a pair of images, we perturb their embeddings along discovered concept directions and measure the resulting change in a chosen similarity function, yielding concept importances. For image pairs, we provide localization with concept attribution maps. We extend this procedure to group-level settings, explaining what drives similarity across a cluster of images rather than a single pair, and further, we introduce Exemplar Retrieval, aiming to recover samples with similar reasons contributing to similarity. Our experiments show that our latent perturbations are more faithful to the underlying data distribution than pixel-space baselines, and that concept importances linearly recover the true similarity score. Qualitative results further confirm the usefulness of our methods in understanding a model's individual and group similarity judgments.
29. 【2607.28362】ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
链接:https://arxiv.org/abs/2607.28362
作者:Jin Cao,Zian Meng,Kaipeng Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
关键词:approach to any-action, frame-level control, dynamics, action, shadow
备注: [this https URL](https://ShadowDancer-1.github.io)
点击查看摘要
Abstract:We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at this https URL
30. 【2607.28341】Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models
链接:https://arxiv.org/abs/2607.28341
作者:Jie Ma,Zhike Qiu,Jie Gao,Jiayi Ji,Qian Chen,Xiaoshuai Sun,Rongrong Ji
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large Language Models, Multimodal Large Language, Language Models, Large Language, perform irreversible filtering
备注:
点击查看摘要
Abstract:While visual token pruning is essential for efficient Multimodal Large Language Models (MLLMs), existing training-free methods suffer from a critical limitation: they rely on static, instantaneous heuristics to perform irreversible filtering. This approach ignores the hierarchical nature of MLLMs, where token importance often evolves dynamically rather than remaining fixed across layers. Consequently, tokens essential for deep-layer reasoning are often prematurely discarded by shallow-layer estimates. To address this, we propose Trend-aware Pruning, a novel framework that elevates pruning from a local snapshot decision to a temporal trajectory modeling problem. Instead of relying on isolated scores, our method captures the momentum of attention flow. This enables a dynamic rectification mechanism that selectively reactivates "late-blooming" tokens, those initially undervalued but exhibiting rising semantic importance, thereby preventing the loss of critical visual cues. Extensive experiments demonstrate that our approach achieves a superior efficiency-performance trade-off across diverse multimodal tasks. Notably, it reduces visual tokens by over 77.8%, retaining only approximately 23 tokens in the final layer while maintaining competitive performance, offering a robust and reversible solution for high-efficiency multimodal inference.
31. 【2607.28327】Same Branches, Different Trees: A Bifurcation Connectedness Metric for Coronary Artery Segmentation and FFR-CT Decision Agreement
链接:https://arxiv.org/abs/2607.28327
作者:Maame Owusu-Ansah,Kelvin Lee,Dr Vinod Venugopal,Muhammad Moazzam Jawaid,Wenting Duan,James Brown
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Fractional flow reserve, reach high Dice, flow reserve derived, patient-specific vessel model, Fractional flow
备注: Accepted at STACOM 2026 (MICCAI workshop). 11 pages, 3 figures, 2 tables
点击查看摘要
Abstract:Fractional flow reserve derived from CT angiography (FFR-CT) simulates flow through a patient-specific vessel model, so its accuracy depends on the connectedness of the segmented tree, not only on volumetric overlap: a segmentation can reach high Dice yet sever a bifurcation, dropping the downstream subtree and reversing the treatment decision. Topology-aware losses such as clDice and Skeleton Recall act on the global centreline and can miss localised breaks. We study the Bifurcation Connectedness Score (BCS), which scores connectedness at each ground-truth bifurcation, and soft-BCS, its differentiable training surrogate. BCS captures a property of segmentation quality the standard metrics miss: it responds strongly to breaks in connectedness while staying largely unchanged under connectedness-preserving narrowing. Higher BCS accompanies closer agreement between the FFR-CT decisions a solver makes on predicted versus ground-truth geometry, most clearly in severe disease (OR 2.16, CI [1.23, 4.18]). Both decisions come from the same solver, so this reflects geometric, not clinical, fidelity. In training, soft-BCS and Skeleton Recall recover the same branches but build different trees. Recovering branches and keeping them connected are separable properties, so we recommend reporting a measure of each.
32. 【2607.28320】AdaAnchor4D: Anchor-Conditioned Spatiotemporal Feature Aggregation for Monocular UAV 4D Reconstruction
链接:https://arxiv.org/abs/2607.28320
作者:Peiyi Xu,Junpeng Zhang,Guanbin Li,Ronghua Shang,Mingtao Feng,Le Dong,Weisheng Dong,Guangming Shi,Jie Feng
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:videos provide valuable, provide valuable observations, UAV videos provide, Monocular UAV videos, complex urban scenes
备注: 9 pages, 4 figures
点击查看摘要
Abstract:Monocular UAV videos provide valuable observations for dynamic reconstruction of complex urban scenes. However, such scenes exhibit pronounced spatiotemporal heterogeneity: different regions follow distinct temporal activity patterns, while the motion states of some dynamic regions may further evolve over time. Although dynamic Gaussian methods based on decomposed shared spatiotemporal feature fields have achieved efficient and accurate reconstruction in object-centric or relatively compact scenes, their commonly adopted fixed plane-wise feature combination mechanisms are less suited to the heterogeneous local dynamics of UAV scenes, often leading to ghosting artifacts and blurred dynamic details. To address this challenge, we propose AdaAnchor4D, an adaptive anchor deformation framework for monocular UAV dynamic scene reconstruction. At its core, Anchor-Conditioned Feature Aggregation (ACFA) adaptively aggregates shared spatiotemporal features using anchor-specific aggregation embeddings and temporal information, allowing different local units to obtain dynamic representations tailored to their local and temporal states. Decoupled Local Geometry Deformation (DLGD) separates anchor-state deformation from local Gaussian geometry deformation, while Density-Adaptive Coordinate Warping (DACW) reparameterizes feature-query coordinates according to the axis-wise anchor distributions, alleviating the mismatch between non-uniform geometric sampling and uniform grid parameterization. Experiments on UAV-Arc4D, VisDrone, and UAVDT show that AdaAnchor4D achieves higher rendering quality than representative dynamic Gaussian methods while maintaining real-time rendering performance. The code will be made publicly available.
33. 【2607.28312】ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding
链接:https://arxiv.org/abs/2607.28312
作者:Mingkang Dong,Muxin Pu,Jie Li,Bohan Guo,Songruo Chen,Bin Ren,Xu Zheng,Chen Zhao,Tianwen Qian,Mohamed Elhoseiny,Yuqian Fu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:video understanding requires, Streaming video understanding, continuously retain, future questions, understanding requires models
备注: 9 pages
点击查看摘要
Abstract:Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal redundancy, or segment-level relevance, but rarely organize evidence around objects that persist and evolve over time. Thus, in this paper, we introduce ObjectStream, a training-free framework that treats latent objects as memory anchors for streaming video understanding. ObjectStream induces spatially coherent latent objects directly from frozen Video-LLM representations, links them across frames into persistent anchors, and maintains their histories under a bounded memory budget, without requiring external object detectors or segmentation models. Built on these anchors, ObjectStream preserves three complementary forms of evidence: persistent object histories, transient object changes, and recent visual context. This design enables existing Video Large Language Models (Video-LLMs) to reason over object identities, interactions, and state changes while leaving the underlying model unchanged. Extensive experiments on online streaming and offline long-video benchmarks demonstrate both effectiveness and efficiency. In online streaming evaluation, ObjectStream improves Qwen2.5-VL-7B by 10.0 points on OVO-Bench Real-Time Visual Perception, while reducing peak GPU mem-ory and TTFT by approximately 50%. On offline long-video benchmarks, it surpasses the full-token baseline while discarding 82.5% of visual tokens. These results highlight latent objects as a practical and effective organizing principle for compact streaming video memory.
34. 【2607.28300】MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians
链接:https://arxiv.org/abs/2607.28300
作者:Pouya Ardekhani,Zahra Dehghanian,Morteza Abolghasemi,Hamid R. Rabiee
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:next-generation interactive systems, navigate reconstructed environments, Open vocabulary, scene understanding, interactive systems
备注:
点击查看摘要
Abstract:Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features. We present a novel, training-free pipeline that fundamentally reimagines this paradigm by explicitly decoupling 3D geometric reconstruction from semantic integration. Given a standard monocular video sequence as input, our method efficiently outputs a compact, highly interpretable, and fully searchable object-level semantic Gaussian map. Rather than entangling heavy language embeddings within the mapping loop, we extract geometry independently and ground semantics through a lightweight, modular post-processing framework. Extensive evaluations on the Replica dataset demonstrate that this decoupled architecture preserves strong rendering fidelity and competitive segmentation accuracy. Crucially, by replacing dense per-Gaussian storage with modular, object-level semantic embeddings, our approach delivers an order-of-magnitude reduction in memory usage compared to SOTA baselines. This provides a highly efficient, scalable, and practical solution for open-vocabulary 3D retrieval and question answering directly from everyday monocular video.
35. 【2607.28293】Filling the Pareto-Optimal Front for Affordance Segmentation on Embedded Devices Using RGB-D Cameras
链接:https://arxiv.org/abs/2607.28293
作者:Edoardo Ragusa,Giovanni Paolo Canuti,Simone Lugani,Rodolfo Zunino,Paolo Gastaldo
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Robotics (cs.RO)
关键词:complement RGB data, RGB data, complement RGB, wearable robots, remain underexplored
备注:
点击查看摘要
Abstract:While depth sensors have the potential to complement RGB data for affordance segmentation in wearable robots, their usage seems to remain underexplored. The paper proposes two approaches: a reformulated version of hardware-aware neural architecture search, endowed with a newly designed search space to integrate depth (D) information into small-sized deep networks, and a dedicated fine-tuning approach, including a preprocessing layer to merge depth information with RGB data and make it compatible with conventional architectures. In both cases, those methods aim to generate solutions that benefit from modern (portable) hardware accelerators and overcome existing tiny-like approaches, which often fail to tackle critical scenarios due to the severe constraints set by the supporting hardware. Extensive experiments on a pair of real-world datasets demonstrate the effectiveness of the proposed method as compared with existing solutions. The approach presented in the paper generates, in most cases, solutions that identify the Pareto optimal front to balance generalization performance and hardware requirements. The paper also describes the supporting prototype, including a Jetson Nano board and a RealSense RGB-D camera. When considering the energy profile of the device, the overall system can attain real-time performances within an energy budget that is compatible with standard batteries, such as those used in smartphones.
36. 【2607.28287】ycho: Active Abstraction with Programmatic World Models for ARC-AGI-3
链接:https://arxiv.org/abs/2607.28287
作者:Jens Lehmann,Andrei Aioanei,Sahar Vahdati
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Symbolic Computation (cs.SC)
关键词:skill acquisition, RHAE, action efficiency, turns abstraction, Opus
备注: 52 pages, 18 figures, 17 tables. Open-source implementation: [this https URL](https://github.com/NIMI-research/Tycho)
点击查看摘要
Abstract:ARC-AGI-3 turns abstraction into an interactive problem of skill acquisition. A player must infer an unfamiliar game's rules, hidden state, and goal while maintaining action efficiency because every move counts. We formalize these environments as parameterized rendered deterministic Moore machines and introduce Tycho, a coding-agent system that constructs and uses game-specific models during interaction. Tycho separates actionable observations from intermediate animation, level-completion, and game-over frames. From this structured history, an agent can model, test, plan with, repair, or bypass a free-form executable hypothesis. In one matched public-set run per policy, we compare four orchestration policies on all 25 public games using Claude Opus 4.8 under matched inference budgets. Actor-requested delegation to a model builder obtains the highest observed mean Relative Human Action Efficiency (RHAE), 88.49. With this selected policy, GPT-5.6 Sol and Opus 5 both reach 100.00 RHAE and complete all 183 levels. Their game-balanced first-run human-replay midranks are 98.5 and 100.0. Opus 5 uses 61% fewer scored actions than the aggregate official human baselines. Automatic repair after verification failures produces models that reproduce observed transitions much more accurately, yet reaches only 83.07 RHAE. Transition match indicates whether a simulator reproduces observed dynamics, not whether it has identified the objective or improves the next action. Strong play also requires deciding when to construct, repair, use, or bypass a model. We call this joint problem active abstraction: generating a testable model from costly interaction and deciding when acquiring or using it is worth its cost.
Comments:
52 pages, 18 figures, 17 tables. Open-source implementation: this https URL
Subjects:
Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Symbolic Computation (cs.SC)
ACMclasses:
I.2.6; I.2.8
Cite as:
arXiv:2607.28287 [cs.AI]
(or
arXiv:2607.28287v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2607.28287
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
37. 【2607.28285】Beyond Visual Ambiguity: Guiding Robust Monocular Depth Estimation in Challenging Scenarios via Detailed Long Captions
链接:https://arxiv.org/abs/2607.28285
作者:Junrui Zhang,Jiaqi Li,Yiran Wang,Liao Shen,Zhiguo Cao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Monocular depth estimation, single-image limited information, faces challenges, Monocular depth, inherent in single-image
备注: Accepted to ACM MM 2026
点击查看摘要
Abstract:Monocular depth estimation (MDE) faces challenges with non-Lambertian surfaces and adverse weather conditions due to the visual ambiguities inherent in single-image limited information. Existing works address them in isolation via image inpainting or augmentation, yielding limited robustness gains. Language, as a powerful complementary modality to vision, is demonstrated to enhance the visual perception capabilities of vision-language models (VLMs) via detailed long captions. However, prior language-integrated MDE methods fail to fully harness this potential due to short text input with limited information, coarse global text feature learning, and limited language guidance during depth decoding. To address these limitations, we propose CapDepth, a novel framework for robust MDE that leverages guidance from detailed long captions to alleviate visual ambiguities in both challenging scenarios. First, we design a detailed long caption input template that explicitly conveys rich spatial relationships among multiple atom sentences. Second, a dynamic caption encoder is introduced to extract fine-grained depth-relevant text features via progressive masked attention. Finally, we propose a text-adaptive decoder that guides enhanced depth decoding with text features via stable adaptive layer normalization. Extensive experiments validate the efficacy of CapDepth, which outperforms state-of-the-art methods, achieving depth error reductions of 25.0% on non-Lambertian surfaces and 22.0% under adverse weather conditions.
38. 【2607.28277】MSCM-net: A hyperspectral image classiffcation method based on multi-scale convolution and Mamba
链接:https://arxiv.org/abs/2607.28277
作者:Jianjun Chen,Linlin Wang,Lifang Chang,Limin Huo,Shujiang Song,Yanjia Zhao,Mingwei Shao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:sensing and engineering, remote sensing, multi-scale CNN, Hyperspectral imaging, CNN and Transformer-based
备注:
点击查看摘要
Abstract:Hyperspectral imaging is widely used in remote sensing and engineering. Therefore, research on its classification methods is crucial. While CNN and Transformer-based methods have advanced, they still face locality constraints and high computational complexity. To address these issues, we propose an innovative hyperspectral image classification model, MSCM-net. Specifically, first of all, a model architecture combining multi-scale CNN and Mamba is proposed. It consists of a multi-scale feature extraction module (MCSE) and multiple stacked Mamba blocks, which integrates the local feature extraction capability of multi-scale CNN and the long sequence modeling advantage of Mamba. Secondly, the proposed MCSE module consists of multi-scale convolution and SENet. Convolution kernels of different scales extract local information with different receptive fields, enhancing the fusion of spatial and spectral information. Meanwhile, the SENet enables the model to automatically learn the importance of each channel in the multi-scale features. Furthermore, we also propose a dual-branch feature aggregation module, which further effectively extracts and integrates the spectral information contained in the central pixel and the spatial information in the surrounding pixels. Our model has undergone numerous experiments on three widely used benchmark datasets. The experimental results show that MSCM-net can achieve advanced classification performance while reducing computational complexity.
39. 【2607.28269】heia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation
链接:https://arxiv.org/abs/2607.28269
作者:Simone Giano,Lorenzo Severini,Alessandro Galdelli,Adriano Mancini
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Multimedia (cs.MM)
关键词:management requires high-quality, requires high-quality multimodal, disaster management requires, high-quality multimodal datasets, deployment of Vision-Language
备注:
点击查看摘要
Abstract:The deployment of Vision-Language Models (VLMs) in critical domains like disaster management requires high-quality multimodal datasets, especially for transferring knowledge via Data-Free Knowledge Distillation (DFKD). However, existing datasets in this domain either entirely lack descriptive text, such as Incidents1M, or suffer from severe text-image semantic misalignment, such as CrisisMMD. In this work, we present a novel methodology to construct and automatically validate a large-scale multimodal dataset for disaster response. Starting from the vision-only Incidents1M, we successfully recovered 100,000 images and generated high-fidelity textual descriptions using two distinct Qwen3.5 architectures: a 4B dense model and a 35B Mixture-of-Experts (MoE) model. To ensure the generated captions provide reliable semantic anchoring for DFKD, we introduce an image-blind LLM-as-a-Judge validation pipeline leveraging Qwen3.5-9B. By intentionally obscuring the original image from the judge, this evaluator accurately simulates the modality gap of the student model during data-free distillation. Our evaluation across 173,179 label pairs demonstrates a high semantic agreement (78.65/100) between the two architectures. Furthermore, the automated evaluation reveals a conservative captioning behaviour, characterized by a high Precision (77.6%) and low Recall (46.0%). This minimizes the false positive noise, while simultaneously exposing underlying human annotation inconsistencies in the original ground truth. This work provides a scalable, LLM-validated multimodal dataset and a reproducible framework to advance cross-modal knowledge distillation.
40. 【2607.28261】ARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
链接:https://arxiv.org/abs/2607.28261
作者:Jiwen Liu,Shujuan Li,Xiaohan Li,Zijie Meng,Xinyue Liu,Yulong Xu,Yan Zhou,Guoxin Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Video re-shooting aims, aims to regenerate, Video re-shooting, regenerate videos, camera
备注: 8 pages, 5 figures
点击查看摘要
Abstract:Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: this https URL
41. 【2607.28247】Space2Ground 2.0: A Multi-Source Dataset and Framework for Agricultural Monitoring through Fusion of Street-Level and Satellite Imagery
链接:https://arxiv.org/abs/2607.28247
作者:Iason Tsardanidis,Alkiviadis Koukos,George Choumos,Vasileios Sitokontantinou,Charalampos Kontoes
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:cloud-induced temporal gaps, Accurate and scalable, satellite Earth Observation, monitoring remains challenging, Earth Observation
备注: This paper has been accepted for presentation at the 45th EARSeL Symposium, Athens, Greece
点击查看摘要
Abstract:Accurate and scalable parcel-level agricultural monitoring remains challenging because satellite Earth Observation alone provides only an overhead perspective of agricultural parcels, while optical observations are further affected by cloud-induced temporal gaps. This paper presents Space2Ground 2.0, a multi-source framework integrating Sentinel-1 SAR and Sentinel-2 multispectral time series with geo-tagged street-level imagery acquired using vehicle-mounted cameras and shared through the Mapillary platform. A largely automated processing pipeline performs semantic filtering, image quality assessment, viewpoint-based parcel association, and dataset refinement, transforming large volumes of crowdsourced imagery into parcel-linked, analysis-ready data. Applied over Cyprus during the 2022 growing season, the pipeline produced a curated dataset of 46,050 annotated street-level images, selected from an initial collection exceeding 900,000 images and linked with satellite information for 8,581 agricultural parcels. The practical value of the dataset was assessed through parcel-level crop classification experiments using both single- and multi-source observations. The results demonstrate that street-level imagery provides complementary fine-scale visual information that enhances classification when integrated with satellite time series. Overall, Space2Ground 2.0 provides an openly available benchmark dataset and a reproducible methodology for multimodal agricultural monitoring, with potential applications in visual verification, reduced reliance on costly field inspections, and data-driven agricultural policy implementation.
42. 【2607.28243】EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
链接:https://arxiv.org/abs/2607.28243
作者:Zexuan Yan,Yuzhou Wu,Yue Ma,Zonghang He,Kaibo Yin,Xiaobing Tu,Yinggui Wang,Jinkui Ren,Xiantao Zhang,Shijian Wang,Jinghong Liu,Linfeng Zhang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:embodiments remains costly, offers rich manipulation, rich manipulation experience, collecting diverse egocentric, video offers rich
备注: project page: [this https URL](https://egogenesis.github.io/)
点击查看摘要
Abstract:Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.
43. 【2607.28227】Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
链接:https://arxiv.org/abs/2607.28227
作者:Hanzhang Zhou,Panrong Tong,Xu Zhang,Quyu Kong,Chenglin Cai,Tianyu Xia,Gongjie Zhang,Jianan Zhang,Long Li,Long Chen,Lei Wang,Gaole Dai,Pengxiang Li,Liangyu Chen,Yue Wang,Steven Hoi
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:general purpose executor, existing digital devices, general purpose, purpose executor, executor over existing
备注:
点击查看摘要
Abstract:GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tasks, proactively initiate useful services, and autonomously improve their capabilities with minimal human effort. Guided by this vision, we present Qwen-UI-Agent, a real-world centric foundation GUI agent spanning mobile, computer-use, web, and DeepSearch environments. Qwen-UI-Agent combines diverse sandbox environments with a large-scale real-device mobile runtime. Its unified action space interleaves GUI operations with CLI execution and generates batched actions in a single model turn. An AutoResearch-style data flywheel uses agents to construct tasks and environments, diagnose failures, and plan subsequent iterations. Online RL supports training on trajectories exceeding 100 turns, with over 10,000 concurrent environments accelerating rollout. A lightweight harness layer supports proactive service initiation and stateful workflows across mobile and computer. Across a broad suite of evaluations, Qwen-UI-Agent sets state-of-the-art performance on mobile-use benchmarks while delivering competitive performance on computer- and browser-use tasks against frontier models, including Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol. On mobile use, it achieves 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily. On computer use, it achieves 79.5% on OSWorld-Verified and a 40.0% partial-progress score on OSWorld-v2. On browser use and GUI grounding, it achieves 73.6% on WebArena and 81.5% on ScreenSpot-Pro, respectively.
Subjects:
Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2607.28227 [cs.AI]
(or
arXiv:2607.28227v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2607.28227
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
44. 【2607.28225】FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
链接:https://arxiv.org/abs/2607.28225
作者:Haoqing Wang,Xingrun Xing,Wei Xia,Ziheng Li,Yehui Tang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Agentic vision-language models, Agentic vision-language, code-based image manipulation, interleave textual reasoning, interpretable multimodal reasoning
备注:
点击查看摘要
Abstract:Agentic vision-language models (VLMs), which interleave textual reasoning with explicit tool calls such as cropping and code-based image manipulation, have emerged as a compelling paradigm for reliable and interpretable multimodal reasoning. However, recent studies have revealed that such models often use tools unfaithfully. Many process images are irrelevant to the question (e.g., the tool crops the wrong region or misses the queried target), yet the call still receives full credit and the model still answers correctly. Such decorative or misaligned tool calls waste computation and reveal that the model leans on prior knowledge or the original image rather than the evidence it retrieves. This may stem from two limitations of prevailing methods: the tool reward fails to distinguish useful from useless calls, and tool feedback carries no signal of usefulness. To this end, we introduce FaithEyes, a multi-agent self-judging framework. Concretely, we use a VLM to judge whether each process image helps answer the question. The judgement is injected into the reasoning context as part of the tool observation to help subsequent reasoning, and meanwhile is used to scale the tool reward by the helpful-tool ratio to suppress reward hacking. To keep judgement available at evaluation and thus ensure train-test consistency, we further design a multi-agent framework where the model itself serves as a subagent to judge the tool calls from main agent, eliminating any dependence on an external model at inference. Training via a two-stage SFT + RL pipeline on adapted open-source data, FaithEyes attains competitive or superior accuracy across visual perception and reasoning benchmarks, while markedly improving tool faithfulness. The homepage is at this https URL.
45. 【2607.28211】Scaling Vision-Language Models Is Not Enough to Mitigate Bias
链接:https://arxiv.org/abs/2607.28211
作者:Ioannis Sarridis,Ioannis Kompatsiaris,Symeon Papadopoulos
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remains poorly understood, spurious correlations remains, correlations remains poorly, multimodal systems, foundational to multimodal
备注:
点击查看摘要
Abstract:Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases). Across these settings, the Spearman correlation between model scale and performance weakens as evaluation shifts from ImageNet ($\rho{=}0.68$) to single-attribute ($\rho{=}0.48$) and further to multi-attribute ($\rho{=}0.05$) bias benchmarks. In contrast, properties of the training data (size and quality) show more consistent relationships with worst-group accuracy across both bias benchmarks. Notably, curated datasets yield improvements of up to 25% over uncurated alternatives at a comparable scale. Finally, the effect of architectural choices (e.g., patch size, image resolution) is highly context-dependent, varying with the nature of the benchmark, including the type of bias and its spatial distribution within images.
46. 【2607.28198】UniCross: Unified Cross-Skill Dexterous Manipulation Synthesis
链接:https://arxiv.org/abs/2607.28198
作者:Hui Zhang,Julian Ferchow,Jie Song,Mirko Meboldt
类目:Robotics (cs.RO); Computer Vision and Pattern Recognition (cs.CV)
关键词:remain securely held, remain securely, securely held, manipulation tasks require, skills
备注: Project page: [this https URL](https://zdchan.github.io/UniCross/)
点击查看摘要
Abstract:Many dexterous manipulation tasks require the object to remain securely held throughout the interaction. From the perspective of hand-object relational motion, such manipulation comprises four canonical skills: grasping, relocation, in-hand rotation, and in-hand translation. Human hands flexibly compose these skills to accomplish complex tasks. Existing approaches, however, model these skills separately with skill-specific action constraints, objectives, or even dedicated hand morphologies, which breaks the compatibility and continuity required for long-horizon composition. In this work, we present a unified framework that models all four skills in a single formulation that shares the same state and action spaces and a common objective structure. This formulation enables straightforward distillation of a single cross-skill policy that performs strongly on every skill, generalizes to unseen objects, stays robust to disturbances, and chains skills seamlessly into long-horizon manipulation. The framework also transfers effectively across different hand morphologies. Overall, our results suggest that different dexterous manipulation skills can be viewed as instantiations of a shared task formulation, revealing the intrinsic consistency across different behaviors.
47. 【2607.28186】hink with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain
链接:https://arxiv.org/abs/2607.28186
作者:Haiyang Wu,Weiliang Mu,Zhuofei Du,Dandan Zhong,Kaijie Shi,Haifeng Li,Chao Tao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:sufficient visual evidence, farmland remote sensing, remote sensing image, Existing farmland remote, remote sensing
备注:
点击查看摘要
Abstract:Existing farmland remote sensing image (FRSI) segmentation follows a "Think with Intra-Image" paradigm, assuming that the current image contains sufficient visual evidence for reliable segmentation. Yet farmland appearance varies with phenology and spatial context and is often confused with other land-cover, making instantaneous, local observations inadequate. Thus, segmentation ambiguity stems not only from limited model representation, but more fundamentally from the required spatio-temporal information lying beyond the current image. Based on this insight, we redefine FRSI segmentation from an information bottleneck perspective as a dynamic decision process driven by task-relevant extra spatio-temporal information gain. We further propose FarmSeeker, a dynamic FRSI segmentation agent that identifies ambiguous regions, reasons about their causes, and queries extra spatio-temporal information on demand for accurate segmentation. To evaluate FarmSeeker, we construct GSFS-Bench, the first global-scale, high-resolution FRSI segmentation benchmark that supports reasoning-querying. Experiments show that FarmSeeker achieves more stable segmentation performance than existing methods. The project is publicly available at: this https URL
48. 【2607.28164】S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single Image
链接:https://arxiv.org/abs/2607.28164
作者:Hail Song,Seokhwan Yang,Jiwon Yang,Woojin Cho,Woontack Woo
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:Gaussian Splatting, lifelike Virtual Reality, generating photorealistic, strategies for animating, Virtual Reality
备注: 15 pages, 12 figures
点击查看摘要
Abstract:We propose S-Avatar, a novel method for generating photorealistic 3D head avatars from a single image using a diffusion-guided 3D model generation module and strategies for animating 3D Gaussian Splatting (3DGS). While single-image head avatar reconstruction is crucial for lifelike Virtual Reality (VR) applications, existing approaches often struggle to preserve 3D consistency under unseen viewpoints. S-Avatar addresses this limitation through a three-stage pipeline. First, a high-resolution 3DGS is synthesized directly from a single image using a diffusion-based Gaussian splat generation module. Next, the parametric head model FLAME is aligned with the generated 3DGS by optimizing its parameters and spatial transformations. Finally, to adapt the 3DGS to FLAME variations, we construct a binding template that encodes the spatial relationship between the initial splats and FLAME. The dynamic 3D head avatar can then be rendered in real time by deforming the 3DGS with the binding template. By combining diffusion-guided canonical 3DGS generation with FLAME-based control, our method achieves efficient and accurate reconstruction with enhanced 3D consistency. Evaluations on public datasets demonstrate that S-Avatar outperforms state-of-the-art methods in novel-view and expression generation, achieving superior realism and consistency. Consequently, our approach represents a significant advance in accessible avatar creation, applicable to a wide range of VR/AR applications. The project page is available at this https URL.
49. 【2607.28154】OPLD: On-Policy Latent Distillation for Multimodal Reasoning
链接:https://arxiv.org/abs/2607.28154
作者:Shoutai Zhu,Tianyang Xu,Bin Sun,Mingyuan Xu,Yu Liu,Qinzhen Guo
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Interleaved multimodal, incorporating auxiliary visual, reasoning, improves visual reasoning, visual
备注:
点击查看摘要
Abstract:Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by internalizing intermediate computation into continuous representations. Nevertheless, existing visual-latent methods mainly supervise latent states through alignment with compressed auxiliary visual features, treating them as proxies for visual observations rather than active reasoning states. Consequently, they capture the provided evidence but fail to fully internalize the abstract reasoning process induced by multimodal CoT. In this paper, we propose OPLD (On-Policy Latent Distillation), a simple framework that transfers the reasoning capability induced by privileged multimodal CoT into latent reasoning representations. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods and achieves state-of-the-art performance on multiple benchmarks. The results suggest that supervising latent representations at the reasoning-process level provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.
50. 【2607.28148】What Makes Deep Learning Work for Traditional Chinese Medicine Tongue Diagnosis? A Comprehensive Ablation Study
链接:https://arxiv.org/abs/2607.28148
作者:Longxia Gao,Linan Wang,Yuhe Han,Junze Geng,Meng Zhang,Hanqing Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:traditional Chinese medicine, space remains underexplored, Chinese medicine, automated tongue diagnosis, Deep learning
备注: 30 pages, 8 figures, 9 tables
点击查看摘要
Abstract:Deep learning has shown promise for automated tongue diagnosis in traditional Chinese medicine (TCM), yet the design space remains underexplored. We conducted a systematic ablation study spanning 20+ model versions under rigorous 5-fold cross-validation on TongueDx2 (5,109 images, 976 expert-annotated) and a merged dataset of 11,101 samples. We compared six backbone architectures, four loss functions, five augmentation strategies, and six training strategies. The best 976-sample model achieved weighted-F1 of 0.6625 using ConvNeXt-Tiny with restrained augmentation and weak-group ensemble, while the best 11,101-sample model reached weighted-F1 of 0.7761. Six key design principles emerged: (1) ConvNeXt-Tiny offers optimal parameter efficiency; (2) BCE substantially outperforms Asymmetric Loss (+2.7%); (3) restrained color augmentation is critical; (4) weak-group ensemble replacement (+2.1%) outperforms probability averaging; (5) data scaling yielded +20.6% improvement; (6) expanding from 13 to 45 label dimensions caused catastrophic collapse (0.78 to 0.22). These principles are generalizable to multi-label medical image classification with class imbalance.
51. 【2607.28132】Convolutional Neural Shading for High-Quality 3D Reconstruction from Multi-View Images
链接:https://arxiv.org/abs/2607.28132
作者:Juheon Hwang,Taewan Kim,Heeseok Oh,Jiwoo Kang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:convolutional neural shading, reconstruct high-quality, pipeline to reconstruct, CNS, neural shading
备注:
点击查看摘要
Abstract:We propose a convolutional neural shading (CNS), a novel pipeline to reconstruct high-quality 3D shapes from multi-view images. Several recent studies have used neural radiance fields and other neural differentiable rendering methods to understand 3D geometry. However, these approaches rely on single-point geometric information, such as positions and normals of the surface, leading to a lack of detailed local geometry. Our approach addresses the inherent limitations of single-point information by leveraging a neural shader to capture variations even in dark and textureless regions with a convolutional neural shader, resulting in far more accurate geometry predictions. Additionally, our method mitigates surface irregularities at image boundaries by introducing a fine-detail displacement network, which utilizes spatial information of surface geometry and learns fine displacement details by correlating neighboring values in the rendering coordinates. Through extensive experiments, our proposed method has demonstrated significant quality improvements in the reconstructed shapes and rendered images over current state-of-the-art methods.
52. 【2607.28130】Collaborative Feature Aggregation for Face Super-Resolution and Robust Re-Identification
链接:https://arxiv.org/abs/2607.28130
作者:Juheon Hwang,Taewan Kim,Jiwoo Kang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:facial, sequential or multi-view, collaborative approach, robust person re-identification, robust person
备注:
点击查看摘要
Abstract:We propose a novel collaborative approach for face super-resolution (SR) and robust person re-identification from sequential or multi-view facial images. Traditional SR methods often suffer from blurring and distortion in faces recovered from poor-quality images due to low resolution. Image- and video-based facial SR methods using facial landmarks or segmentation also have similar challenges. To overcome these limitations, we leverage multiple correlated facial observations, across time or viewpoints, by introducing a transformer-based collaborative feature aggregation method that unifies identity features from multi-sequence or multi-view data. This allows faces in multiple sequences of an individual to contribute to accurately estimating common facial features. Furthermore, we propose a cascade SR network to progressively restore the high-resolution image of the target's face with gradual facial feature unification. The unified identity representation is further utilized in person re-identification scenarios, enabling accurate matching even under severe image degradation. The exhaustive experimental results and comparisons show that our method outperforms other state-of-the-art methods, demonstrating consistent improvements in both face super-resolution and re-identification performance. Our work highlights the effectiveness of joint identity reconstruction and progressive image restoration from multiple facial inputs in enhancing downstream visual recognition tasks.
53. 【2607.28129】Face and Voice Cross-modal Association with Learning Convex Feature Embedding
链接:https://arxiv.org/abs/2607.28129
作者:Taewan Kim,Jiwoo Kang
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:deep learning, association learning, cross-modal, cross-modal association tasks, learning
备注:
点击查看摘要
Abstract:Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull. Moreover, the incorporation of cross-modal attention mechanisms with convex embedding techniques represents a highly effective strategy for the attenuation of false positives and false negatives, accomplished via the minimization of inter-class discrepancies. We exhaustively evaluated our method for cross-modal verification, matching, and retrieval tasks on the large-scale VoxCeleb dataset. Extensive experimental results demonstrate that the proposed method achieves notable improvements over existing state-of-the-art methods.
54. 【2607.28125】owards Practical Algorithm Selection for Unsupervised Domain Adaptation in Medical Imaging
链接:https://arxiv.org/abs/2607.28125
作者:Yiheng Xiong,Luisa Gallée,Daniel Santak Wolf,Heiko Hillenhagen,Michael Götz
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Numerous unsupervised domain, unsupervised domain adaptation, domain prevents direct, prevents direct evaluation, Numerous unsupervised
备注:
点击查看摘要
Abstract:Numerous unsupervised domain adaptation (UDA) algori-thms exist, but for clinical practice, selecting the best-suited one along with proper hyperparameters often remains unclear, as the unlabeled deployment (target) domain prevents direct evaluation. We propose a label-free criterion that jointly selects the algorithm and hyperparameters for UDA. Given a pool of candidate models from multiple algorithms trained with different hyperparameters, our approach scores each candidate against an agreement reference, and selects the one with the highest score. The agreement reference is constructed in two levels without using target labels. First, we leverage multiple label-free selection signals, using each to nominate a model within every algorithm. Second, the nominated models are aggregated across algorithms to form a reference prediction for each unlabeled target sample. The candidate whose predictions agree most with this reference is then selected for deployment. Experimental results on four brain MRI and four chest X-ray datasets across seven clinically relevant transfer scenarios show that our method achieves better selection performance than other methods and remains effective across different algorithm pools. Our approach takes a step towards practical, label-free algorithm selection for clinical deployment of UDA.
55. 【2607.28108】mmRadarTwin: A Measurement-Calibrated Signal-Level Digital Twin Platform for Indoor mmWave Radar
链接:https://arxiv.org/abs/2607.28108
作者:Jianyi Zhou,Chenghao Zhang,Yanli Li,Dong Yuan
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Indoor mmWave radar, mmWave radar perception, scene geometry, perception is difficult, difficult to reproduce
备注: 7 figures, 4 tables
点击查看摘要
Abstract:Indoor mmWave radar perception is difficult to reproduce because measured range-angle responses depend on scene geometry, material response, multipath, hardware conventions, and signal processing. Existing ray-tracing and digital-twin tools often expose rendering, channel, or path-level quantities, while radar sensing requires complex signal products that can be processed and compared in the same domain as real FMCW measurements. We present mmRadarTwin, a signal-level and path-attributed digital-twin platform for indoor mmWave radar. mmRadarTwin links a real radar measurement branch with an Unreal Engine scene-simulation branch through a shared receive-channel and range-angle processing interface. The simulator writes complex multi-channel receive grids and exports per-path contribution records that identify the actor, material tag, propagation event, and output-bin support of each simulated return. We evaluate mmRadarTwin in an office deployment using a commodity monostatic mmWave radar and mobile scene-capture hardware. Across 154 measured poses spanning 22 radar locations, the current physics-only path-basis simulator recalls 70.8% of measurement-active geometry-supported response regions in the central usable field of view while exposing residuals caused by weak or missing path support, shifted responses, unsupported anchors, and missing physical mechanisms. Rather than claiming complete radar-map reconstruction or cross-room generalization, mmRadarTwin establishes a practical systems workflow for constructing, comparing, and diagnosing indoor radar digital twins.
56. 【2607.28073】GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios
链接:https://arxiv.org/abs/2607.28073
作者:Yiming Xu,Jihua Kang,Chunsai Du,Qifan Zhang,Wangqiu Zhou,Yiting Wu,Tianqi Li,Qi Song
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:high cognitive load, demanding professional environments, meeting review scenarios, cognitive load, imposes a high
备注:
点击查看摘要
Abstract:In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transforming verbose text into logically clear diagrams is essential. Scalable Vector Graphics (SVG) provide an effective representation for this purpose due to their editability and resolution independence. However, current research on Text-to-SVG generation remains hindered by three major challenges: (1) the scarcity of datasets for complex, logic-rich diagrams; (2) the absence of explicit layout priors, which leads to chaotic spatial arrangements; and (3) the lack of fine-grained visual feedback to validate rendered outputs and correct aesthetic defects. To address these challenges, at the data level, we introduce DocMeetSVG-100K, a large-scale SVG dataset tailored for document authoring and meeting review scenarios. At the model level, we propose GVR-Coder, a novel framework designed to generate high-quality logical diagrams from lengthy professional texts. Specifically, we adopt a curriculum-driven rejection sampling fine-tuning to progressively enhance the model's capability in modeling complex structures, while explicitly incorporating layout constraint knowledge during training. In addition, we introduce reinforcement learning from dual rendering feedback, a mechanism that provides implicit feedback through reward signals to jointly optimize structural complexity and visual aesthetics. Furthermore, we design a generate-verify-repair agent loop, which improves generation quality through explicit, fine-grained feedback and targeted refinement. Extensive experiments demonstrate that GVR-Coder outperforms competitive baselines and reliably produces logically coherent and visually appealing diagrams. Code and data are available at this https URL.
57. 【2607.28065】BladeYOLO: Wind Turbine Blade Defect Detection with Limited Annotations and Weak-Saliency Awareness
链接:https://arxiv.org/abs/2607.28065
作者:Yabin Xu,Fangtao Zhang,Fan Wang,Zhan Wang,Honghua Chen,Mingqiang Wei,Haoran Xie,Sam Kwong
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:remains highly challenging, real-world inspection scenarios, inspection scenarios due, detection remains highly, limited on-site data
备注: Accepted to IEEE TGRS, Code: [this https URL](https://github.com/zhangfangtao/BladeYOLO)
点击查看摘要
Abstract:Wind turbine blade defect detection remains highly challenging in real-world inspection scenarios due to limited on-site data and the subtle visual characteristics of defects. In practice, blade defects are often small-scale, low-contrast, and difficult to distinguish from complex backgrounds, which significantly limits the robustness of existing detectors. To address these challenges, we propose BladeYOLO, a defect detection framework for wind turbine blades. Specifically, we integrate a Vision Transformer (ViT) backbone initialized with DINOv3 self-supervised pre-trained weights into YOLOv12-L, enabling the transfer of large-scale generic visual priors to blade defect detection and improving feature representation under limited training annotations. To enhance the perception of subtle defects, we further develop a Mamba-guided Weak-Defect Enhancement module, which consists of a Detail-Enhanced Multi-scale Branch for preserving high-frequency structural cues and a Cross-Mamba module for progressively propagating high-level semantic guidance to shallow features. In addition, we introduce a lightweight Style-Injector module that captures environment-related style information via Fourier decomposition and injects it into selected ViT self-attention layers, thereby improving robustness against environment-induced appearance variations. Extensive experiments demonstrate that BladeYOLO achieves superior performance on the WTBlade-Defect dataset, with additional annotation-budget experiments showing its favorable performance under reduced training annotations. Evaluation on the public Wind Surface Defect dataset further provides supportive evidence for the cross-dataset robustness of BladeYOLO. In particular, on this public dataset, BladeYOLO outperforms the best competing method by 3.5\% in mAP$_{50}$ and 2.5\% in mAP$_{50-95}$.
58. 【2607.28064】Landmark shape spaces with induced metrics
链接:https://arxiv.org/abs/2607.28064
作者:Sarang Joshi,Peter W. Michor,Stefan Sommer
类目:Computer Vision and Pattern Recognition (cs.CV); Differential Geometry (math.DG)
关键词:spaces carrying Riemannian, carrying Riemannian metrics, right-invariant Sobolev metrics, Riemannian metrics descending, landmark shape spaces
备注:
点击查看摘要
Abstract:We present a unification of Kendall's landmark shape spaces, where rigid motions are factored out and scale fixed on landmark configurations equipped with Euclidean geometry, with landmark configuration spaces carrying Riemannian metrics descending from right-invariant Sobolev metrics on the diffeomorphism group. The resulting new landmark shape spaces achieve the defining properties of both approaches: The regularity of the descending metric prevents landmarks from colliding, the metric is defined in the ambient space independent of the number of landmarks, local rigid transformations are preserved, global rigid motions are removed, and scale fixed. To achieve this, we define a particular Sobolev-type operator, the screened elasticity operator, whose null-space consists exactly of the rigid motions, we show how this operator descends to achieve the desired geometry, and we present approaches to solving matching problems and computing geodesics numerically. The resulting construction allows the use of landmark configuration spaces with sufficiently regular metrics in applications while retaining the shape invariances that are a hallmark of Kendall's shape spaces.
59. 【2607.28058】mporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
链接:https://arxiv.org/abs/2607.28058
作者:Henglin Liu,Fangyuan Kong,Jing Wang,Yizhou Lin,Nisha Huang,Chang Liu,Xintao Wang,Pengfei Wan,Kun Gai,Xiu Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:improved visual quality, significantly improved visual, Direct Preference Optimization, Recent advances, diffusion-based video generation
备注: project page: [this https URL](https://henglin-liu.github.io/cIPO_vis/)
点击查看摘要
Abstract:Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit preference with temporally concentrated optimization.
60. 【2607.28039】ongueReenact: Geometry-Anchored Tongue Synthesis for Face Reenactment
链接:https://arxiv.org/abs/2607.28039
作者:MD Wahiduzzaman Khan,Mingshan Jia,Xiaolin Zhang,En Yu,Kaska Musial-Gabrys
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:systems achieve impressive, achieve impressive pose, Modern face reenactment, Modern face, reenactment systems achieve
备注:
点击查看摘要
Abstract:Modern face reenactment systems achieve impressive pose and expression transfer using geometry-driven representations. However, they largely ignore tongue dynamics, leading to anatomically inconsistent mouth interiors during speech and expressive motions. We introduce the first framework for cross-identity tongue dynamics transfer in face reenactment. We propose a foundation-model-assisted bootstrapping pipeline that produces a dedicated tongue segmentation model for in-the-wild reenactment without curated annotations. We further introduce a spatially constrained latent masked diffusion model for realistic tongue synthesis, with adaptive mask dilation for seamless mouth boundary transitions. Extensive experiments demonstrate improvements of more than two times over all baselines on every tongue-specific metric. We additionally propose a VLM-based evaluation protocol that replicates expert annotation at scale, confirming perceptual superiority across all ablation variants.
61. 【2607.28032】Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head Avatars
链接:https://arxiv.org/abs/2607.28032
作者:MD Wahiduzzaman Khan,Mingshan Jia,Xiaolin Zhang,En Yu,Kaska Musial-Gabrys
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Creating photorealistic animatable, digital human synthesis, Creating photorealistic, photorealistic animatable head, single image remains
备注:
点击查看摘要
Abstract:Creating photorealistic animatable head avatars from a single image remains a fundamental challenge in digital human synthesis. While recent 3D Gaussian Splatting methods have achieved promising results, they rely on external tracking pipelines whose latency is excluded from inference measurements. Furthermore, they adopt unified representations that entangle geometrically distinct facial regions, limiting both expressiveness and rendering fidelity. We propose SpiD (Split and Drive), a single-image Gaussian head avatar framework built on two disentanglement axes. The compute axis internalizes per-frame driving, eliminating external tracking dependency at inference. The feature axis decomposes the avatar into three specialized Gaussian branches, each modeling a geometrically distinct facial domain. Extensive experiments demonstrate consistently strong performance against state-of-the-art methods while achieving the fastest inference speed among all compared methods on a single GPU with the complete driving pipeline included.
62. 【2607.28030】MUL-T: Decoding Spatial Cellular Architecture in Multiplexed Tissue Images
链接:https://arxiv.org/abs/2607.28030
作者:Farzaneh Seyedshahi,Kai Rakovic,Adalberto Claudio Quiros,John LeQuesne,Ke Yuan
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Understanding tissue organisation, multiplexed imaging requires, imaging requires modelling, Understanding tissue, spatial context
备注:
点击查看摘要
Abstract:Understanding tissue organisation in multiplexed imaging requires modelling both cellular phenotypes and their spatial context. Existing approaches typically rely on handcrafted features, such as marker intensity statistics or cell-type proportions, which often fail to scale or generalise across cohorts with heterogeneous marker panels. We introduce MUL-T, a lightweight transformer framework that reframes tissue architecture as a masked contextual prediction task over discrete cell tokens. By learning contextualised [CLS] embeddings without task-specific supervision, the model captures higher-order cellular interactions while remaining computationally efficient. We evaluate MUL-T on several clinically relevant downstream tasks, including core-level tumour pattern classification, patient-level grading, PD-L1 positivity prediction, and cross-dataset treatment response prediction. Across tasks, MUL-T consistently outperforms classical feature-based baselines and achieves performance comparable to a foundation ViT model, despite substantially fewer parameters and lower training cost.
63. 【2607.28020】ENCORE: Event-Assisted Complementary Motion Refinement for Learned Video Compression
链接:https://arxiv.org/abs/2607.28020
作者:Shuhan Ye,Hongbin Yu,Chenqi Kong,Pingchuan Ma,Chong Wang,Jun Wan,Qixin Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Learned video compression, accurate temporal modeling, video compression relies, sampled RGB frames, relies on accurate
备注:
点击查看摘要
Abstract:Learned video compression relies on accurate temporal modeling to remove redundancy between adjacent frames. However, most existing codecs infer motion solely from discretely sampled RGB frames, making their estimates vulnerable to fast motion, blur, occlusion, weak texture, low illumination, and abrupt brightness changes. Event cameras asynchronously capture fine-grained intensity changes between RGB timestamps and therefore provide complementary evidence about inter-frame dynamics. We propose ENCORE, an Event-Assisted Complementary Motion Refinement framework for learned video compression. ENCORE first employs Complementary Motion Representation (CMR) to decompose aligned RGB-event features into common and modality-specific motion representations. Spatial Energy and Redundancy-Informed Calibration (SERIC) then identifies event-specific responses that are active and novel relative to RGB, suppresses weak or redundant evidence, and predicts a candidate flow correction. Finally, Energy-Aware Routing (EAR) determines where and how strongly the correction should refine the RGB flow. Events serve solely as an auxiliary modality for motion modeling, while RGB remains the only coding and reconstruction target. Experiments on BS-ERGB, HQ-EVFI, and CED demonstrate consistent gains across datasets and GOP lengths. On BS-ERGB, ENCORE achieves up to 20.80% PSNR-RGB and 22.14% MS-SSIM-RGB BD-rate savings, while retaining clear improvements on the other two datasets.
64. 【2607.28007】Beyond Classification: Pathology Foundation Models as Detection Encoders for Mitotic Figures
链接:https://arxiv.org/abs/2607.28007
作者:Sweta Banerjee,Alireza Teimoury,Nils Porsche,Alexandra K. Stoll,Viktoria Weiss,Niklas Hargarter,Jonas Ammeling,Thomas Conrad,Christoph Stroblberger,Christopher Kaltnecker,Robert Klopfleisch,Christof A. Bertram,Katharina Breininger,Marc Aubreville
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:typically unlabeled data, downstream classification tasks, yield regularized latent, regularized latent spaces, vast amounts
备注:
点击查看摘要
Abstract:Pathology foundation models (FMs) are models trained on vast amounts of typically unlabeled data and have been shown to yield regularized latent spaces that can be used effectively in downstream classification tasks. This is also true for the classification of mitotic figures vs. other cells. However, it is so far unclear if the latent space of current FMs provides features that are discriminant and spatially suitably resolved to also serve as a backbone for dense object detection paradigms. In this work, we investigate this question for common current pathology FMs (UNI, UNI2-h, Virchow, Virchow2, H-optimus-0, H-optimus-1) and compare their performance against a fully end-to-end trained baseline based on a ResNet50 architecture. We combine FM backbones with representatives of single stage, dual stage and self-attention-based detectors (RetinaNet, Faster R-CNN, Deformable DETR respectively) on the multi-domain MIDOG++ dataset, and on the TUPAC16 dataset as an out-of-domain case. We show that the H-optimus-0 and Virchow models yielded competitive performance, indicating that the latent spaces of current FMs, all trained on image-level self-supervision, are suitable for direct mitotic figure detection and may be slightly more robust on our out-of-domain test case. All code is made available publicly at this https URL.
65. 【2607.28005】Deep learning-based hierarchical insect classification using camera trap imagery
链接:https://arxiv.org/abs/2607.28005
作者:Zaki Mahfoud,Juan A. Chiavassa,Simon Walther,Florian Haselbeck,Ehsan Yaghoubi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Declining insect populations, time-consuming manual identification, monitoring increasingly urgent, reliable biodiversity monitoring, biodiversity monitoring increasingly
备注:
点击查看摘要
Abstract:Declining insect populations make reliable biodiversity monitoring increasingly urgent, yet monitoring of insect biodiversity is hampered by a lack of standardised data and by costly and time-consuming manual identification by expert entomologists. Deep learning-based image classifiers, processing data from automated non-lethal camera traps, have the potential to transform and scale insect biodiversity monitoring. However, challenges remain in acquiring expert-annotated datasets, developing model architectures that generalise well across diverse taxonomic levels and training models on highly imbalanced data. Hierarchical data also benefits from designing models that default to higher-confidence, coarser-level predictions, when uncertain about finer taxonomic levels. In this paper we address these challenges with a deep learning-based hierarchical classification model. First, we present a manually curated, long-tailed dataset of around one million images of insects, extracted from 1,801 camera-trap video recordings and annotated with a five-level, 34-class hierarchy. Further, we adapt a hierarchical classification model architecture to a five-level variable-depth hierarchy, with class-balanced weighting. Our model improves on non-hierarchical classifiers by leveraging biological taxonomy to extract granularity-specific visual features and makes hierarchy-consistent predictions to the deepest taxonomic level that meets a confidence threshold (T = 0.6). Our model achieved a per-level accuracy of 80-99% on test data, across five levels of hierarchy. Furthermore ...
66. 【2607.27982】ViP-Rig: Visual-Prompted Controllable Rigging
链接:https://arxiv.org/abs/2607.27982
作者:Zihan Qin,Mingze Sun,Yifan Mao,Jialei Xu,Jingfeng Guo,Changrong Hu,Wenbo Zhao,Junjun Jiang,Xianming Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:deformation behavior, inherently task-dependent, mesh may require, animation tasks, deformation behaviors
备注: 8 pages, 4 figures. Zihan Qin and Mingze Sun contributed equally. Xianming Liu is the corresponding author
点击查看摘要
Abstract:Rigging is inherently task-dependent because the same mesh may require different skeletons and deformation behaviors across animation tasks. In practice, artists often inspect an initial rig and repeatedly edit its skeletal structure and deformation behavior to meet specific animation requirements. Existing automatic methods primarily generate a plausible rig from geometry, offering limited explicit control over the resulting skeleton and deformation behavior. In this work, we present ViP-Rig, a visual-prompted framework that supports both prompt-first rigging and result-guided editing by injecting features extracted from user-drawn or edited 2D skeletal and rigidity prompts into frozen pretrained backbones. Specifically, ViP-Rig consists of two stages, Skeleton Generation and Skinning Prediction. In the first stage, the skeletal sketch is processed by the Dense-to-Compact Visual Prompt Encoding to produce compact, fixed-length conditioning tokens. The resulting tokens are injected into a frozen pretrained autoregressive generator through gated adapters to control joint placement and branching structure while preserving the generator's geometric prior. In the second stage, the rigidity map is processed using the same visual encoding design, while the pretrained skinning backbone remains frozen. The resulting tokens are symmetrically injected into the point and joint streams to modulate point-joint compatibility and the resulting skinning weights. Experiments on Articulation-XL2.0 and zero-shot evaluation on ModelsResource show that ViP-Rig more accurately recovers target skeletons and skinning weights than geometry-conditioned baselines under prompt-guided evaluation. Qualitative results further demonstrate explicit and localized control in both prompt-first rigging and result-guided editing.
67. 【2607.27974】Now You Have My Healthy Attention: A U-DiT for Brain-MRI Inpainting
链接:https://arxiv.org/abs/2607.27974
作者:Danilo Danese,Angela Lombardi,Tommaso Di Noia
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:BraTS Local Synthesis, ASNR-MICCAI BraTS Local, Local Synthesis, anatomically plausible completion, ASNR-MICCAI BraTS
备注:
点击查看摘要
Abstract:The ASNR-MICCAI BraTS Local Synthesis (Inpainting) task asks for the anatomically plausible completion of healthy brain tissue within a masked region of a T1-weighted MRI, providing a tumor-free anatomical reference for downstream analysis. As the task is scored by distortion metrics (SSIM, PSNR, MSE), we build a deterministic regression model and focus on giving it inductive biases tailored to inpainting. Our network follows the U-DiT principle of performing self-attention on a downsampled token grid: a volumetric encoder-decoder imports long-range context through a downsampled global self-attention block with three-dimensional rotary position embeddings, while convolutions and skip connections preserve high-frequency detail. Two ideas drive our results. First, we constrain the attention so that occluded ("void") tokens attend only to known-healthy tokens of the same volume, with a learned bias toward each query's contralateral homologue, forcing the completion to be inferred from observed anatomy rather than from other unknown regions. Second, we add a contralateral-symmetry input that supplies the mirrored healthy hemisphere as a patient-specific prior; since the brain is approximately bilaterally symmetric and lesions are typically unilateral, this prior improves the distortion metrics at matched structural similarity. On the official BraTS-2026 validation leaderboard our submission reaches a mean healthy-region SSIM of $0.864$, PSNR of $24.7$\,dB and MSE of $4.6{\times}10^{-3}$ over $219$ cases. We further analyse the residual smoothness inherent to distortion-optimal regression and discuss its implications for anatomical realism.
68. 【2607.27969】FootprintNet: State-Transition-Guided Dynamic Footprint Learning for Multi-temporal Remote Sensing Change Detection
链接:https://arxiv.org/abs/2607.27969
作者:Haotian Zhang,Hao Chen,Han Guo,Zhengxia Zou,Zhenwei Shi
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:entire observation period, existing MTCD methods, existing MTCD, remote sensing multi-temporal, entire observation
备注:
点击查看摘要
Abstract:Despite substantial progress in remote sensing multi-temporal change detection (MTCD), most existing MTCD methods still represent the dynamic process at each spatial location over the entire observation period using a single change category associated with the final observation. This implicit single-change assumption limits their ability to characterize regions of recurrent change closely related to human activities. To address this limitation, we introduce Urban Building Dynamics Detection (UBDD), which identifies building-change dynamic footprints, i.e., the temporal intervals in which changes occur, from multi-temporal imagery and produces pixel-wise classification masks. For regions undergoing two or more changes, UBDD introduces an independent multi-change class for unified representation, thereby enabling unified modeling of single- and multi-change processes. Furthermore, we propose FootprintNet, which abstracts building-change processes as interactions between latent states and actions, and imposes state-action transition constraints to guide the learning of causally coherent change trajectories. It further exploits temporal change-boundary cues to enhance feature contrast across boundary sides, thereby improving the discrimination among different dynamic footprints and enabling accurate detection of dynamic footprints. Moreover, we introduce the Building Change Dynamics Score (BCDS) to address the inability of conventional metrics to reflect the temporal proximity between predicted footprints and labels. It evaluates predictions according to their preservation of change semantics and temporal offsets from the corresponding labels. Extensive experiments on TSCD, MUDS, and WUSU demonstrate that FootprintNet outperforms current state-of-the-art methods. The code is available at this https URL.
69. 【2607.27959】FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
链接:https://arxiv.org/abs/2607.27959
作者:Bohan Hou,Haoqiang Lin,Xuemeng Song,Haokun Wen,Meng Liu,Yupeng Hu,Xiangyu Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR)
关键词:Multimodal Large Language, Large Language Models, Large Language, strong generalizable multimodal, generalizable multimodal processing
备注:
点击查看摘要
Abstract:Due to their strong generalizable multimodal processing and reasoning capabilities, Multimodal Large Language Models (MLLMs) have demonstrated significant potential as universal image retrievers, effectively addressing diverse real-world image retrieval tasks. Nevertheless, pioneering studies, while promising, overlook the potential of fine-grained context modeling and disentangled fine-tuning objectives in enhancing MLLMs' retrieval performance, particularly for complex tasks such as long-text-to-image retrieval, visual dialog retrieval, and composed image retrieval (CIR). Therefore, in this work, we propose an automated fine-grained multimodal quintuple dataset construction pipeline and a novel two-stage fine-grained multimodal fine-tuning strategy. The dataset generation pipeline produces a comprehensive CIR dataset with fine-grained image captions and modification text, facilitating fine-grained context modeling. Beyond the previously entangled fine-tuning paradigm, our approach separates the fine-tuning process into two distinct stages: (1) fine-grained context reasoning-oriented fine-tuning and (2) fine-grained retrieval-oriented fine-tuning. These stages aim to sequentially enhance the model's context understanding and query-target alignment capabilities, thereby improving retrieval performance. Extensive experiments across five datasets encompassing diverse and complex image retrieval tasks demonstrate the remarkable superiority of our method over existing approaches in zero-shot retrieval settings, even with a more lightweight MLLM backbone compared to those methods.
70. 【2607.27952】LAST: The Last Query Token Guides Visual Token Pruning for Edge-Cloud Collaborative MLLM Inference
链接:https://arxiv.org/abs/2607.27952
作者:Feng Yang,Xinrui Ju,Keyang Zhang,Xiandong Meng,Rongqun Lin,Howard Leung,Shiqi Wang,Haoliang Li,Chris Xing Tian
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:task-specific feature pipelines, edge devices encode, general-purpose cloud MLLM, devices encode visual, encode visual inputs
备注:
点击查看摘要
Abstract:Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs into tokens for a general-purpose cloud MLLM. However, dense visual-token sequences increase cloud-side inference costs. Existing pruning methods mainly target centralized inference: vision-driven methods can operate before cloud execution but are typically query-agnostic, whereas query-guided methods often rely on internal states of the target MLLM and cannot determine token relevance before transmission. Compact guidance models offer an alternative, but existing designs may require costly attention aggregation or auxiliary generation. We propose LAST, a training-free framework for query-dependent visual token pruning in edge-cloud collaborative MLLM inference. LAST uses a compact edge-side VLM as a guidance proxy and derives a lightweight importance signal from the last query token's attention to visual tokens. Under causal attention, the last query token can attend to the full visual sequence and the entire query context, enabling query-aware pruning without cloud-model access, autoregressive generation, or costly aggregation over multiple query positions. LAST then retains a diverse set of query-relevant visual tokens under a fixed token budget. We evaluate LAST on 11 multimodal benchmarks under multiple token budgets against pruning methods with different guidance strategies. Experiments show that LAST consistently achieves the strongest performance, preserving 95.4% of the full-token accuracy while retaining only 12.5% of the visual tokens, with low edge-side selection overhead and reduced cloud-side computation.
71. 【2607.27927】ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
链接:https://arxiv.org/abs/2607.27927
作者:Dongfu Yin,Rourou Su,Cong Zhao,Fei Yu
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:remains challenging due, Asymmetric Region Denoising, Feature Similarity Matching, detection remains challenging, asymmetric regions
备注:
点击查看摘要
Abstract:Reflection symmetry detection remains challenging due to interference from asymmetric regions and arbitrary orientations of symmetric patterns. Asymmetric regions introduce background clutter that disrupts symmetric pattern matching, whereas conventional convolutional neural networks lack rotation equivariance, leading to inconsistent feature representations under rotational transformations. To address these issues, we propose an Asymmetric Region Denoising (ARD) module and a Rotation Equivariant Feature Similarity Matching (REFSM) module. The ARD module suppresses asymmetric interference to refine symmetric patterns, while the REFSM module enhances rotation equivariance through feature similarity matching between original and rotated images. Specifically, our dual-input REFSM framework leverages rotation loss to maximize consistency between the score maps of original and rotated images, thereby enabling precise prediction of rotation-equivariant symmetry axes. Furthermore, we introduce GMSYM, a new benchmark dataset that categorizes images into diverse scenarios and incorporates various interferences to address the limitations of existing reflection symmetry detection benchmarks. Extensive experiments on four standard datasets (DENDI, NYU, LDRS, SDRW) and our proposed GMSYM dataset demonstrate that our method achieves state-of-the-art performance in both accuracy and robustness.
72. 【2607.27924】ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow
链接:https://arxiv.org/abs/2607.27924
作者:Dongxiu Liu,Haoyi Niu,Peng Cheng,Yuan Gao,Xirui Kang,Sangli Teng,Koushil Sreenath,Xianyuan Zhan
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:world, physical world, latent, space, ODE
备注:
点击查看摘要
Abstract:In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction, thereby exhibiting significant inefficiency in capturing the dynamics of physical world. We introduce Physical-Time Flow (\textbf{PT-Flow}), a novel approach that learns a continuous latent velocity field operating in physical time. Crucially, the underlying dynamics of sequential data are parameterized by an ordinary differential equation (ODE) embedded in a well-structured representation space. Under this paradigm, the prediction of future can be recast as temporal integration via an ODE solver in the compressed latent space. Building upon PT-Flow, we construct \textbf{ODEWorld}, a continuous-time latent world model that is both efficient and versatile. By extracting time-variant features and enforcing ODE properties on both the dynamical representation space and the latent velocity field, ODEWorld effectively addresses the long-standing representation collapse issue in latent world model literature. This also enables high-quality image reconstruction even after long-horizon prediction. Moreover, its continuous nature allows for arbitrary temporal resolution and even backward prediction, which is impossible for most discrete-time models. Lastly, ODEWorld can provide rich planning-oriented information to facilitate downstream policy learning. Comprehensive experiments demonstrate that ODEWorld successfully reconciles planning-conducive dynamics abstraction with visual realism, excelling in both video generation and robotic control. \href{this https URL}{Project Website}.
73. 【2607.27902】One Patch Is Enough: Reinforcement-Optimized Visual Token Grounding for MLLM-Based Scene Text Spotting
链接:https://arxiv.org/abs/2607.27902
作者:Rui Tang,Wentao Yang,Peirong Zhang,Yongxin Shi,Shun Zhang,Huiguo He,Lianwen Jin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Scene text spotting, spotting requires high-precision, text spotting requires, requires high-precision alignment, Scene text
备注: 15 pages, 11 figures. Accepted to ACM Multimedia 2026
点击查看摘要
Abstract:Scene text spotting requires high-precision alignment between textual recognition and spatial localization. While visual-token grounding has emerged as a promising formulation for Multimodal Large Language Models (MLLMs), the previous multi-patch paradigm often introduces redundant noise and localization ambiguity, particularly for dense or small text instances. To address this, we propose Single-Patch Text Spotting (SPaTS), a vision-centric framework that routes each text instance through a single anchor visual token and then recovers geometry via full-image refinement. To accurately identify this anchor without oracle labels, we introduce Single-Patch Selective Optimization (SPaSO), a reinforcement learning framework that optimizes discrete visual-token selection using patch-level rewards. To further improve representation robustness and localization precision, we introduce Directional Embedding Alignment (DEA) to suppress unstable norm bias by decoupling feature magnitude and direction, and Patch-Enhanced Decoding (PED) to fuse the routed anchor with language semantics and cross-attend over the full-image feature map for geometry-aware boundary regression beyond coordinate-space surrogates. Extensive experiments demonstrate that SPaTS consistently and significantly outperforms both frontier closed-source MLLMs and OCR MLLMs. Code will be released soon.
74. 【2607.27898】CoRE-UIR: Prior-guided common and residual experts for efficient all-in-one remote sensing image restoration
链接:https://arxiv.org/abs/2607.27898
作者:Zaiyan Zhang,Qiangqiang Yuan,Jie Li,Ziyang Lihe,Yu Wan,Yuzeng Chen,Xin Su,Liangpei Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Remote sensing images, unmanned aerial vehicles, sensing images acquired, jointly induce global, induce global distribution
备注: Accepted by ISPRS Journal of Photogrammetry and Remote Sensing
点击查看摘要
Abstract:Remote sensing images acquired by unmanned aerial vehicles (UAVs) and satellites are often degraded by adverse weather, illumination variation, and imaging artifacts, which may co-occur and jointly induce global distribution shifts and local structural corruption. Although All-in-One image restoration offers an appealing unified alternative to task-specific pipelines, existing methods still suffer from weak or implicit degradation cues and parameter redundancy caused by full-rank multi-expert designs with overlapping restoration behaviors. We propose CoRE-UIR (Common and Residual Experts for Universal Image Restoration), a prior-guided global-local framework centered on the Common-and-Residual Expert Block (CoRE). CoRE explicitly decomposes restoration capacity into a common dense expert for degradation-invariant restoration and low-rank residual experts for degradation-specific compensation, enabling adaptive specialization without redundant expert replication. Built on this design, Degradation Prior Embedding (DPE) adapts frozen CLIP features into an explicit restoration-oriented prior, while Global Feature Modulation (GFM) aligns global feature statistics before local residual compensation. We also construct MDVD-108K (Multi-Degradation VisDrone), a large-scale UAV restoration dataset covering both single and compound degradations, together with a real-world test set. Extensive experiments on multiple datasets show that CoRE-UIR improves the overall average PSNR by 1.05 dB while running 11.83$\times$ faster and reducing peak memory by 85.3% relative to the strongest baseline, BaryIR, thereby maintaining a favorable quality-efficiency trade-off. Evaluations on downstream tasks and unseen degradation also validate the generalizability of CoRE-UIR. The code and dataset will be released at this https URL.
75. 【2607.27897】Unifying Adversarially Robust Model Experts in Vision-Language Models
链接:https://arxiv.org/abs/2607.27897
作者:Nguyen Duc Thai,Junhao Dong,Sua Qi Rong,Hua Yu,Yew-Soon Ong
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:applications and deployment, problem for real-life, real-life applications, adversarial, CLIP
备注:
点击查看摘要
Abstract:Vision-language models (VLMs), such as CLIP, are vulnerable to adversarial attacks, posing a serious problem for real-life applications and deployment. Adversarial fine-tuning emerges as a prominent defense method; however, different fine-tuning strategies often produce specialized models with distinct robustness characteristics. Each fine-tuned model in turn thrives in some evaluation settings but falters on others, limiting their defensive capabilities. We refer to these specialized fine-tuned models as robust model experts and propose a collaborative adversarial fine-tuning framework: CARE - Collaborative Adversarial Robustness fine-tuning using Embedding alignment. CARE maintains multiple experts during training, enables knowledge exchange through embedding-space harmonization, and consolidates the learned knowledge into a single unified robust model. Experts benefit from one another while preserving their individual specializations, enabling the final model to inherit complementary robustness properties. In this paper, we demonstrate CARE on two different adversarial fine-tuning strategies with complementary robustness behaviors. Extensive experiments on classic image classification and downstream vision-language tasks display the effectiveness of our approach, with CARE being able to outperform individually learned model experts. The results suggest that collaborative learning across model experts is a promising direction for improving adversarial robustness.
76. 【2607.27895】MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
链接:https://arxiv.org/abs/2607.27895
作者:Jinpeng Hu,Erqiang Wang,Shan Wang,Zhuo Li,Peipei Song,Xun Yang,Meng Wang
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Mental health understanding, interpersonal context, latent psychological states, Mental health, health understanding
备注:
点击查看摘要
Abstract:Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.
77. 【2607.27882】DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection
链接:https://arxiv.org/abs/2607.27882
作者:Zihao Cai,Xinghan Li,Ruiyan Yang,Xue Song,Haijun Shan,Jingjing Chen
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:preserving knowledge acquired, AI-generated image detectors, generative models continue, continue to evolve, AI-generated image
备注:
点击查看摘要
Abstract:As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge acquired from previous ones. This continual learning setting is particularly challenging because forensic traces are often subtle and generator-specific, making detectors highly vulnerable to catastrophic forgetting. Existing methods primarily address this problem by stabilizing feature representations, implicitly treating forgetting as a representation-level issue. In this paper, we show that this perspective is incomplete. We demonstrate that even when feature representations remain discriminative, the decision boundary can progressively drift as the classification head is continually optimized on new domains. These two effects jointly give rise to a compound failure mode, termed Dual Degradation. To overcome this challenge, we propose DECODE, a decoupled continual detection framework that jointly mitigates representation- and decision-level forgetting. Specifically, we introduce Subspace Diversity Regularization (SDR) to preserve diverse forensic representations and Closed-Form Decision Alignment (CDA) to recalibrate the shared classification head after each adapter merge without manual hyperparameter tuning. Extensive experiments on 19 generative domains show that DECODE achieves an average accuracy of 99.36% with only 0.39% forgetting, while further generalizing to 11 unseen generators with 95.36% accuracy.
78. 【2607.27865】Learning to Understand Body Language from Flight through Robust 3D Avatar Placing
链接:https://arxiv.org/abs/2607.27865
作者:Dragos Costea,Alina Marcu,Cristina Lazar,Marius Leordeanu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:intelligent aerial robots, socially intelligent aerial, Perceiving human motion, Perceiving human, aerial robots
备注:
点击查看摘要
Abstract:Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We introduce Drones2BodyLanguage, a dataset grounding human motion in real UAV footage: avatars manifesting ten communicative intents are placed into unmodified 4K drone scenes with metrically correct position, scale and orientation, maintained over hundreds of frames of camera motion. Enabling it is a lightweight geometric world model of the local scene - semantically selected anchors lifted to 3D through streaming monocular depth - in which a placement point is predicted as an affine anchor combination with provably rigid-invariant weights, and re-rendered under an SVD-fitted ground rotation. Across twelve architectures on scene- and motion-disjoint splits, training on placed data lifts mean intent accuracy by a wide margin for real, retargeted and generated motion alike, with gains confirmed on two in-the-wild scenes.
79. 【2607.27857】EEG-EditBench: Probing Visual Information in EEG-Image Retrieval Models with Controlled Image Edits
链接:https://arxiv.org/abs/2607.27857
作者:Kaifan Zhang,Lihuo He,Yuqi Ji,Junjie Ke,Lukun Wu,Tianhao You,Xinbo Gao
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:semantically diverse candidates, achieved strong performance, diverse candidates, performance in identifying, semantically diverse
备注: Main paper with supplementary material. Code: [this https URL](https://github.com/XiaoZhangYES/EEG-EditBench) . Dataset: [this https URL](https://huggingface.co/datasets/xiaozgg/EEG-EditBench)
点击查看摘要
Abstract:Recent EEG-to-image retrieval models have achieved strong performance in identifying viewed images from semantically diverse candidates. Yet such success does not reveal what visual information supports the match. A model may readily identify a cheetah among tools, plants, and vehicles, but can it still distinguish the viewed cheetah from the same scene with the cheetah replaced by a dog? Motivated by this question, we introduce EEG-EditBench, a diagnostic benchmark that examines this question through controlled edits of object identity, attributes, background, and object presence. Built from the 200 THINGS-EEG2 test images, EEG-EditBench contains 2,137 quality-controlled edits and evaluates eight representative EEG visual decoding models. Our results show that strong standard retrieval does not consistently transfer to edit-based evaluation, with fine-grained attribute changes presenting the greatest challenge. EEG-EditBench reveals model behavior hidden by aggregate retrieval accuracy and provides a controlled basis for studying what visual information EEG-image models preserve. The code and complete dataset are publicly available.
80. 【2607.27856】Benchmarking Foundation and Large Language Models for Few-Shot Medical Image Segmentation
链接:https://arxiv.org/abs/2607.27856
作者:Jinghong Liu,Yuchuan Deng,Fanping Liu,Meng Huang,Xirong Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:FS-MIS solutions, existing FS-MIS solutions, aims to segment, regions of interest, FS-MIS solutions span
备注:
点击查看摘要
Abstract:Few-shot medical image segmentation (FS-MIS) aims to segment novel regions of interest (ROIs) from a few annotated support examples. Despite rapid progress, existing FS-MIS solutions span diverse paradigms but are evaluated under inconsistent settings, leaving their relative effectiveness unclear. We introduce FAME, a unified benchmark for evaluating FS-MIS solutions, covering specialists, SAM-based methods, CLIP-based methods, and MLLM-based methods. FAME contains 14,958 test samples across 7 anatomical sites, 9 imaging modalities, and 14 ROI categories, and evaluates models under zero-shot and ten-shot settings with additional assessment of target-absence recognition and generalization under covariate and semantic shifts. Our evaluation reveals several findings. First, effective few-shot segmentation depends on how models exploit support examples: direct visual adaptation generally outperforms prompt-based strategies. Second, increasing support examples improves performance only when models can effectively utilize them. Third, semantic transfer remains substantially more challenging than imaging-domain adaptation, and strong localization ability does not necessarily imply reliable target-absence recognition. We hope FAME provides a comprehensive understanding of current FS-MIS solutions and facilitates the development of more effective and reliable few-shot medical segmentation methods.
81. 【2607.27854】Simplifying Neural Networks During Training
链接:https://arxiv.org/abs/2607.27854
作者:Lorenzo Sciandra,Samuele Fonio,Roberto Esposito
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:modern machine learning, neural networks remains, overparameterized deep neural, Understanding and exploiting, deep neural networks
备注: Preprint, submitted to a journal
点击查看摘要
Abstract:Understanding and exploiting the training dynamics of overparameterized deep neural networks remains a central challenge in modern machine learning. Recent evidence on Neural Collapse (NC) shows that class representations and classifiers exhibit highly structured geometry, while the Tunnel Effect suggests that only a subset of layers is essential for feature extraction. We combine these two perspectives and propose an NC-inspired training framework for simplifying deep networks during training. Our method monitors representation dynamics through the Inverse Fisher Criterion, a stable and efficient proxy for the variability collapse behavior, to identify both the split point between feature extraction and classification and the training stage at which simplification becomes viable. We then replace the trailing layers with a lightweight classification head and continue training the reduced model. Experiments on image-classification benchmarks across MLP, VGG, and ResNet architectures show that the proposed method achieves substantial parameter reductions while maintaining accuracy comparable to that of the full model. Code to reproduce the experiments can be found at: this https URL.
82. 【2607.27843】VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection
链接:https://arxiv.org/abs/2607.27843
作者:Songsong Duan,Xi Yang,Nannan Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Camouflaged Object Detection, segment camouflaged objects, Camouflaged Object, camouflaged objects, segment camouflaged
备注: Accepted by ECCV 2026
点击查看摘要
Abstract:Camouflaged Object Detection (COD) aims to identify and segment camouflaged objects in complex environments, which are often concealed because their color and texture are similar to the background. Several existing COD methods introduce depth maps to boost detection performance via learning complementary RGB-D features, ignoring modality-specific characteristics of concealed objects in the depth domain. To address this issue, we propose a depth collaborative network, called VCP-DCN, to mine distinguishable multi-modality features beyond visual concealed prototype in depth domain. Specifically, VCP-DCN progressively performs multi-modality alignment, interaction, and fusion for the COD task. In the \textbf{alignment} stage, we propose a Separable Prototype Embedding (SPE) module to learn modality-consistency and modality-specific RGB/depth prototype tokens through prototype contrastive learning. Furthermore, we develop a Multi-modality Dual Attention (MDA) module to enhance the cross-modal feature representation through local response maps between modality-consistency RGB/depth prototype tokens and visual tokens on the \textbf{interaction} stage. Finally, we design a Depth Adaptive Injection (DAI) module to adaptively measure contribution of RGB/depth features with a decision-making mechanism, which calculates similarity distance between RGB/depth modality-specific prototype tokens and modality-consistency ones on the \textbf{fusion} stage. Extensive experiments demonstrate the effectiveness of our VCP-DCN on three authoritative datasets.
83. 【2607.27842】FeatFix: Reuse What You Verify through Local Exact-Feature Correction for Faster Cached Diffusion Inference
链接:https://arxiv.org/abs/2607.27842
作者:Hanshuai Cui,Zhiqing Tang,Zhi Yao,Qianli Ma,Fanshuai Meng,Weijia Jia
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:remains computationally intensive, iterative denoising process, denoising process remains, process remains computationally, generate high-quality images
备注:
点击查看摘要
Abstract:Diffusion models are widely used to generate high-quality images and videos, but their iterative denoising process remains computationally intensive. A growing class of training-free accelerators reduces this cost by reusing cached intermediate features or forecasting future ones. To control draft drift, these methods sometimes compute an exact block feature for verification. Yet the resulting exact feature is typically used only to measure discrepancy or guide a later decision and is then discarded. We find that this previously computed feature can instead be reused for correction. Forwarding it at the verification site resets the local draft residual and reduces downstream feature error. Based on this observation, we introduce FeatFix, a local exact-feature correction method for cached diffusion inference. FeatFix operates at a fixed sparse set of layer--timestep sites. At each selected site, it replaces the complete draft block output with the exact output computed from the same incoming state, avoiding token- or channel-level partial replacement and full-timestep recomputation. Experiments across four image and video backbones show that FeatFix consistently accelerates generation, achieving a speedup of up to $6.70\times$ over Vanilla while maintaining competitive output quality.
84. 【2607.27835】SAFViT: Spatial Attention Fusion Gating for Vision Transformer-Based Nucleus Segmentation and Classification
链接:https://arxiv.org/abs/2607.27835
作者:Harshit Mittal,Arash Rabbani
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Accurate cell segmentation, enabling quantitative tissue, quantitative tissue analysis, Accurate cell, digital pathology
备注:
点击查看摘要
Abstract:Accurate cell segmentation and classification are foundational to digital pathology, enabling quantitative tissue analysis for diagnosis and treatment planning. Encoder-decoder architectures that fuse multi-scale features through skip connections have become the dominant paradigm for this task, yet standard direct skip connections treat every spatial location equally, which leads to redundant and potentially conflicting information reaching the decoder. To overcome this problem, various gating mechanisms have been introduced, but most of them operate solely on filtering encoder information, neglecting the benefit of global contextual information from the decoder. This study proposes replacing conventional skip connections in a CellViT-based model with a novel Spatial Attention Fusion (SAF) Gating module. Each SAF gate concatenates the encoder skip and upsampled decoder features, compresses them through two pointwise convolutions with an intermediate ReLU, and applies a channel-wise softmax to produce a per-pixel "heatmap of trust" that sums to unity at every spatial location, allowing the network to learn where each source is most trustworthy. The resulting fused features improve the model's ability to detect the minority "Dead" class, which in turn enhances the multi-class panoptic quality (mPQ) on the PanNuke dataset. SAF Gating is compared against six gating alternatives including no gating, attention gates, squeeze-and-excitation, CBAM, cross-attention, and attentional feature fusion on PanNuke and MoNuSeg datasets. SAF Gating achieves the highest mPQ (0.471), a gain driven primarily by a 14.5-point improvement in Dead-class F1 score compared to ungated CellViT baseline.
85. 【2607.27830】hinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
链接:https://arxiv.org/abs/2607.27830
作者:Zhongkuan Mao,Xianjie Liu,Tianyu Meng,Yidong Wang,Wenzhuo Zhao,Ronghao Xian,Yao Jiang,Fei Shen,Junfeng Fang,Yong Dai,Yi Zhang,Keren Fu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:High-resolution visual question, failing multimodal large, multimodal large language, visual question answering, insufficient evidence acquisition
备注:
点击查看摘要
Abstract:High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a \textbf{training-free, single-visual-pass} evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on V$^*$Bench, HRBench-4K, and HRBench-8K by \textit{+3.1}, \textit{+3.0}, and \textit{+2.7} points while reducing the average peak memory by about 4,GB. On Qwen2.5-VL-7B, it improves the three benchmarks by \textit{+9.9}, \textit{+4.6}, and \textit{+5.5} points, raising the cross-benchmark mean from 72.5 to 79.1. With the ZwZ-8B base model, Thinking-Once reaches a mean score of 82.7. Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces V$^*$Bench inference time by \textbf{97.2\%} while improving the cross-benchmark mean from 77.8 to 79.1. These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.
86. 【2607.27826】Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding
链接:https://arxiv.org/abs/2607.27826
作者:Shiwei Gan,Lichen Wang,Xiao Liu,Yafeng Yin,Kuizhuang Liu,Sanglu Lu,Lei Xie
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Recent advances, sign language, Sign Language Question, led to remarkable, remarkable progress
备注:
点击查看摘要
Abstract:Recent advances in sign language (SL) understanding (SLU) have led to remarkable progress in tasks such as continuous SL recognition and SL translation. However, these tasks are designed with predefined objectives, requiring models to learn a fixed mapping from sign videos to glosses or spoken-language sentences. As a result, they provide only a limited assessment of whether a model truly understands the semantic content of SL videos. To address this limitation, \textbf{we first propose a new task, Sign Language Question Answering (SLQA)}, which evaluates SL understanding by requiring models to answer arbitrary natural language questions about SL videos. Unlike previous SLU tasks, SLQA provides a more flexible and comprehensive evaluation framework that assesses multiple reasoning capabilities beyond recognition and translation. To facilitate this task, \textbf{we further construct two SignQA benchmarks} based on PHOENIX14T and CSL-Daily by automatically generating question-answer pairs from existing gloss and sentence annotations using carefully designed templates. The resulting datasets cover five complementary question categories, including position reasoning, structural reasoning, visual search, gloss recognition, and translation understanding. \textbf{Finally, we propose a simple yet effective baseline model} equipped with a Question-Conditioned Modulated Temporal Downsampling module and an in-domain knowledge transfer strategy, enabling effective knowledge transfer from existing SLU tasks while enhancing question-aware temporal feature modeling. Extensive experiments demonstrate that our baseline consistently outperforms representative vision-language models across all question categories, establishing a strong benchmark for future research on SLQA. Datasets are available at:{this https URL}.
87. 【2607.27823】Hallucinations Leave a Grounding Signature:Verifier-Guided Decoding for Selective Object Correction
链接:https://arxiv.org/abs/2607.27823
作者:Lei Yang,Xinze Liu,Dayan Wu,Ding Wang,Hengjie Zhu,Zihao Zhang,Tianzhu Hu,Hanqi Wu,Peng Fu,Zheng Lin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large vision-language models, Large vision-language, Intrinsic Grounding Signature, object, Large
备注:
点击查看摘要
Abstract:Large vision-language models (LVLMs) often hallucinate objects that are absent from an image. Despite recent progress, existing mitigation methods still lack reliable object-level grounding diagnostics and therefore tend to apply coarse-grained interventions, which can impair visual understanding, shorten responses, and reduce coverage of genuinely grounded objects. The key challenge is thus to detect, during generation, whether each emerging object mention is supported by reliable visual evidence, so that hallucination can be mitigated selectively. Yet output confidence reflects next-token plausibility rather than visual support, allowing language priors to make absent objects appear certain. We show that the missing diagnostic evidence is encoded in an Intrinsic Grounding Signature (IGS), a distributed signed attention pattern that remains informative for such confident hallucinations. Based on IGS, we propose Verifier-Guided Decoding (VGD), a decoding framework in which a lightweight verifier examines each emerging object mention, rolls back the KV cache when the mention is identified as high risk, suppresses the object and its synonyms, and regenerates the affected continuation. Because VGD intervenes only on object mentions identified as high risk, it reduces object hallucination while preserving the model's original visual understanding and grounded object coverage. Experiments on CHAIR and AMBER-G show that VGD achieves state-of-the-art object hallucination reduction: at @rec90, it cuts AMBER-G CHAIR by 43.6\% while retaining 99.6\% of grounded-object coverage, and reduces CHAIR-MSCOCO CHAIR$_i$/CHAIR$_s$ by 37.0\%/30.4\% without shortening captions.
88. 【2607.27811】SPFM-Net: Semantic-Prior-Guided Frequency-Constrained Mamba for Invisible Watermark Attack
链接:https://arxiv.org/abs/2607.27811
作者:Chunpeng Wang,Yanan Shi,Zhiqiu Xia,Jidong Yang,Suo Gao,Qi Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:constrained restoration networks, predefined signal-processing operations, locally constrained restoration, Existing watermark attacks, attacks typically rely
备注:
点击查看摘要
Abstract:Existing watermark attacks typically rely on predefined signal-processing operations or locally constrained restoration networks, making it difficult to capture the long-range dependencies of globally distributed watermark signals and resulting in an unfavorable trade-off between removal effectiveness and visual fidelity. In this paper, we propose SPFM-Net, a semantic-prior-guided and frequency-constrained Mamba framework for invisible watermark attack. SPFM-Net first employs high-ratio masking to disrupt the spatial coherence of invisible watermark signals, and then utilizes a partially fine-tuned pretrained Masked Autoencoder to reconstruct semantically consistent image from sparse observations while suppressing watermark-related information. A Multi-scale Residual Frequency Feature Interaction module subsequently aggregates watermark-related residual features across multiple receptive fields, while adaptively suppressing responses from watermark-irrelevant regions. To further capture the long-range dependencies of globally distributed watermark signals, a lightweight Mamba-based Global State-space Feature Modeling (GSFM) unit is introduced to separate watermark-related features from natural image content and suppress the remaining watermark traces. In addition, SPFM-Net is optimized using a multi-level objective that jointly imposes spatial-, frequency-, and edge-domain constraints, enabling effective watermark suppression while preserving perceptual quality. Extensive experiments on representative spatial-domain, transform-domain, orthogonal moment-based, and deep learning-based watermarking schemes demonstrate that SPFM-Net achieves a favorable trade-off between watermark attack effectiveness and perceptual fidelity.
89. 【2607.27806】LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
链接:https://arxiv.org/abs/2607.27806
作者:Zhilin Wu,Zhangkai Ni,Chengmei Yang,Longzhen Yang,Yihang Liu,Ying Wen,Lianghua He
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:undergo multiple imaging, multiple imaging examinations, clinical practice, successive visits, undergo multiple
备注: 23 pages, 17 figures, 7 tables. Code and data: [this https URL](https://github.com/pepperbubble/LoMeVQA)
点击查看摘要
Abstract:In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response. However, despite the rapid advancement of multimodal large language models (MLLMs), longitudinal medical visual reasoning remains largely underexplored. To fill this gap, we propose LoMeVQA, a comprehensive benchmark consisting of 206K longitudinal visual question answering (VQA) pairs for temporal medical image analysis. LoMeVQA covers five tasks: progress classification, progress description, progress report generation, differential region grounding, and differential region description. To construct the dataset, we develop an automated pipeline that (1) organizes patient records chronologically, (2) extracts clinically meaningful entities via a medical knowledge graph, and (3) models their temporal evolution to guide large language models in generating high-quality longitudinal VQA pairs. Extensive evaluations demonstrate that both general-purpose and medical-domain MLLMs perform poorly on LoMeVQA, revealing substantial limitations in temporal reasoning. To address these limitations, we introduce MedLong-8B, which achieves state-of-the-art performance across all tasks. Beyond benchmarking, we conduct detailed analyses that uncover key failure modes and shed light on how to improve longitudinal medical visual reasoning. Our data is available at: this https URL
90. 【2607.27800】FDDWAN: A Frequency-Decoupled Diffusion Network for Watermarking Attack
链接:https://arxiv.org/abs/2607.27800
作者:Chunpeng Wang,Yuxin Li,Xiaoyu Wang,Jidong Yang,Suo Gao,Qi Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Existing invisible watermark, Existing invisible, invisible watermark removal, Watermark Attack Network, Preliminary Attack Module
备注:
点击查看摘要
Abstract:Existing invisible watermark removal methods often struggle to accurately capture the watermark-bearing features, leading to an unfavorable trade-off between watermark suppression and perceptual fidelity. In this paper, we propose the Frequency-Decoupled Diffusion Watermark Attack Network (FDDWAN), a coarse-to-fine framework that performs watermark removal through wavelet-domain decomposition and residual diffusion refinement. In the initial stage, the Wavelet-based Frequency-domain Preliminary Attack Module (WFPAM) decomposes the watermarked image into low- and high-frequency subbands and applies frequency-specific attack strategies tailored to their respective contributions to watermark robustness and perceptual quality. In the next stage, the Frequency-domain Residual Diffusion Attack Module (FRDAM) separately models the residual distributions between the preliminarily attacked outputs and the corresponding watermark-free references during training. Rather than reconstructing the entire image, FRDAM selectively refines frequency-domain residuals, directing the diffusion process toward the remaining watermark related discrepancies while minimizing modifications to image content. Extensive experiments on CelebA and ImageNet across four representative watermarking schemes demonstrate that FDDWAN achieves a more favorable trade-off between watermark removal effectiveness and visual fidelity than conventional and learning-based attack methods.
91. 【2607.27779】CXR-Retrieve: Compositional Text-to-Image Retrieval in Chest Radiography
链接:https://arxiv.org/abs/2607.27779
作者:Tomer Erez,Moshe Kimhi,Chaim Baskin,Ehud Rivlin
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large chest radiography, chest radiography archives, Large chest, structured clinical annotations, radiography archives
备注:
点击查看摘要
Abstract:Large chest radiography archives are difficult to search because most studies are paired only with free-text reports rather than structured clinical annotations. Vision-language models offer a natural interface for text-to-image retrieval, but current biomedical models are primarily optimized for report-to-image matching rather than for satisfying short clinical search queries. This creates an objective mismatch: a model may retrieve images related to words in the query while failing to satisfy the full clinical constraint, especially for conjunctions and negations such as ``atelectasis and no pneumonia.'' We introduce CXR-Retrieve, a structured benchmark for compositional chest X-ray text-to-image retrieval. The benchmark contains 5,159 test images from the official test-split of MIMIC-CXR-JPG and 145 textual queries spanning single and conjunction findings, both positive and negative. Relevance is defined by whether a retrieved image satisfies all asserted pathology constraints, rather than by whether it matches a paired report. We further propose a label-aware contrastive fine-tuning objective for clinical retrieval. Our method attracts image-text pairs with compatible asserted pathology constraints, including shared confirmed absences, while explicitly repelling contradictory pairs. Starting from the in-domain CXR-CLIP checkpoint, our method improves Precision@5 over CXR-CLIP by 8.5 percentage points on two-pathology conjunctions and by 22.0 percentage points on negation queries. These results show that reliable chest X-ray retrieval requires training objectives that model not only which findings are mentioned, but also how they are clinically asserted.
Subjects:
Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2607.27779 [cs.CV]
(or
arXiv:2607.27779v1 [cs.CV] for this version)
https://doi.org/10.48550/arXiv.2607.27779
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
92. 【2607.27764】Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation
链接:https://arxiv.org/abs/2607.27764
作者:Shuhuan Chen,Xiangyu Zhu,Weisong Zhao,Siran Peng,Tianshuo Zhang,Haoyuan Zhang,Haichao Shi,Xiao-Yu Zhang,Zhen Lei
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Publishing private face, Publishing private, expose identity information, Private Face Distillation, training dataset publication
备注:
点击查看摘要
Abstract:Publishing private face recognition~(FR) training datasets is privacy-sensitive because faces expose identity information. Private FR training dataset publication mitigates this risk by releasing protected proxies as substitutes for private training faces. However, training FR models with such data introduces an identity paradox: \emph{the identity cues that make released faces useful for recognition supervision are also the cues that make them linkable to real individuals.} A protected face should be decoupled from the original identity, yet still behave as a reliable identity sample for training. Removing these cues too aggressively may destroy the class structure needed for recognition learning, whereas preserving them too faithfully may increase source-identity linkability. We argue that this paradox stems from conflating source-aligned identity semantics with recognition-useful proxy identity geometry. The former should be suppressed to reduce linkage to private individuals, while the latter should be preserved for FR learning. Based on this insight, we propose \textbf{Private Face Distillation}, an identity-decoupling and geometry-preserving framework. It uses Orthogonal Geometry Preservation to construct decoupled proxy identities from private identity representations while maintaining hyperspherical geometry, and Relational Topology Alignment to preserve identity relations for recognition learning. Experiments across multiple domain-shifted FR scenarios show that Private Face Distillation achieves stronger utility than the evaluated publication baselines. On IJB-C surveillance, it improves $\mathrm{TAR}@\mathrm{FAR}{=}1\text{e-}{3}$ by 3.94\% over the baseline while reducing source-identity linkability. These results suggest that private FR training dataset publication should decouple source-identity correspondence while preserving proxy identity geometry.
93. 【2607.27763】DS@GT ARC at ImageCLEFmedical 2026: Architectural Diversity for Concept Detection and Foundation-Model Scaling for Caption Prediction in Medical Image Analysis
链接:https://arxiv.org/abs/2607.27763
作者:Bowen Wang,Youwen Zhang,Ritesh Mehta
类目:Computer Vision and Pattern Recognition (cs.CV); Information Retrieval (cs.IR); Machine Learning (cs.LG)
关键词:Concept Unique Identifiers, UMLS Concept Unique, assigning UMLS Concept, generating natural-language captions, Caption Prediction
备注: 21 pages, 9 figures
点击查看摘要
Abstract:We describe the DS@GT submissions to the ImageCLEFmedical Caption 2026 challenge, which continues a long-running benchmark on the ROCOv2 dataset with two tracks: Concept Detection (Task 1), assigning UMLS Concept Unique Identifiers (CUIs) to radiology images, and Caption Prediction (Task 2), generating natural-language captions. For Task 1, our primary submission was a three-way late-fusion ensemble of ConvNeXt-V2, BiomedCLIP ViT-B/16, and DenseNet-169 with a regularized ''Honest Threshold Tuning'' procedure designed to avoid validation overfitting on rare concepts; this submission ranked first on the official submission with a primary $F_1$ of $0.5790$ and a secondary $F_1$ of $0.9657$. In parallel, we submitted a training-free KNN retrieval pipeline over frozen BiomedCLIP embeddings, which reached a primary $F_1$ of $0.5780$ and a secondary $F_1$ of $0.9599$-essentially matching the fine-tuned ensemble on the primary track at a fraction of the cost. For Task 2, our submissions included a fine-tuned Gemma-3 27B model (overall $0.3571$, ranking third in the official submission), a fully fine-tuned BLIP pipeline with custom Vizwins merging ($0.3564$), and a zero-shot MedGemma-4B run with a PubMed-style prompt ($0.3186$), spanning a wide range of model scales and training costs. Code: this https URL.
94. 【2607.27761】DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement
链接:https://arxiv.org/abs/2607.27761
作者:Shubin Ma,Liang Zhao,Chuanye He,Zhenjiao Liu,Liang Zou,Lin Yuanbo Wu,Yu Shao
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:widespread research interest, attracted widespread research, recent years, research interest, attracted widespread
备注: 8 pages, 4 figures. Accepted by ACM Multimedia 2026
点击查看摘要
Abstract:In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across different views often suffer from misalignment, leading to the partial view alignment problem (PVAP). To mitigate the impact of view asymmetry and irrelevant samples, this paper proposes a framework for partial multi-view clustering via dual alignment and structure enhancement (DAS-PMVC), which leverages view structure consistency and semantic relevance. Specifically, DAS-PMVC includes three parts: \textbf{anchor graph structure alignment}, where sample joint embedding representations with consistent latent space are derived from anchor point relationships for initial view alignment; \textbf{structure-enhanced feature learning}, where the model learns view structure information through pretraining and combines multi-view graph convolutional networks to further extract deep latent features from the aligned graph structure to improve the discriminative power of representations; and \textbf{a dual alignment strategy}, where initial alignment is performed through the anchor graph in the pretraining phase, and contrastive learning loss and the Hungarian algorithm are introduced in the training phase to further optimize the alignment of latent features. Experimental results on various datasets demonstrate that the DAS-PMVC framework outperforms existing state-of-the-art methods in clustering performance, showcasing its effectiveness and superiority.
95. 【2607.27755】EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder
链接:https://arxiv.org/abs/2607.27755
作者:Jaehun Jung,Wonjun Kim
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:address the problem, problem of recovering, head pose, full-body poses, head
备注: 18 pages, 6 figures, Accepted to ECCV 2026
点击查看摘要
Abstract:We address the problem of recovering the full-body mesh from only the head pose. This task has become essential for various applications based on head-mounted devices or smart glasses. The challenge of this task lies in estimating the pose information of unobserved body parts based solely on a single joint (i.e., head) trajectory. Several studies have begun to adopt head-conditioned generative models, however, such previous methods are costly and time-consuming due to the diffusion-based iterative process. As an alternative, we propose a simple yet novel method that leverages the latent space of the guidance network, which is designed as a variational autoencoder taking full-body poses as inputs. By enforcing latent distributions of this guidance network and our head-to-motion network to be similar, latent features sampled from the 'guided' distribution, i.e., distribution learned in our head-to-motion network, can be reliably decoded for natural representations of full-body poses even only with the head pose. One important advantage of the proposed method is that one-step sampling scheme achieves remarkably fast inference (more than 50 times faster) compared to diffusion-based approaches. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of ego-body mesh reconstruction.
96. 【2607.27749】Articulated Object Reconstruction from Rest-State Observation
链接:https://arxiv.org/abs/2607.27749
作者:Daeun Lee,Jaeah Lee,Woosung Kim,Haebeom Jung,Jaesik Park
类目:Computer Vision and Pattern Recognition (cs.CV); Robotics (cs.RO)
关键词:Building interactive digital, interactive digital twins, twins requires recovering, digital twins requires, Building interactive
备注: ECCV 2026
点击查看摘要
Abstract:Building interactive digital twins requires recovering both 3D geometry and the kinematic structures that govern how objects articulate. Yet existing methods for articulated object reconstruction require explicitly observable motion from multiple articulation states. We introduce a rest-state formulation that reconstructs articulated objects from a single closed configuration, an inherently ill-posed setting where geometry, semantics, and motion priors compensate for the absence of motion cues. Our framework adopts an explicit mesh as an intermediate representation for cross-model verification and fusion, reconciling noisy outputs from vision-language and segmentation models into spatially consistent part structures. To estimate joint parameters without observed motion, we use a video diffusion model to synthesize articulation hypotheses and validate them through geometric consistency. Our approach achieves accurate part decomposition and physically plausible articulation, performing competitively with motion-observing reconstruction-based, generation-based, and modular pretrained-model baselines.
97. 【2607.27729】PrintAnything: Learning an Intermediate Representation for 3D printing G-code Generation
链接:https://arxiv.org/abs/2607.27729
作者:Sangmin Hong,Daniel Sungho Jung,Heewon Kim,Kyoung Mu Lee
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Point clouds, Point, clouds, fundamental and widely, basic geometric representation
备注: European Conference on Computer Vision (ECCV) 2026
点击查看摘要
Abstract:Point clouds are one of the most fundamental and widely used 3D representations, serving as the most basic geometric representation of 3D shapes. Nevertheless, most existing 3D printing pipelines require a watertight mesh as input, preventing the direct use of point clouds for fabrication. A common workaround is to reconstruct meshes from point clouds; however, the resulting meshes often contain geometric artifacts, such as incorrect faces or topological inconsistencies, that are difficult to repair and may lead to printing failures. To overcome these limitations, we propose PrintAnything, a novel framework that learns to produce executable 3D printing G-code directly from 3D point clouds without requiring mesh reconstruction. To enable point clouds to serve as direct input for slice-wise toolpath generation, we introduce a slice-wise point projection strategy that transforms unstructured 3D point clouds into slice-aligned 2D representations consistent with layer-by-layer nature of fused deposition modeling in 3D printing. To eliminate mesh dependency and provide a unified representation that bridges point clouds and G-code, we propose Geometric plan (G-plan) map, a compact 2D representation composed of occupancy, region, and flow maps that encode the geometric and extrusion properties required for toolpath synthesis in 3D printing. As a result, our proposed method accurately generates printable G-code directly from point clouds, enabling a practical and fully mesh-free pipeline for 3D printing. The code is publicly available at \href{this https URL}{this https URL}.
98. 【2607.27700】Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
链接:https://arxiv.org/abs/2607.27700
作者:Jiasheng Li,Zhong Ji,Yan Zhang,Huihui Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Large Vision-Language Models, Large Vision-Language, prohibitive inference overhead, inference overhead due, suffer from prohibitive
备注:
点击查看摘要
Abstract:Large Vision-Language Models (VLMs) suffer from prohibitive inference overhead due to long sequences of visual tokens. However, existing visual token reduction methods mainly improve efficiency by pruning or compressing redundant tokens without examining whether the resulting representation remains semantically consistent with the original representation. Mapping the original N-token visual sequence to K tokens may discard, dilute, or misassign critical visual cues, triggering severe semantic drift that deviates the VLM's understanding. In this paper, we first introduce the principle of 'Calibrate Before Reason' to visual token reduction and propose CaRe, a training-free robust framework that calibrates compact visual representations before reasoning to preserve semantic fidelity in VLMs. CaRe consists of two mutually complementary modules: 1) Perturbation-Robust Calibration Anchoring, which identifies calibration anchors with stable model-side influence under multi-directional perturbations; 2) Confidence-Gated Token Calibration, which extracts reliable calibration signals from unselected tokens and injects them into anchors. Extensive evaluations across diverse VLM architectures and benchmarks verify that CaRe outperforms state-of-the-art token reduction baselines. While pruning 94.4% of visual tokens, our method retains 96.4% of the original full-token performance, delivering up to 2.30 times faster end-to-end inference speed relative to unpruned vanilla models.
99. 【2607.27699】RefineSVG: Visual Feedback-Driven Reinforcement Learning for Image-to-SVG Generation
链接:https://arxiv.org/abs/2607.27699
作者:Shaobo Liu,Feiqiao Mao,Shuaishuai Zhou,Yan Zhan,Weiqi Tan,Zhiqiong Lu,Zhengping Liang
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:enables multimodal large, multimodal large language, large language models, visual feedback framework, perform high-fidelity
备注: 17 pages, 5 main-paper figures. Accepted at the 34th ACM International Conference on Multimedia (ACM MM 2026). Includes the complete supplementary material
点击查看摘要
Abstract:We propose RefineSVG, a single-step closed-loop visual feedback framework that enables multimodal large language models (MLLMs) to perform high-fidelity image-to-SVG generation through self-correction. Existing MLLM-based approaches rely on single-pass open-loop inference, where the model receives visual input only once and must generate thousands of SVG code tokens without intermediate verification. This paradigm inevitably leads to geometric drift, error accumulation, and visual hallucination on complex images. RefineSVG overcomes this limitation by invoking an external rendering engine after an initial SVG generation pass to compare the rendered output against the target image. The comparison yields a multi-dimensional visual residual map (Diff-Map) that is fed back to the model as a ReAct-style correction signal, driving a targeted correction step. To support this render-observe-correct interaction, we further introduce an SVG-oriented semantic vocabulary that compresses token sequences by over 52%. A progressive training pipeline spanning supervised fine-tuning, rejection-sampling cold-start data construction, and end-to-end agentic reinforcement learning aligns the model with closed-loop visual correction. Extensive experiments show that RefineSVG consistently outperforms existing baselines in reconstruction fidelity, structural accuracy, and code this http URL is available at this https URL.
100. 【2607.27670】JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
链接:https://arxiv.org/abs/2607.27670
作者:Shawn Li,Wei Yang,Jike Zhong,Jiate Li,Jiawei Yang,You Qin,Ryan Rossi,Franck Dernoncourt,Roger Zimmermann,Yue Wang,Zhengzhong Tu,Vicente Ordonez,Mohit Bansal,Yue Zhao
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Jigsaw puzzle solving, solving requires jointly, create ambiguous ground, ambiguous ground truth, puzzle solving requires
备注:
点击查看摘要
Abstract:Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.
101. 【2607.27667】Witness Evidence Portfolios: Single-Prefill Risk Detection for Closed Multimodal Answers
链接:https://arxiv.org/abs/2607.27667
作者:Fexiang Liu,Shiye Wang,Qiang Qiu,Zheng Wang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:multimodal large language, large language models, requires deciding, deployment of multimodal, multimodal large
备注: 22 pages, 6 figures; includes supplementary material. Code: [this https URL](https://github.com/SouthWinter/WEP)
点击查看摘要
Abstract:Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or routed to a stronger system. Confidence scores capture candidate margins, but not where the estimated signed visual readouts associated with those margins come from or how they are distributed. We study inference-time risk detection for closed visual answers using the same white-box prefill path that produces the answer. Witness Evidence Portfolios (WEP) first estimates, layer by layer, which visual contributions support or contradict the predicted candidate. It summarizes these contributions through two interpretable route families: question-related evidence provenance and signed evidence concentration. Nested grouped validation chooses the more reliable family and a sparse top-k route portfolio, which is fused with candidate confidence. WEP needs no image perturbation, decoding change, backward pass, or external verifier. Across three MLLMs and four binary-answer benchmarks, WEP improves mean error AP by 0.134. All 12 model--dataset gains are positive, and image-cluster bootstrap intervals are strictly positive on 10 pairs. WEP targets white-box closed-answer systems and uses a labeled calibration slice.
102. 【2607.27660】Understanding Submodular Information Measure Based Objectives for Representation Learning: A Variance and Separation Perspective
链接:https://arxiv.org/abs/2607.27660
作者:Rishabh Iyer,Truong Pham,Anay Majee
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Submodular Information Measures, Submodular Information, Information Measures, representation learning, recently emerged
备注:
点击查看摘要
Abstract:Submodular Information Measures (SIMs) have recently emerged as a powerful framework for representation learning and multimodal learning. In particular, the SCORE framework~\cite{majee2024score} demonstrated that SIMs can serve as effective objectives for supervised contrastive learning. Despite their empirical success, however, the geometric and statistical properties induced by different submodular information measures remain poorly understood. In this work, we develop a unified theoretical framework connecting SIMs to classical concepts in representation learning and statistical pattern recognition. We show that Total Information (TI) objectives characterize intra-class structure: Graph Cut TI recovers within-class variance, LogDet TI recovers generalized variance and covariance volume, and Facility Location TI induces imbalance-aware separation that emphasizes rare and confusable classes. We further show that Mutual Information (MI) objectives capture complementary notions of inter-class structure: Graph Cut MI is closely related to centroid separation and Fisher-style discrimination, LogDet MI captures covariance-aware separation through Mahalanobis distance, and Facility Location MI measures nearest-mode representational overlap. We validate these theoretical characterizations using controlled synthetic experiments that independently vary variance, covariance, class imbalance, class separation, and multimodal overlap. Across all settings, the empirical behavior closely matches the proposed theory. Our results provide the first unified geometric and statistical understanding of submodular information measures and offer principled guidance for selecting and designing SIM-based objectives for representation learning.
Subjects:
Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
Cite as:
arXiv:2607.27660 [cs.LG]
(or
arXiv:2607.27660v1 [cs.LG] for this version)
https://doi.org/10.48550/arXiv.2607.27660
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
103. 【2607.27659】Learning Color Grading, No Photo Sharing: Federated Aesthetic Preference Learning for Personalized Image Enhancement
链接:https://arxiv.org/abs/2607.27659
作者:Chuanzhi Xu,Ziyuan Tao,Jean Julien KNell,Yanrong Chen,Haolan Guo,Xuanhua Yin,Adnan Mahmood,Weidong Cai
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
关键词:individual aesthetic taste, reflect individual aesthetic, preferences commonly depends, centralized collection, reflect individual
备注:
点击查看摘要
Abstract:Personalized image enhancement should reflect individual aesthetic taste, yet learning such preferences commonly depends on private photos and ratings that are unsuitable for centralized collection. The task must infer preference from sparse, heterogeneous feedback and translate it into natural-looking color transformations on resource-constrained user devices. We introduce FedPAIE, a federated personalized aesthetic image enhancement framework for user-adaptive color grading without centralizing raw photos or ratings. FedPAIE trains a lightweight dual-cue aesthetic scorer, calibrates it into a personalized scorer on a small local support set, and freezes it to guide regularized adaptation of a lightweight CLUT enhancer from unpaired local photographs. Fidelity constraints and an excess-gap penalty regularize scorer-guided adaptation to limit proxy-score over-optimization while preserving content and natural appearance. Training remains lightweight throughout the pipeline: scorer learning updates at most 0.787M parameters, enhancer adaptation updates 0.265M, and inference retains only a 0.293M-parameter personalized enhancer. Experiments on MIT-Adobe FiveK and Flickr-AES demonstrate effective open-world personalization and a favorable balance between user preference and image fidelity. FedPAIE thus connects decentralized preference learning with efficient personalized image transformation without requiring paired user retouches.
104. 【2607.27637】MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
链接:https://arxiv.org/abs/2607.27637
作者:Wenjie Zhu,Yabin Zhang,Wenjun Zeng,Lei Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multimodal Large Language, Large Language Models, Multimodal Large, Large Language, achieved strong performance
备注: project page: [this https URL](https://zhuwenjie98.github.io/MMOOC-project-page/)
点击查看摘要
Abstract:Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable MLLM should refuse truly out-of-context (OOC) questions with subject-level context shifts while still answering shifted in-context (Shifted IC) questions with non-subject context shifts. Existing benchmarks mainly target OOC or visually unanswerable questions, but overlook answerable Shifted IC cases and cover limited OOC shifts. To fill this gap, we present MMOOC, a large-scale benchmark for evaluating refusal and robust answering abilities of MLLMs. MMOOC contains over 41K image-question pairs, including answerable Shifted IC cases and unanswerable OOC cases, spanning three question formats, eight shift types and six visual scenarios, with data quality ensured through MLLM-based filtering and human verification. We evaluate model responses using Accuracy and Refusal Rate, and further introduce an LLM-as-a-Judge metric to assess the correctness of model reasoning. Experiments on diverse MLLMs show that current models still struggle to balance answer-ability and refusal under shifted contexts. We further analyze key failure patterns and show that post-training can improve robustness. MMOOC will be made publicly available.
105. 【2607.27634】4DHumanDiff: Direct Text-to-4DGS Generation for Consistent 360-Degree Dynamic Humans
链接:https://arxiv.org/abs/2607.27634
作者:Renlong Wu,Haoran Chen,Yuxiang Wei,Xiaowei Jin,Wangmeng Zuo,Hui Li
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generating high-quality, text prompts, dynamic human assets, dynamic humans, Generating
备注: 14 pages
点击查看摘要
Abstract:Generating high-quality 360-degree dynamic human assets from text prompts is challenging. Existing methods usually synthesize monocular or multi-view videos first and then fit a 4D representation, which is expensive and often causes incomplete geometry or view-inconsistent renderings. We present 4DHumanDiff, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting (4DGS) from text prompts. By modeling the structured 4D representation space end-to-end, 4DHumanDiff avoids video pre-generation and per-scene reconstruction, making it better suited for view-consistent and temporally coherent asset generation. The model uses a 3D U-Net backbone with temporal attention for motion-aware generation. We further construct a large-scale text-to-4DGS dataset with 60,000 high-quality pairs, and introduce 2D regularization and training-free 4D interpolation to improve rendering quality and motion smoothness. Experiments show that 4DHumanDiff generates consistent 360-degree dynamic humans within one minute, achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.
106. 【2607.27628】BlindPSNR: A No-Reference Fidelity Predictor for Low-Light Image Enhancement
链接:https://arxiv.org/abs/2607.27628
作者:Mingzhe Lyu,Jinqiang Cui,Hong Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:involve tunable parameters, methods involve tunable, applied across scenes, involve tunable, leading to performance
备注:
点击查看摘要
Abstract:Low-light image enhancement (LLIE) methods involve tunable parameters that are typically fixed, often leading to performance degradation when applied across scenes. Manually selecting the best configuration, however, can be time-consuming and not always practical. Peak signal-to-noise ratio (PSNR) is the natural fidelity criterion for automating parameter selection, yet it requires a ground-truth reference that is typically unavailable. To our knowledge, no learning-based method addresses no-reference PSNR prediction for low-light image enhancement; the natural surrogate, no-reference image quality assessment (NR-IQA), targets perceptual quality rather than signal fidelity, and all seven baselines we test achieve 0% top-1 selection accuracy on our benchmark. With paired training data, the ground-truth PSNR is analytically computable, providing exact supervision without a separate teacher network. Building on this, we propose BlindPSNR, a lightweight no-reference network that fuses the enhanced image with the degraded low-light input via windowed cross-attention and estimates PSNR through heteroscedastic regression. While a scalar-regression baseline achieves top-1 accuracy of 54.4%, BlindPSNR raises this to 89.5% with regret dropping from 1.62 dB to 0.026 dB, and generalizes to unseen datasets (SRCC = 0.61-0.67).
107. 【2607.27620】MedXplore: Towards Reliable and Unbiased Generalized Category Discovery in Medical Imaging
链接:https://arxiv.org/abs/2607.27620
作者:Jianwei He,Kailin Lyu,Junhao Dong,Long Xiao,Wenjie Hou,Jingze Lu,Di Wu,Lin Shu,Jie Hao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Generalized Category Discovery, medical image analysis, existing methods rely, shown strong potential, clinical practice
备注: accepted by ACM MM 26
点击查看摘要
Abstract:Deep learning has shown strong potential in medical image analysis, but most existing methods rely on large-scale annotations and a closed-world assumption that rarely holds in clinical practice. Although Generalized Category Discovery (GCD) has advanced rapidly on natural images, it remains underexplored in medical imaging. To address this issue, we propose MedXplore, a unified framework for reliable and unbiased medical GCD, optimizing from both perceptual and decision levels. Specifically, at the perceptual level, taking a frequency domain perspective, Frequency-SNR Adaptive Attention and Consistency (FAAC) performs learnable full-spectrum filtering and global-local energy contrast activation to not only highlight local abnormal signals relative to the global context, but also provide reliable semantic anchors for patch consistency learning. At the decision level, Adaptive Cosine-Angular Margin (ACAM) adjusts angular margins using semantic difficulty and feature confidence to balance intra-class compactness and inter-class separability. Together, the two modules improve lesion-sensitive representation learning and mitigate old-class bias. Experiments on multiple benchmarks show an average \textbf{8.5\%} gain in \textit{All} accuracy over the strongest competing methods. On Kvasir, MedXplore reduces false-old errors from 14.50\% to 0.80\%, demonstrating strong robustness under severe old-new ambiguity.
108. 【2607.27616】MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing
链接:https://arxiv.org/abs/2607.27616
作者:Jiajia Lin,Mingxuan Du,Tuowen Zhou,Benfeng Xu,Hongtao Xie
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:synthesize high-fidelity single-subject, high-fidelity single-subject images, synthesize high-fidelity, high-fidelity single-subject, personalized editing models
备注:
点击查看摘要
Abstract:Text-to-image and personalized editing models now synthesize high-fidelity single-subject images with ease. Yet placing multiple named people into shared contact actions such as embrace, carry, or grapple still exposes major failures: fused limbs, invented extremities, and interpenetrating bodies. Existing evaluations largely overlook these anatomical and geometric issues, and VLM-as-a-judge checklists often saturate on Interaction while the errors remain obvious to humans. We introduce MPIE-Bench, a 2,500-sample benchmark of video-mined editing triplets spanning 405 scenes, 14 interaction categories, and four contact densities (C0-C3). We also propose MPIE-Eval, whose two new axes score contact-time geometry from a frozen public multi-person mesh reconstruction. Anatomy asks whether every human-like mass is explained by a complete set of reconstructed bodies, and Interaction asks whether the penetration and surface distance between those bodies match the contact the instruction asked for. Across ten editors, mesh Anatomy tops out at 0.65 and mesh Interaction at 0.72 on two different models, so no single editor is strong on both, while VLM checklists rate the same images above 0.95. A five-rater study confirms that both axes track human judgement more closely than a zero-shot VLM judge, and the rankings hold under ablation of every weight and threshold.
109. 【2607.27592】MeshFM: 2D Features Are All You Need for 3D Shape Understanding
链接:https://arxiv.org/abs/2607.27592
作者:Jinfan Zhou,Richard Liu,Itai Lang,Rana Hanocka
类目:Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:efficient feedforward framework, extracting rich features, framework for extracting, extracting rich, efficient feedforward
备注: Project page: [this https URL](https://threedle.github.io/MeshFM/)
点击查看摘要
Abstract:We present MeshFM, an efficient feedforward framework for extracting rich features from 3D inputs. Our method distills 2D features from visual foundation models into 3D. We train a feedforward network to directly predict 3D features without requiring optimization during inference. The approach utilizes a two-stage training strategy. First, we optimize a feature field in 3D using only 2D feature supervision. Second, we train a network to regress this feature field. The entire procedure requires no 3D annotation, instead relying on the powerful information in 2D foundation models. We demonstrate that our learned features can be immediately applied to downstream tasks, including part segmentation, dense correspondence, and mesh deformation. Extensive experiments show that MeshFM, trained solely with 2D supervision, performs on par with methods trained explicitly with 3D supervision, even without task-specific fine-tuning. Moreover, our model is trained to be robust to extreme rotations of the input objects. Project page: this https URL
110. 【2607.27585】ZMIS-SAM: Segment Anything Model Enhanced with Wavelet Transform for Zooplankton Microscopy Image Instance Segmentation
链接:https://arxiv.org/abs/2607.27585
作者:Dekun Yuan,Zhongwei Li,Zheng Qiao,Jie Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:marine food chain, marine ecological balance, maintaining marine ecological, food chain, ecological balance
备注:
点击查看摘要
Abstract:As primary consumers in the marine food chain, zooplankton play a crucial role in maintaining marine ecological balance. However, the Segment Anything Model (SAM) exhibits limited performance in microscopic image instance segmentation due to its lack of zooplankton-specific domain knowledge. To address these challenges, we propose a novel instance segmentation model based on SAM and wavelet transform (ZMIS-SAM), effectively tackling issues such as inaccurate classification, discontinuous segmentation of slender appendages, and incomplete boundary segmentation. Our framework incorporates three core innovations: ZM-ViT enhances SAM's capability to model zooplankton morphology and image intensity distributions through two lightweight adapters, the Neighboring Feature Aggregation Module (NFAM) improves continuous segmentation of semi-transparent slender appendages by integrating general-purpose and domain-specific features, and the Wavelet-based Multi-scale Multi-directional Feature Enhancement (WM2FE) module effectively recovers high-frequency details to refine boundary segmentation completeness. Extensive experiments demonstrate that ZMIS-SAM achieves state-of-the-art instance segmentation performance on the zooplankton dataset and exhibits strong generalization capability across multiple public cross-domain datasets.
111. 【2607.27566】Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
链接:https://arxiv.org/abs/2607.27566
作者:Site Li,Jianyi Hao,Xiaofeng Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multi-frame medical VQA, static hard-negative mixing, Multi-frame medical, medical VQA, controller-style inference
备注: Presented at the CVPR 2026 Workshop on Multimodal Foundation Models for Biomedicine: Challenges and Opportunities
点击查看摘要
Abstract:Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negative mixing, and staged continuation all appear plausible from first principles. We test a simpler competing hypothesis on MedFrameQA: methods that remain tightly aligned with the benchmark's final answer objective should be the strongest \emph{robust} adaptation family once evaluation is controlled across fixed splits, matched budgets, repeated seeds, and calibration. We compare controller-based methods, scaffold evolution, static mixed supervision, continuation-heavy variants, and direct answer-only supervised fine-tuning (SFT). The strongest robust family is direct decoder-only answer SFT on MedGemma-1.5-4B. Empirically, this family yields substantial improvements in held-out report accuracy over frozen baselines while remaining remarkably stable across repeated seeds and matched controls, ensuring our claims reflect true family-level robustness rather than an isolated hyperparameter peak. Furthermore, post-hoc calibration effectively repairs confidence estimation without compromising accuracy, and the core approach transfers consistently to secondary backbones like Qwen2.5-VL-3B. The main result is therefore not that a complex auxiliary mechanism wins, but that objective-aligned direct answer SFT is the strongest robust adaptation family we found for MedFrameQA. By establishing this strong, minimalist baseline, we hope to redirect community focus toward fundamentally robust optimization rather than architectural complexity.
112. 【2607.27564】Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning
链接:https://arxiv.org/abs/2607.27564
作者:Site Li,Jianyi Hao,Xiaofeng Liu
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Multi-image medical VQA, medical VQA, prompt-length problem, fundamental challenge, Multi-image medical
备注: Presented at the CVPR 2026 Workshop on Multi-Modal Reasoning for Agentic Intelligence
点击查看摘要
Abstract:Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents must aggregate evidence across ordered images, remain robust to answer-order perturbations, and avoid overfitting to noisy search-time feedback. We study MedFrameQA through a controlled comparison of five inference-time agentic strategies, optimized using the same high-budget ShinkaEvolve configuration and evaluated on a reproducible internal frozen split (1,331 evolution, 665 holdout, 855 final test). Across five independent repeated runs, the strongest method emerges as the simplest robust aggregator: the \textbf{order-vote} policy achieves $57.89 \pm 0.65\%$ final-test accuracy, significantly outperforming the fixed baseline ($52.73 \pm 0.42\%$) and the more complex, albeit brittle, order-rerank variant ($55.79 \pm 0.43\%$). Paired bootstrap analysis confirms these significant gains. Extending the evolutionary search budget from 50 to 100 generations yields no generalization benefit: while holdout performance marginally increases, final-test accuracy drops from $57.89\%$ to $56.02\%$. Our findings suggest that for multi-image medical reasoning, defining the correct agentic decision rule is substantially more impactful than expanding the optimization search budget.
113. 【2607.27558】Drawing-Recode: Annotation Grounding for Parametric CAD Code Generation from Raster 2D CAD Drawings
链接:https://arxiv.org/abs/2607.27558
作者:Mingi Kim,Yongjun Kim,Hyungki Kim
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:Recovering Parametric CAD, Parametric CAD sequences, Parametric CAD, Structured Parametric CAD, Parametric CAD Code
备注:
点击查看摘要
Abstract:Recovering Parametric CAD sequences from raster-format 2D Computer-Aided Design (CAD) drawings accumulated prior to digital transformation is important for part reproduction and manufacturing process automation. However, existing studies either process only vector drawings or are limited to specific domains, and fail to explicitly connect dimensional annotations to geometric information, limiting their use of dimensional information for 3D Parametric CAD sequences recovery. We propose Drawing-Recode, a framework that generates Parametric CAD sequences as CAD code from raster 2D CAD drawings. Drawing-Recode extracts geometric features via an image encoder and recognizes annotations through a separate text recognition module, then explicitly grounds annotations to geometric information using cross-attention and our proposed Annotation Grounding Loss (AGL). The resulting features are fed into a Large Language Model (LLM) to generate CAD code in the Structured Parametric CAD Code (SPCC) format. Experiments show that Drawing-Recode outperforms existing baselines and remains robust on scanned drawings resembling industrial conditions. We expect Drawing-Recode contributes to digitizing raster 2D CAD drawings in industrial settings and to part reproduction and manufacturing automation.
114. 【2607.27549】Cross-Embodiment Transfer via Behavior-Aligned Representations
链接:https://arxiv.org/abs/2607.27549
作者:Ajay Sridhar,Jensen Gao,Jonathan Yang,Jean Mercat,Suneel Belkhale,Dorsa Sadigh
类目:Robotics (cs.RO); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
关键词:large-scale imitation learning, Recent progress, imitation learning, driven by leveraging, wide range
备注: Project page: [this https URL](https://ajaysridhar.com/barx/)
点击查看摘要
Abstract:Recent progress in large-scale imitation learning for robot manipulation has been driven by leveraging datasets across a wide range of robot embodiments. However, achieving significant cross-embodiment transfer is often still challenging. In this work, we study the role of using behavior-aligned representations (e.g., object bounding boxes, language motions, end-effector traces of robot motion) in vision-language-action (VLA) models to promote cross-embodiment transfer. We hypothesize that by possessing invariances across embodiments while being predictive of robot actions, these representations can help unify large-scale cross-embodiment data to enhance transfer. To assess our hypothesis, we develop a simulation-based benchmark designed to assess transfer with diverse cross-embodiment data to new embodiments. Using this benchmark, we compare different representations and ways of incorporating them. We identify that end-effector traces can be particularly beneficial for transfer, representations are generally more useful with larger prior datasets, and can be used to benefit from action-free data. We also demonstrate that they can enhance sim-to-real cross-embodiment transfer, improving task completion progress of real robot policies pre-trained on simulation data by 28%. We provide videos of our evaluations at our website: this https URL.
115. 【2607.27537】ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction
链接:https://arxiv.org/abs/2607.27537
作者:Dexuan Ding,Yuankai Qi,Luping Zhou,Jian Yang,Quan Z. Sheng,Ming-Hsuan Yang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:specific anatomical regions, Predicting future structural, anatomical regions, Predicting future, confined to specific
备注:
点击查看摘要
Abstract:Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, while most subject-specific brain structure remains stable over time. An effective model should therefore preserve global brain structural consistency while remaining sensitive to fine-grained disease progression. Existing latent-space-based methods improve computational efficiency, but suffer from information loss during their compression-reconstruction procedure. In contrast, direct voxel-space methods avoid latent reconstruction but commonly use a unified prediction pathway to model brain structure and progression-related changes. Subtle local changes may therefore be overshadowed by the dominant stable brain structure. To address these challenges, we propose ProgFormer, a hierarchical voxel-space Diffusion Transformer for longitudinal brain MRI prediction. ProgFormer uses a coarse pathway to perform the primary volumetric prediction from 3D patch tokens. This pathway models overall brain structure and longitudinal context. The fine pathway then uses the coarse representations as spatio-temporal grounding for voxel-level refinement within individual patches. The two pathways jointly estimate a velocity field directly in voxel space through conditional flow matching, enabling end-to-end prediction without a separately learned image autoencoder. The predicted future scan is then generated from Gaussian noise by integrating the estimated velocity field over a sequence of Euler steps. Extensive experimental results on three widely used benchmarks, ADNI, AIBL, and OASIS, under both pairwise and trajectory settings demonstrate favourable performance compared against several state-of-the-art methods.
116. 【2607.27465】IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks
链接:https://arxiv.org/abs/2607.27465
作者:Mengqi He,Jing Zhang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:transferable adversarial perturbations, evaluating transfer attacks, dense prediction models, adversarial perturbations, evaluating transfer
备注: 8 pages, 3 figures
点击查看摘要
Abstract:Semantic segmentation models are vulnerable to transferable adversarial perturbations, yet evaluating transfer attacks on dense prediction models can be computationally expensive. Existing ensemble attacks often rely on multiple surrogate models, increasing the computation cost, even harder for segmentation. This paper studies an efficient single-source alternative for transferable attacks on semantic segmentation. We formulate transferable attack composition as a chained computation over differentiable attack components, allowing the expensive source-model gradient computation to be shared. To reduce the update instability introduced by chained composition, we further use an integrated-gradient-style path-averaged direction as an empirical stabilization heuristic. Experiments on Pascal VOC and Cityscapes evaluate the resulting transferability efficiency trade-off across CNN- and transformer-based segmentation models. IGME achieves competitive transferability compared with single-source baselines and favorable runtime compared with model-ensemble attacks, while requiring access to only one source model.
117. 【2607.27380】VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
链接:https://arxiv.org/abs/2607.27380
作者:Haodong Li,Tianfei Ren,Xiaoxiao Ma,Chunmei Qing,Zhen Fang,Sipeng He,Ziyu Guo,Haoyu Wu,Juanxi Tian,Yihang Zou,Ruichuan An,Dongzhi Jiang,Boxue Yang,Ji Xie,Xu Huang,Wenhao Yan,Jialv Zou,Zhengrong Yue,Yaxin Luo,Xiaotong Li,Yuzhu Wang,Junyan Ye,Jinjing Zhao,Zehui Chen,Lin Chen,Renye Yan,Feng Zhao,Pheng-Ann Heng
类目:Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
关键词:remarkable visual quality, achieved remarkable visual, highly compressed text, compressed text prompt, models have achieved
备注: 15 pages, 3 figures, and 3 tables
点击查看摘要
Abstract:Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.
118. 【2607.27378】PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
链接:https://arxiv.org/abs/2607.27378
作者:Xiaohan Li,Xinyu Liu,Chang Liu,Sum Wing Au Yeung,Jun Liu,Yixuan Yuan,Hui Chen
类目:Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
关键词:multimodal large language, dental panoramic radiography, Accurate evaluation, large language models, panoramic radiography
备注:
点击查看摘要
Abstract:Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-grained, clinically reliable benchmarks that reflect expert interpretation. This work introduces PanDent, a large-scale, clinically grounded OPG benchmark built upon fine-grained, expert-validated tooth-level annotations. The dataset comprises 9,524 high-quality OPGs, each associated with comprehensive structured annotations produced by experienced dentists and further validated by an oral and maxillofacial radiologist, providing clinically reliable supervision for tooth-level diagnosis and reasoning. Clinically consistent radiology reports are constructed from expert-validated findings using clinician-defined reporting logic, establishing explicit correspondence between structured clinical evidence and free-text descriptions. This design enables evaluation of whether MLLMs generate reports that are not only linguistically coherent but also clinically consistent with expert-validated tooth-level findings. Experiments are conducted on diverse MLLMs, including state-of-the-art (SOTA) proprietary models, general-domain open-source models, and medical-specific models. Results show that current MLLMs can generate fluent reports, yet fail to produce clinically consistent descriptions, exhibiting substantial errors in fine-grained localization and tooth-level diagnosis. Fine-tuning on PanDent significantly improves structure-language consistency, substantially enhancing visual localization accuracy and diagnostic correctness, and bringing model outputs closer to expert dental interpretation. These results establish PanDent as a rigorous benchmark for evaluating tooth-level clinical reasoning in MLLMs and a valuable resource for clinically grounded dental AI.
119. 【2607.27372】Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
链接:https://arxiv.org/abs/2607.27372
作者:Alexi Gladstone,Heng Ji,Yilun Du
类目:Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
关键词:deep learning revolution, training beats decomposing, learning revolution, hand-designed stages, Generative modeling
备注:
点击查看摘要
Abstract:The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generative modeling, however, has remained the exception-despite generative models being remarkably capable, they are still not trained end-to-end. This is because, at its core, generative modeling is about handling distributions with many modes, and existing scalable approaches handle this the same way, by factoring the generation procedure, which prevents end-to-end generation. In this work, we introduce Explorative Modeling, a new paradigm that instead factors the training loop, exploring K candidate matches between model generations and data, and training on the best, so predictions commit to modes rather than blurring them. We find Explorative Models (XMs) useful in two settings. First, increasing exploration adds a third pretraining axis beyond parameters and data for existing generative models-where scaling exploration monotonically improves performance across both continuous and discrete domains (images, video, and language). Notably, gains from exploration increase with scale, climbing from 7% to 36% as data scales and from 13% to 23% as models grow, with efficiency gains more than doubling at 3x the compute. Concretely, exploration improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, parameter efficiency by 47%, lifts the strongest of image-generation recipes to a near-state-of-the-art 1.43 FID on ImageNet without guidance, enables scaling how end-to-end existing models are, and unlocks scaling generalization. Second, XMs enable end-to-end reconstructive generative modeling, matching diffusion on control tasks with 16-256x fewer inference steps. Together, these results establish XMs as both a new pretraining axis for existing generative models and a standalone end-to-end generative modeling paradigm.
120. 【2607.27357】Shared Semantic Codebook Distillation for Unpaired Cross-Modal Medical Classification
链接:https://arxiv.org/abs/2607.27357
作者:Dillan Imans,Phuoc-Nguyen Bui,Duc-Tai Le,Hyunseung Choo
类目:Computer Vision and Pattern Recognition (cs.CV); Image and Video Processing (eess.IV)
关键词:Cross-modal knowledge distillation, costly teacher modality, deployable student modality, transfer diagnostic knowledge, Cross-modal knowledge
备注: 16 pages, 2 figures, 4 tables
点击查看摘要
Abstract:Cross-modal knowledge distillation can transfer diagnostic knowledge from a strong but costly teacher modality to a cheaper and more deployable student modality. In medical image analysis, however, the two modalities are often unpaired: they are collected from different patient cohorts and occupy geometrically incompatible feature spaces. This makes instance-level distillation invalid and direct feature matching unreliable. To address these challenges, we propose Shared Semantic Codebook Distillation (SSCD), which compares teacher and student representations through a shared discrete codebook. Each image is represented as a distribution over a common, modality-agnostic vocabulary, and knowledge is transferred by aligning these distributions across modalities, both globally and class-conditionally, without requiring paired samples or directly comparable raw features. The codebook is evolved online by exponential moving average and kept diverse through entropy regularization and dead-code restart. At inference, all teacher-side and codebook modules are discarded, leaving only the student encoder and classifier. On two heterogeneous unpaired settings, OCT-to-fundus retinal disease classification and CT-to-chest-X-ray pneumonia classification, SSCD improves the student from 64.5 to 70.2 macro-F1 and from 73.8 to 76.3 macro-F1, respectively, outperforming all evaluated distillation baselines on both settings. Code and pretrained models are available at this https URL
121. 【2607.27348】Bunraku: Turning a Single Illustration into an Editable Live2D Character
链接:https://arxiv.org/abs/2607.27348
作者:Junhao Chen,Jingjia Mao,Dayong Li,Chenghai Li,Saining Zhang,Zhihao Li,Hao Zhao,Yufei Wang,Ruqi Huang
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:RGBA layers driven, per-layer mesh deformation, character-animation format, format for anime, driven by per-layer
备注: Project page: [this https URL](https://bunraku-live2d.github.io/)
点击查看摘要
Abstract:Live2D is the dominant 2D character-animation format for anime characters and virtual avatars, representing each character as a stack of RGBA layers driven by per-layer mesh deformation. Despite its wide use in virtual streaming, mobile games, and interactive characters, authoring a Live2D model still demands weeks of manual layer separation, occlusion completion, mesh placement, and keyframing, and no prior generative method produces such a structured asset end-to-end. We present the first system that, from a single illustration, generates all the structured information a Live2D runtime consumes: ordered RGBA layers, a deformation mesh per layer, and the parameter-driven keypose vertex offsets that make the character move. Stage 1 casts layered decomposition as a layered diffusion process under a Live2D-aware organ-level taxonomy, producing an ordered RGBA stack with hidden-region completion. Stage 2 builds a content-conforming triangle mesh for each layer from its alpha channel alone, then predicts the keypose displacement field of all layers jointly: every vertex of every layer is one token, self-attention spans layer boundaries, and each displacement is factorised into a bounded direction and a log-magnitude. Joint rather than independent prediction is what makes the result a coherent character instead of separately plausible parts, and is our largest gain; scaling the network 112x yields none. On 50 held-out characters, under true generation with no teacher forcing, Stage 2 attains a per-vertex direction cosine of 0.768 (median 0.828). Because a layer's mesh derives from its alpha channel, a clothing layer can be re-textured from a natural-language instruction while the mesh and predicted animation are reused byte-for-byte. We further contribute Live2D-Bench, the first standardized benchmark for the task, and an 8,884-model Live2D corpus with layer and animation supervision.
122. 【2607.27304】Position, Not Provenance: Separating Reasoning Mediation from Sycophancy in Medical Vision-Language Models
链接:https://arxiv.org/abs/2607.27304
作者:Supratik Bhowal,Subhrajyoti Basu,Aritra Gir Mahanta,Anik Pal Chowdhury
类目:Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV)
关键词:predictions remains unclear, causally influences predictions, influences predictions remains, answering clinical questions, reasoning causally influences
备注:
点击查看摘要
Abstract:Medical vision-language models (VLMs) generate chain-of-thought (CoT) reasoning before answering clinical questions, but whether this reasoning causally influences predictions remains unclear. We present CoT-Mediate, a behavioral framework that perturbs a single clinically meaningful attribute within a model's own generated reasoning and measures whether the resulting prediction follows the edited reasoning. Our framework combines a dual-arm protocol comparing re-prompted evidence with prefix-forced continuation, together with a provenance-controlled intervention that varies only the attributed source of identical reasoning to disentangle reasoning mediation from sycophancy. We evaluate LLaVA-Med and MedGemma on 1,000 VQA-RAD samples each. Prefix-forced continuation consistently yields higher mediation faithfulness than re-prompting, while the provenance analysis reveals distinct model-specific deference behaviors. Across both models, removing visual evidence increases reliance on injected reasoning, whereas laterality is the least faithfully tracked clinical attribute. These results show that the mechanism used to inject reasoning substantially affects measured faithfulness and that contextual position, rather than stated provenance, is the primary determinant of whether medical VLMs use their generated reasoning.
123. 【2607.27292】VETO: Towards Protecting Images From Frontier AI Editing
链接:https://arxiv.org/abs/2607.27292
作者:Jonas Grebe,Hossein Shakibania,Tobias Braun,Marcus Rohrbach,Anna Rohrbach
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:accessible image-editing models, brought high-fidelity editing, rise of powerful, accessible image-editing, broad reach
备注:
点击查看摘要
Abstract:The rise of powerful, accessible image-editing models such as FLUX.2 has brought high-fidelity editing within broad reach. Their capabilities now extend beyond localized modifications to extracting and recontextualizing objects and identities in entirely new scenes. By allowing prompt and generation tokens to attend directly to reference-image tokens, modern models blur the boundary between conventional editing and text-to-image synthesis. This expanded generative freedom also broadens the space of potential misuse, as harmful transformations are no longer confined to a predictable set of localized edits. Existing anti-edit defenses are designed to disrupt the semantic bottleneck of the reference-image encoding in legacy diffusion pipelines. However, newer editors distill reference information through joint-attention blocks, thereby often circumventing these protections. We therefore introduce VETO, a subtle anti-edit cloak that disrupts this inner mechanism through which modern models read the source image. Additionally, as existing editing benchmarks leave comprehensive recontextualizations largely untested, we introduce VetoBench, which evaluates defenses not only on conventional localized edits but also on broader contextual shifts. Across two contemporary editing models and three benchmarks, VETO consistently outperforms existing defenses while providing a stronger protection-fidelity trade-off.
124. 【2607.27278】OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
链接:https://arxiv.org/abs/2607.27278
作者:Kaiyu Li,Zepeng Xin,Zixuan Jiang,Jing Fu,Lanxuan Xue,Lingyu Zhang,Xiangyong Cao
类目:Computer Vision and Pattern Recognition (cs.CV)
关键词:Open-vocabulary Earth observation, localize geospatial concepts, Earth observation, fixed label set, Open-vocabulary Earth
备注:
点击查看摘要
Abstract:Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms. To fill this gap, we introduce OVEarth-Bench, which extends existing evaluation in two directions: category breadth, through broad hierarchical category coverage with positive and negative expressions, and query diversity, through vocabulary, referring, and reasoning queries. The benchmark supports mask and box localization under a unified zero-shot protocol. We evaluate a broad set of general and EO-specific methods. The evaluation reveals that: (1) the performance of current methods remains limited, while broader category coverage yields more stable model rankings; (2) MLLM-based methods achieve the strongest overall performance; and (3) EO-specific methods generally underperform general models and rarely match the strongest methods. These findings provide guidance for future open-vocabulary EO method design and highlight the importance of developing more realistic, diverse, high-quality, and large-scale benchmarks for reliable evaluation. Our data and evaluation package are released at this https URL.
125. 【2607.27266】heatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography
链接:https://arxiv.org/abs/2607.27266
作者:Diego Belzarena(UDELAR, CB),Seginus Mowlavi(CB),Paula Casariego Castiñeira(ROMA TRE),Alejandra Ulla Lorenzo(USC),Gregory Randall(UDELAR),Jean-Michel Morel(LU - Hong Kong)
类目:Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV)
关键词:printed historical books, methodology that quantifies, quantifies the similarity, historical books, statistical methodology
备注:
点击查看摘要
Abstract:We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books. This provides a tool that accelerates philological analysis. Using character prototypes derived from clustering and aligning automatically extracted character images, the method defines a typeface distance between any two books. To produce actionable outputs, we develop an a contrario statistical framework to interpret the significance of the computed typeface distances. We apply the method to the philological study of 17 th -century Spanish printed theatre chapbooks in a quantity that exceeds the capabilities of systematic visual inspection by human experts. Our method enables the automatic comparison of Roman and Italic types extracted from different books. After validation by human experts, our method has led to new printer attributions being discovered, and former printer attributions being revised. This success strongly suggests that our method has the potential to enable digital bibliography on a larger scale than was previously possible.
126. 【2607.27235】RadHarmony: Radiological Data Handling in the Era of Agentic AI
链接:https://arxiv.org/abs/2607.27235
作者:Frank Li,Bardia Khosravi,Mohammadreza Chavoshi,Theo Dapamede,YoungSeok Jeon,Janice Newsome,Hari Trivedi,Judy Gichoya
类目:Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV)
关键词:Training deep learning, radiological images requires, Training deep, images requires integrating, deep learning models
备注:
点击查看摘要
Abstract:Training deep learning models on radiological images requires integrating heterogeneous datasets across different sources, file formats, directory layouts, label schemas, and annotation types. We present RadHarmony, an open-source Python library that provides a unified API for loading, harmonizing, and augmenting radiological datasets, with a primary focus on chest radiographs and early support for computed tomography (CT) and magnetic resonance imaging (MRI). RadHarmony standardizes metadata from 24 public datasets into a single tabular format, wraps MONAI's map-style datasets for deep-learning-ready sample delivery with optional on-disk caching, and supports classification labels, segmentation masks, bounding boxes, and radiology report text through a single interface, with an interactive visualization tool for dataset exploration and verification. To lower the barrier for integrating new datasets, RadHarmony introduces an AI-agent skill that guides the full integration workflow from raw data inspection through code generation and testing. We demonstrate the library's utility by pretraining RadHarmony-ViT, a reference vision transformer baseline that combines three heterogeneous chest radiograph datasets with no dataset-specific code. The code and pretrained model weights are available at this https URL.
127. 【2607.28144】ReGenVC: End-to-End Real-Time Generative Video Coding at Ultra-Low Bitrate
链接:https://arxiv.org/abs/2607.28144
作者:Zheyuan Zhang,Johnson Wu
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:compresses talking-head video, generative video codec, real time, compresses talking-head, talking-head video
备注: 12 pages, 5 figures, 5 tables
点击查看摘要
Abstract:We present ReGenVC, an end-to-end generative video codec that compresses talking-head video to an ultra-low bitrate and decodes it in real time. The encoder reduces a source clip to a compact bitstream -- a neurally compressed first frame, per-frame pose keypoints, and metadata -- totaling about 26 kB for a 77-frame sequence. The decoder is a four-step distilled diffusion transformer that reconstructs the video conditioned on the transmitted pose and reference frame. Compared with x264/x265, ReGenVC reduces the bitrate to roughly one tenth of that required by traditional codecs (about 26 kB vs. 250--280 kB for essentially artifact-free reconstruction); at a matched ultra-low bitrate, conventional codecs collapse into blocking artifacts while ReGenVC stays sharp by exploiting a strong generative prior. The central obstacle to deploying such a codec is decoder latency: multi-step sampling with transformer and VAE components is too slow for interactive use. We make the decoder real-time through four-step distillation and three model-preserving system techniques: (i) eight-GPU unified sequence parallelism (Ulysses Ring), (ii) a spatially-split VAE, and (iii) a three-stage overlapped pipeline; an analytical timing model characterizes the real-time feasibility region. On an 8-GPU node, the system sustains 24 fps output (972 ms per 25-frame window, within the 1000 ms budget), enabling a live browser stream without observed frame underruns. A hybrid CPU-GPU deployment further runs the encoder on the CPU at 24 fps and offloads the decoder-side one-shot conditioning encoders to the CPU, reducing the per-GPU memory peak from 21.1 GB to about 7.7 GB. To our knowledge, ReGenVC is the first end-to-end generative video codec to combine ultra-low-bitrate encoding with real-time decoding on an 8-GPU system.
128. 【2607.28049】SOG: A Format For Temporally And Spatially Ordered Gaussians
链接:https://arxiv.org/abs/2607.28049
作者:Shady Gmira,Evangelos Alexiou,Emmanouil Potetsianakis,Emmanuel Thomas
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR)
关键词:Spatially Ordered Gaussians, Spatially Ordered, Ordered Gaussians, Gaussian Splatting, propose Temporally
备注: 2026 IEEE International Conference on Image Processing (ICIP). IEEE, 2026
点击查看摘要
Abstract:We propose Temporally and Spatially Ordered Gaussians (TSOG), a format for efficient representation of 4D Gaussian Splatting (4DGS) content. TSOG extends the Spatially Ordered Gaussians (SOG) framework to the temporal domain by introducing a timeline attribute and temporal parameterization of geometry and appearance attributes. Similar to SOG, TSOG is a lossy format that assigns each Gaussian a unique index and encodes attribute values as index-aligned image data. TSOG is model-agnostic, extensible, and compatible with both discrete and continuous 4DGS representations. Evaluation using a PLYs sequence and FreeTimeGS as baselines, serving as simplistic and state-of-the-art 4DGS representations respectively, shows file size reductions exceeding 90%, with PSNR differences ranging between -0.42 and +0.85 dB. These results demonstrate substantial file size savings with minimal quality degradation, enabling efficient representation, storage, and delivery of dynamic scenes for next-generation 4D content.
129. 【2607.27825】Endo-NeRF++: Uncertainty-Aware Neural Rendering with Multi-Resolution Hash Encoding for Dynamic Surgical Scene Reconstruction
链接:https://arxiv.org/abs/2607.27825
作者:Gousia Habib,Laura Ruotsalainen
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:minimally invasive surgery, Reconstructing dynamic surgical, robot-assisted minimally invasive, dynamic surgical scenes, Reconstructing dynamic
备注:
点击查看摘要
Abstract:Reconstructing dynamic surgical scenes is crucial for robot-assisted minimally invasive surgery; however, it continues to be difficult because of tissue deformation, occlusions, specular reflections, and restricted viewpoints. In this study, we introduce Endo-NeRF++, a neural rendering framework that accounts for uncertainty in the reconstruction of dynamic surgical scenes. Expanding on EndoNeRF, the suggested approach incorporates multi-resolution hash-grid encoding, temporal feature merging, and uncertainty-informed adaptive sampling to enhance reconstruction accuracy and temporal coherence in deformable endoscopic this http URL multi-resolution hash-grid representation within the framework effectively captures both coarse and fine anatomical details, while temporal feature blending ensures stable reconstruction during tissue deformation and surgical tool occlusions. Additionally, uncertainty-driven adaptive sampling assigns more samples to uncertain areas to enhance rendering quality and geometric coherence. Experiments on robotic surgical video sequences demonstrate that the proposed uncertainty-guided adaptive sampling improves PSNR by up to 1.22\,dB (4.3\%), increases SSIM by up to 5.3\%, and reduces LPIPS by up to 55.1\% compared with the EndoNeRF baseline.
130. 【2607.27741】hree-Photon Bayesian Imaging of Ortho-Positronium
链接:https://arxiv.org/abs/2607.27741
作者:L. Raczynski,W. Krzemien,A. Coussat,M. Bala,B.C. Hiesmayr,K. Klimaszewski,M. Obara,R. Y. Shopa
类目:Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV); Computational Physics (physics.comp-ph)
关键词:PET provides functional, functional images relying, relying on two-photon, positron-electron annihilation, PET
备注: 17 pages, 4 figures
点击查看摘要
Abstract:PET provides functional images relying on two-photon coincidences from positron-electron annihilation. In human tissue, about 40\% of annihilations are preceded by Ps formation, of which o-Ps component partially decays into three photons, with the remainder annihilating via pick-off or spin-exchange into two photons. This three-photon channel carries additional information about the surrounding micro-environment, including the three-to-two-photon yield ratio as a potential diagnostic marker. We propose the TRIO algorithm, a novel three-photon event-by-event image reconstruction algorithm formulated as a Bayesian maximum a posteriori inference problem. TRIO unifies time-based trilateration, energy-based reconstruction and, for the first time, a physics-informed prior derived from the QED description of Ps decay within a single probabilistic framework. In contrast to positronium lifetime imaging, which requires a prompt photon and is therefore restricted to specific radionuclides, TRIO relies solely on the three photons and is fully compatible with standard radionuclides such as 18F. Monte Carlo simulation modelled after the Siemens Biograph Quadra scanner demonstrates a mean position error of 1.62~cm, improving by approximately a factor of two over the time-based trilateration (3.05 cm) and by about an order of magnitude over energy-based reconstruction alone (18 cm). More importantly, the proposed Bayesian approach is compatible with existing TOF-PET scanners that can register three-photon annihilation coincidences.
Comments:
17 pages, 4 figures
Subjects:
Medical Physics (physics.med-ph); Computer Vision and Pattern Recognition (cs.CV); Computational Physics (physics.comp-ph)
Cite as:
arXiv:2607.27741 [physics.med-ph]
(or
arXiv:2607.27741v1 [physics.med-ph] for this version)
https://doi.org/10.48550/arXiv.2607.27741
Focus to learn more
arXiv-issued DOI via DataCite (pending registration)</p>
131. 【2607.27286】oward Multi-Modal Deep Learning for Pulmonary Disease Classification: A Texture-Based Machine Learning Pilot Study on Public Chest X-Ray Data
链接:https://arxiv.org/abs/2607.27286
作者:Yogisri Pujitha Chinthoti
类目:Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
关键词:widely studied application, Automated classification, Image Data Collection, classification of pulmonary, pulmonary disease
备注: 6 pages, 2 figures
点击查看摘要
Abstract:Automated classification of pulmonary disease from chest radiographs is a widely studied application of machine learning in medical imaging. This paper presents a pilot study evaluating classical texture- and gradient-based feature representations for distinguishing COVID-19 from other forms of pneumonia using the publicly available COVID-19 Image Data Collection (668 posteroanterior/anteroposterior radiographs from 408 patients). Using histogram of oriented gradients (HOG) and gray-level co-occurrence matrix (GLCM) texture descriptors with classical classifiers (logistic regression, random forest, and support vector machine), evaluated under patient-level 5-fold stratified cross-validation to prevent data leakage, we obtain a best mean accuracy of 75.4% and AUC of 0.755, modestly exceeding the 71.6% majority-class baseline. We report these results transparently, including their limitations, and use them to motivate and scope a proposed multi-modal deep learning architecture -- combining convolutional and transformer-based encoders across imaging modalities -- as a direction for future work requiring access to larger, multi-institutional, ethically sourced datasets.

