Physical AI Brief
Daily cross-source signals for the Physical AI supply chain — silicon photonics, CPO, VLA models, humanoid hardware, embodied AI. Three streams, one page, zero filler.
1222 items today · 1166 arxiv · 2 SEC 8-K · 54 humanoid · 0 CN photonics
01 ARXIV · PHYSICAL AI PAPERS
1166 items- arxiv:2609.35765 · cs.CLRetrieving Biblical Intertextual References in Karen Blixen's Seven Gothic TalesAndrás Kovács, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen
Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Ta
benchmark - arxiv:2609.35764 · cs.CVReliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D PoseZhilin Guo, Boqiao Zhang, Oszkár Urbán, Josef Bengtson +7
Sparse inertial pose estimation promises camera-free motion capture from consumer devices, but consumer sensors are unreliable: firmware-fused orientations are biased, mounting varies between sessions, and streams drift or drop out. On a new 35-take single-subject benchmark pairing an earbud head in
benchmark - arxiv:2609.35761 · cs.RODexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human DemonstrationsRui Zhou, Yibo Yuan, Junkai Zhao, Fangyuan Zhao +3
Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease t
vlamanipulationdexterouspi0gr00t - arxiv:2609.35760 · cs.LGTokenCast: Forecasting Token Consumption During LLM Agent ExecutionChaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo +6
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call.
agentllm agent - arxiv:2609.35759 · cs.CLScaling Long-Form Story Generation via Narrative State TrackingZhennan Wan, Jianfei Chen
LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-generation methods typically focus on stories of up to about ten thousand words, leavin
agentagenticbenchmark - arxiv:2609.35752 · cs.LGNeural Harmonic Measure OperatorJinjin He, Sinan Wang, Yuchen Sun, Bo Zhu
We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on
benchmark - arxiv:2609.35750 · cs.LGKV-streams for Efficient Compaction in Agentic Reinforcement LearningEmiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer +14
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling
memoryagenticpost-training - arxiv:2609.35748 · cs.LGImproving Test-Time Scaling with Adaptive Looped TransformersYichen You, Tianyu Fu, Aosong Feng, Xingtai Lv +3
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs gro
post-trainingbenchmark - arxiv:2609.35744 · cs.AIFinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research AgentsHoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho +16
Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric
agentbenchmark - arxiv:2609.35743 · cs.CVInfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric VideoKerui Ren, Kaiwen Song, Weiguang Zhao, Yuxi Wang +7
World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe c
memorybenchmark - arxiv:2609.35741 · cs.AIShockingly Simple Self-retrospection Improves Agentic Models Without RLJonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer +6
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this que
agentagentic - arxiv:2609.35738 · cs.LGHarness Learning Enables Generalizable Test-Time AdaptationAlvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar +5
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at h
agenttool use - arxiv:2609.35734 · cs.CVGeoVerse: World-Consistent Novel View Synthesis in Geometric Latent SpaceKerui Ren, Tao Lu, Linning Xu, Changjian Jiang +4
Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unse
memory - arxiv:2609.35732 · cs.AIFailure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language ModelsJunru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding +3
Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA
agentbenchmark - arxiv:2609.35728 · cs.CVFlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent PlanningZiyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang +4
We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini
humanoidagent - arxiv:2609.35718 · cs.CVHard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer VisionHanoona Rasheed, Mohammed Irfan Kurpath, Bin Ren, Hisham Cholakkal +2
Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We
benchmark - arxiv:2609.35717 · cs.RORoboCompiler: Graph-Native Compilation of Closed-Chain Robots for Consistent Modeling, Control, and SimulationMehdi Heydari Shahna, Joongheon Kim, Jouni Mattila
Robots with kinematic loops, coupled actuators, and changing contacts require consistent models of configuration, motion, force, and dynamics. Yet these interfaces are often reconstructed separately for control and simulation, making closure and actuation consistency difficult to maintain. This pape
franka - arxiv:2609.35715 · cs.ROX-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment ResetsPrithwish Dan, Chenyang Ma, Wei Zhan
Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom
manipulationdexteroussim-to-realgrippergrasp - arxiv:2609.35709 · cs.ROHumanoid Loco-Manipulation With Discrete VLA ModelWenxin Shao, Siqi Chai, Kun Li, Kerou Zhang +3
Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, trainin
vision-language-actionvlavla modelmanipulationhumanoidteleoperation - arxiv:2609.35707 · cs.LGScAn-Bench: Evaluating Scaling Analysis MethodologyArtin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog +3
Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the metho
benchmark - arxiv:2609.35706 · cs.AIReinforcing Agentic Creativity in Scientific Ideation with Night SciencePriyanka Kargupta, Silviu Cucerzan, Shweti Mahajan, Allen Herring +3
Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to l
agentic - arxiv:2609.35704 · cs.CVDynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test TimeZiqi Ma, Hongqiao Chen, Georgia Gkioxari
Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle
world model - arxiv:2609.35703 · cs.LGA Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic ExpansionFred Xu, Thomas Markovich, Florence Regol, Yizhou Sun
Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filter
benchmark - arxiv:2609.35701 · cs.LGMeqMuon: Matrix-Equilibrating Muon for LLM PretrainingChang-Wei Shi, Xu Wang, Wu-Jun Li
The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining
memory - arxiv:2609.35700 · cs.ROLQR-ArUco Fusion: Robust Hierarchical Control for Navigation and Asymmetric Manipulation in Two-Wheeled RobotsAnupam Chatterjee, Arpita Kumari
We propose a hierarchical control framework to address severe dynamic instabilities and navigational drift that arise when a two-wheeled inverted pendulum (TWIP) robot attempts asymmetric object manipulation. While two-wheeled platforms are highly manoeuvrable, their constant balancing adjustments m
manipulation - arxiv:2609.35698 · cs.LGProvable Benefits of Regularization: Fast Rates for Adversarial Imitation LearningHanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang
We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components
agent - arxiv:2609.35694 · cs.AIReasoning with Continuous Latent DiffusionXiang Cheng
Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We
iterative refinement - arxiv:2609.35692 · cs.AIReport: Progressive Disclosure of Agent SkillsGuilin Zhang, Kai Zhao, Priyanka Mudgal, Waleed Ammar +3
Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost. P
agent - arxiv:2609.35690 · cs.ROAgent Priors-guided Policy LearningPuming Jiang, Tianrun Hu, Haozhe Du, Yibo Li +3
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost b
grippergraspagent - arxiv:2609.35686 · cs.LGRethinking Circuit Evaluation: Do Circuits Explain Model Errors?Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang +3
Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to r
benchmark - arxiv:2609.35685 · cs.CLQuanReview: Offline, Auditable Reconciliation of Human and LLM Span AnnotationsMatteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda, Naomi Couriel +3
Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanR
benchmark - arxiv:2609.35673 · cs.CVFlowTool: Controlling Tool Parameter in Image Retouching via Flow MatchingThanh-Long V. Le, Steven Walton, Seunghyun Yoon, Branislav Kveton +3
Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a
memorypost-training - arxiv:2609.35671 · cs.AIPhoneCLI: From App Interfaces to Callable Commands for Mobile AgentsYangqin Jiang, Lingrui Xu, Chao Huang
Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeate
agent - arxiv:2609.35664 · cs.AIMS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal ResolutionPrasoon Dev, Anirudh Sankar, Vasudeva Varma
Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously
memorylong-contextbenchmark - arxiv:2609.35658 · cs.CVMany Eyes, One World: Feed-Forward 3D Reconstruction from Mixed CamerasQiaoge Li, Yifan Zhan, Haijun Yang, Haiyang Liu +2
Real-world capture is heterogeneous: perspective, fisheye, and $360^\circ$ panoramic images can coexist within a single reconstruction task, yet most feed-forward 3D reconstruction models assume perspective imagery and a uniform input representation. Recent models handling several camera types are e
benchmark - arxiv:2609.35657 · cs.AICMDO: A Cognitive Memory-Driven Optimization Algorithm for Adaptive Population-Based SearchMohammed Yusuf Mujawar, Shahram Rahimi, Noorbakhsh Amiri Golilarz
Population-based optimization methods often use previous search information through successful solutions, parameter adaptation, or operator performance, but they rarely retain the context in which a search behavior succeeded or failed. We introduce Cognitive Memory-Driven Optimization (CMDO), a deri
memorybenchmark - arxiv:2609.35652 · cs.ROMM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and ImaginingQiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao +5
Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base act
manipulationliberoworld model - arxiv:2609.35645 · cs.LGCoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise SettingsShama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin +1
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how
agentbenchmarkevaluation framework - arxiv:2609.35641 · cs.CVVerifiable Visual Rewards Transfer from Synthetic Scenes to Natural PromptsShuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer
Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR)
post-trainingbenchmark - arxiv:2609.35639 · cs.AIGPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics SimulationYuchen Sun, Jinjin He, Sinan Wang, Bo Zhu
Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks co
benchmark - arxiv:2609.35629 · cs.LGSANTA++: Sampling Attention through Representative KeysKyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee +3
Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection wit
memoryretrieval-augmented - arxiv:2609.35627 · cs.CLCan LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical DiagnosisKehua Feng, Yunsheng Lu, Yitong Qiao, Tiantian He +4
A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidenti
benchmark - arxiv:2609.35622 · cs.LGElicitation and Decision Geometry in Single-Index BanditsSakshi Arya, Cheng Soon Ong
We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (
benchmark - arxiv:2609.35621 · cs.LGCartridges++: KV Cache Compression without Off-Context DerailmentSonia Laguna, Joao Monteiro, Marco Cuturi, Pierre Ablin +1
Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all,
memorylong-context - arxiv:2609.35620 · cs.LGAttention Graphons: A Graph Limit Perspective on Graph TransformersCaio F. Deberaldini Netto, Moshe Eliasof, Luana Ruiz
Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-depe
benchmark - arxiv:2609.35619 · cs.ROCollisionSplatting: Collision-Aware Motion Planning in 3DGS Scenes with Image-Conditioned Objectives and Adjustable ConservatismR. Khorrambakht, Joaquim Ortiz-Haro, Stephan Weiss, Ludovic Righetti
Incorporating dense visual information into motion planning remains challenging, as geometric planners rely on abstracted scene representations that discard visual richness, while learned visual models often lack geometric interpretability and computational efficiency. This paper introduces Collisio
manipulation - arxiv:2609.35616 · cs.CVEvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations UnfoldJunjie Chen, Fei Wang, Kun Li, Yiqi Nie +4
Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAva
benchmark - arxiv:2609.35615 · cs.LGBehavioral Foundation Models for Quality DiversityNazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud
Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting t
manipulationbenchmark - arxiv:2609.35614 · cs.LGEvE: An Alternate Optimizer to AdamShashank Raj, Kalyanmoy Deb
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutiona
benchmark - arxiv:2609.35611 · cs.LGOn-Policy Self-Distillation for Multi-Turn Image EditingLiangbing Zhao, Le Zhuo, Mohamed Elhoseiny
Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure
benchmark - arxiv:2609.35606 · cs.AITCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer ScienceChutong Yang, Xiyuan Zhang, Yu Huang, Boran Han +6
Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether m
agentagenticbenchmark - arxiv:2609.35603 · cs.LGControl-Geometry Straightening for Sampling-Based Latent PlanningZiang Fu, Ning Ning
Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly represe
world model - arxiv:2609.35596 · cs.AISEABench: Benchmarking Endogenous Misalignment In Self-Evolving AgentsSaswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang +2
Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates
memoryagentllm agentagenticself-evolvingbenchmark - arxiv:2609.35586 · cs.AIIMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory ComputingYung-Chin Chen, Chia-Yu Chen, Naveen Verma
Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision
memory - arxiv:2609.35583 · cs.CVWhat Paired Evaluations Reveal under Visual PerturbationsYongda Wei, Chen Zhang, Yifei Wang, Xinyu Wang +3
Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggr
benchmark - arxiv:2609.35581 · cs.LGQC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing TasksPranav Gupta
We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find
benchmark - arxiv:2609.35578 · cs.AIFactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language ModelsBowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat
memory - arxiv:2609.35576 · cs.LGShare-Borne AI Virus: Memory-Hopping Attacks Across LLM AgentsSidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav +3
Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent as
persistent memoryllm agent - arxiv:2609.35575 · cs.ROF4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-ImprovementZhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren +12
The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered fai
vision-language-actionmanipulationsim-to-realagentself-improvement - arxiv:2609.35570 · cs.ROEdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation ModelRithvik Jonna, Man Namgung, Aakash Gurram, Tinoosh Mohsenin
Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preservin
memory - arxiv:2609.35569 · cs.LGBeyond Energy: When Sustainability Dimensions Reshape LLM Serving DecisionsTianyao Shi, Xipeng Shen, Yi Ding
Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different optimization decisions. We present PRISM,
embodied - arxiv:2609.35568 · cs.LGFrom Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel SynthesisLongxiao Fan, Tao Zhang, Han Yan, Jiajun Li +4
High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models
memoryexternal memoryagentself-improvingpost-training - arxiv:2609.35564 · cs.AIAlmieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech RecognitionOmid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir +36
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unsee
benchmark - arxiv:2609.35561 · cs.AIRSI-Master: Structuring Experiments to Guide Autonomous Model ImprovementYaxin Du, Xiyuan Yang, Zhifan Zhou, Yujie Ge +9
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit
agentself-improvementpost-trainingbenchmark - arxiv:2609.35560 · cs.CVWorldPlay2: Extending Real-Time Interactive World Models in Control and HorizonHaiyu Zhang, Wenqiang Sun, Tengfei Wang, Junta Wu +4
Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time respons
world modelmemory - arxiv:2609.35559 · cs.AIFrom Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor MiningKangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian +5
Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and perfo
agentllm agent - arxiv:2609.35557 · cs.AIThe Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding AgentShobhan Roy
The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program
agent - arxiv:2609.35553 · cs.LGSimplex Diffusion ModelsJustin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang +3
Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information co
iterative refinement - arxiv:2609.35551 · cs.AIBaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent ConsultationPeilin Feng, Zhengyang Huang, Soujanya Poria
In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advis
memoryagentmulti-agentagent systembenchmark - arxiv:2609.35549 · cs.AIRareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease DiagnosisBo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong +5
Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates,
post-trainingbenchmark - arxiv:2609.35545 · cs.LGGraph World Models for Constrained Epidemic Policy PlanningYiqi Su, Rashed Shelim, Lingyi Wang, Walid Saad +1
Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We presen
world modelaction-conditioned - arxiv:2609.35540 · cs.AIContinuous Context ManagementWilliam Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan
Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the act
memoryagent - arxiv:2609.35539 · cs.CVLearning to Reason with Persistent Object States for Video Instance SegmentationYongxue Xu, Boxue Yang, Ziqian Liu, Shaoqiu Zhang +2
Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSRe
persistent statebenchmark - arxiv:2609.35537 · cs.LGOptimal Networks for Agentic Information AggregationMohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz +1
We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a subset of the features and its parents'
agentagentic - arxiv:2609.35536 · cs.CVLook Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake DetectionChia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin, Tai-Ming Huang +7
Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evid
manipulation - arxiv:2609.35532 · cs.AIARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement LearningKun Feng, Yuchen Fang, Yiyang Tan, Shuqi Gu +5
As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training prioriti
agentagenticagent benchmarkbenchmark - arxiv:2609.35530 · cs.CVAutoRef: Harness Optimization for Agentic Multi-Reference Image GenerationYuta Oshima, Ku Onoda, Yusuke Iwasawa, Masahiro Suzuki +2
Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally past
agentagenticbenchmarkevaluator - arxiv:2609.35525 · cs.LGDeep Epistemic Value Functions for Optimistic ExplorationLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central chall
agent - arxiv:2609.35517 · cs.LGReward-Aligned Reweighting for On-Policy DistillationHaofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou +1
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's
benchmark - arxiv:2609.35515 · cs.LGMechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding +1
Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery lar
benchmark - arxiv:2609.35507 · cs.CVReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question AnsweringZhen Yao, Likai Wang, Yuming Yang, Zhihao Zheng +8
Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images t
benchmark - arxiv:2609.35505 · cs.LGAn RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM ReasoningShangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao +3
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that
benchmark - arxiv:2609.35504 · cs.CVSolveEdit: Benchmarking Visual Problem Solving in Generative ModelsWenjie Shu, Yexin Liu, Harold Haodong Chen, Xuerui Qiu +8
Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing th
benchmark - arxiv:2609.35502 · cs.LGStructured Latent Modeling for Supervised Multimodal Information DecompositionWanting Huang, Sanvesh Srivastava, Weiran Wang
Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way
benchmark - arxiv:2609.35501 · cs.LGSRHarness: A Harness for Agentic Symbolic RegressionZihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding +1
Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runt
agentic - arxiv:2609.35497 · cs.CVSprout: Building Dynamic Memory While Reasoning for Agentic Video UnderstandingWei Chen, Xuanyu Zheng, Yancheng Long, Haoyang Xu +5
Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by seve
memoryagentagenticbenchmark - arxiv:2609.35493 · cs.ROTerrain-Aware Autonomous Planetary Exploration for Exteroceptive-Proprioceptive Mapping with Quadruped ScoutsAlberto Sanchez-Delgado, João Carlos Virgolino Soares, Victor Barasuol, Claudio Semini
Autonomous planetary exploration requires robots to navigate unknown, uneven terrain while assessing risk, traversability, and energetic cost. Quadruped scouts are well suited for this task because they can traverse irregular surfaces and gather mobility-relevant information during locomotion. This
quadruped - arxiv:2609.35491 · cs.CVFrom Scores to Samples: Elastic Forcing for Autoregressive Video GenerationChi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An +4
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-traini
post-training - arxiv:2609.35486 · cs.CVWho Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position ReasoningYingjin Song, Denis Paperno, Albert Gatt
High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this proce
benchmark - arxiv:2609.35482 · cs.ROForVis: An In-Field Dataset and Benchmark for VIO Using Under-Canopy UAV Flights in ForestsArman Kiani, Masoud Ataei, Elvis Gyaase, Jeffrey Eiyike +3
Visual-inertial Simultaneous Localization and Mapping (VI-SLAM) for UAVs remains difficult to evaluate in real forest environments, where motion, illumination changes, repetitive vegetation, and vibration can all affect estimation. We present ForVis, an in-field dataset and benchmark for evaluating
benchmark - arxiv:2609.35479 · cs.RORobot Tool Design from Scratch via Behavior-Aware Hierarchical OptimizationYinghan Chen, Xiyao Tian, Yizan Dai, Yuyang Li +1
The ability to design a tool for a task marks a level of intelligence beyond merely understanding, selecting, or using one. Existing methods for robotic tool design typically optimize a tool's continuous shape and action within a structure that is prescribed or generated beforehand, so the structure
tool-use - arxiv:2609.35477 · physics.opticsPlasmon-Enhanced Second-Harmonic Generation in Atomically Thin Crystalline Silver NanostructuresSaad Abdullah, Philipp K. Jenke, Andrew P. Weber, Álvaro Rodríguez Echarri +7
The intrinsically weak nonlinear optical response of existing materials, further constrained by symmetry-forbidden second-order processes in centrosymmetric media, severely limits efficient frequency conversion in deeply subwavelength, ultrathin volumes. Addressing this challenge is crucial for the
quantum photonic - arxiv:2609.35476 · cs.ROCoBrush: A Hierarchical Planning Framework for Human-Robot Co-PaintingDantong Qin, Yike Guo, Qinlin Liu, Alessandro Bozzon +1
Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coheren
embodied - arxiv:2609.35473 · cs.LGHandwritten Text Recognition Lives in the High-Pixel Variance SubspaceCarlos Garrido-Munoz, Jorge Calvo-Zaragoza
In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-
benchmarkevaluation protocol - arxiv:2609.35469 · cs.RORethinking Causal Action Tokenization with Conditional Annealing in Flow MatchingChenyu Zhang, Yuhang Cao, Daru Du, Yingxi Lu +7
Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action token
vision-language-actionvlamanipulationbenchmark - arxiv:2609.35465 · cs.LGTetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per ParameterPier-Jean Malandrino
Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on
memory - arxiv:2609.35463 · cs.AIA.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual HumansAlessandro Emmanuel Pecora, Stefano Calzolari, Francesco Strada, Andrea Bottino
Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions wi
world modeltool calling - arxiv:2609.35462 · cs.LGCLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn ConversationsYusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu +6
Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchm
benchmark - arxiv:2609.35461 · cs.CLAraDynFact: Dynamic Evaluation of Factual Knowledge in ArabicIgnacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi
As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains lar
benchmarkevaluation framework - arxiv:2609.35456 · cs.AIAutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain--Computer InterfacesMuyun Jiang, Yi Ding, Wei Zhang, Jinbo Chen +8
EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection
agentagentic - arxiv:2609.35450 · cs.ROUni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-ManipulationZihao Wang, Shutong Liu, Siqi Zheng, Liu Cao +4
Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions,
vision-language-actionvlamanipulationhumanoidtactile - arxiv:2609.35445 · cs.LGNeuronSifter: Intervention Planning in CNS MicroenvironmentsHaowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie
Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar expo
action-conditionedbenchmark - arxiv:2609.35440 · cs.LGSOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local LearningBojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li
Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict t
memory - arxiv:2609.35433 · cs.LGReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy LearningYihang Chen, Yuanhao Ban, Cho-Jui Hsieh
Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppre
benchmark - arxiv:2609.35432 · cs.ROSelf-Evolving Coding Agents: From Digital Programs to Physical-World IntelligenceHongcheng Gao, Jingjing Zhou, Zelin Zheng, Shijia Ge +9
Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirem
vision-language-actionagentself-evolving - arxiv:2609.35431 · cs.ROMemory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory AggregationChengyang Li, Yujie Wan, Shuai Wang, Kejiang Ye +5
This paper studies low-altitude question answering (LAQA), in which distributed unmanned aerial vehicle (UAV) memories are aggregated at a ground server to answer questions about observations over a long horizon. Unlike conventional resource allocation based on sensing, communication, control, or co
memoryagent memorymulti-agentagent systembenchmark - arxiv:2609.35427 · cs.LGLLMs are General Asynchronous AgentsGeorge Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy +3
Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform a
embodiedmemoryautonomous agentembodied agenttool calling - arxiv:2609.35426 · cs.LGFrontier Learning: Training LLM Reasoners at the Edge of CapabilityRobin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi
Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as
post-training - arxiv:2609.35416 · cs.CVWhen Should the Count Change? Learning State Maintenance for Causal Video CountingPengyiang Liu, Dongyue Lyu, Junbo Niu, Zhongyue Shi +2
Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain vi
benchmark - arxiv:2609.35412 · cs.AISelf-Adapting Group of Experts for Multi-Agent ReasoningMohammad Atif Quamar, Nurbek Tastan, Karthik Nandakumar, Junpei Komiyama
Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often
agentmulti-agentagent systembenchmark - arxiv:2609.35409 · cs.AIAwarenessBench: Assessing Cognitive Capabilities of Language ModelsXiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang +8
As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social aw
benchmark - arxiv:2609.35408 · cs.AI"Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM DeliverablesYage Zhang, Yukun Jiang, Yang Zhang
Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the ed
agentbenchmark - arxiv:2609.35407 · cs.CVBiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete DiffusionWanjiang Weng, Yongliang Wu, Xiaofeng Tan, Xingyu Zhu +2
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poo
self-correction - arxiv:2609.35402 · cs.LGPersistent Partners Raise Prices Among Learning AgentsPaul-Peter Arslan, Yubin Kim, Xiao Xiao
When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand duopoly of Calvano et al., each agent's pr
agent - arxiv:2609.35400 · cs.AIStructural Alignment for Reliable Industrial AI: Bridging Physical Reality, Data, Models, and Human IntentLizhi Xiao, Sihong Wu, Victoria Xiao, Yiqiao Song +4
Artificial intelligence is increasingly deployed in critical industrial domains, including healthcare, energy grids, subsurface exploration, where failures can have severe consequences for human safety, system stability, and economic outcomes. Yet AI is still evaluated primarily through benchmark ac
benchmark - arxiv:2609.35394 · cs.CVRethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong BaselineXiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu +2
Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos. Recent video token compression methods increasingly introduce sophisticated strategies for to
benchmark - arxiv:2609.35390 · cs.LGInductive Feedback for Mixed-Policy DistillationAmir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang
Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, whi
agenticpost-trainingbenchmark - arxiv:2609.35381 · cs.AIMCP Error Messages Written for Developers Hurt the Most Capable Agents MostXiaonan Xu, Wenjing Wu
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 erro
agentleaderboard - arxiv:2609.35378 · cs.AIMultilinguality in Hybrid Attention LLMsLucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz +1
In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations o
agentic - arxiv:2609.35375 · cs.ROFrom Pixel to Poses: Object-centric Tool Manipulation Learning from Human DemonstrationsBangjun Wang, Longyan Wu, Yukun Wei, Shenghe Shao +7
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although cu
manipulationworld modeltool use - arxiv:2609.35367 · cs.CLFrom Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder FeaturesDewen Liu, Zixuan Li, Jonathan Pan, Zhao Wu +3
Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while
agentagentic - arxiv:2609.35366 · cs.AIPlanarian: Managing Agent State with StatepointsJinnan Guo, Hao Mark Chen, Kapil Vaswani, Andrew Paverd +1
LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or re
agentllm agent - arxiv:2609.35362 · cs.LGd-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language ModelsRuitao Liu, Qinghao Hu, Song Han
Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. R
benchmark - arxiv:2609.35357 · cs.AIDo Coding Agents Reuse Existing Code or Reinvent the Wheel?Dongsheng Ma, Sizhe Wang, Xinyi Huang, Zhengren Wang +4
Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents
benchmark - arxiv:2609.35349 · cs.LGQuasi Linear Kernel Attention with Infinite CapacityNicolaj Rux, Johannes Hertrich, Sebastian Neumayer
The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi lin
benchmark - arxiv:2609.35347 · cs.LGBeyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy DistillationXin Li, Hao Jiang, Xin Gao, Annan Wang +5
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists t
benchmark - arxiv:2609.35342 · cs.AIJev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability CalibrationRiccardo Porcedda
The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public
benchmark - arxiv:2609.35341 · cs.CVGenerative Uncertainty as a Self-supervised Signal for Semantic Similarity LearningEnrico Pallotta, Sina Raoufi, Lars Doorenbos, Gianni Franchi +1
Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-tempo
v-jepa - arxiv:2609.35336 · cs.CVTMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem SolvingShengqin Wang, Jie Jin, Yu Cheng, Yihang Chen +3
Despite the promise of Large Language Models (LLMs) in computational chemistry, rigorous combinatorial chemistry problems remain difficult because they require quantitatively constrained molecular modification, candidate validation, and systematic revision after failed attempts. Existing tool-augmen
multi-agentagent frameworktool use - arxiv:2609.35335 · cs.LGLarge Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and EvaluationPetros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann
Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level informatio
benchmark - arxiv:2609.35333 · cs.LGScalable In-Context Reinforcement Learning with Recurrent Algorithm DistillationYuanqing Ma, Zhenrui Zheng, Chenjun Xiao
Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scal
memory - arxiv:2609.35328 · cs.AIHyper Algorithm Design Agent: Evolving Learnable Optimizer from ZeroZipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang +1
Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm's bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While M
agentagent frameworkself-improvement - arxiv:2609.35324 · physics.opticsContactless Continuous-Variable Quantum Optical State Conditioning With Classical MmWave PhaseNiloy Ghosh, Sarang Pendharker
This paper establishes the Heisenberg's picture framework for seamless mmWave-to-photonic data transduction. Based on this framework, contactless modulation of continuous-variable (CV) quantum optical states with digitally modulated classical mmWave beams is shown for the first time. Our analysis re
quantum photonic - arxiv:2609.35319 · cs.LGTeacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous AgentsTong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao +7
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for cor
autonomous agent - arxiv:2609.35318 · cs.RODexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool LibraryYouhui Wang, Yunzhu Li, Li Fei-Fei, Jiajun Wu +1
Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving arti
manipulationdexteroussim-to-realagenticself-evolving - arxiv:2609.35316 · cs.AIReliability Engineering for AI Systems: Challenges, Methods, and DirectionsRong Pan, Yili Hong, Min Xie
AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permission
agentictool useself-evolvingbenchmark - arxiv:2609.35312 · cs.CLMemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMsZineddine Tighidet, Andrea Mogini, Jiali Mei, Patrick Gallinari +1
Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \textit{memorization bias}, where familiar
memorybenchmark - arxiv:2609.35311 · cs.RORoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model RolloutsJin Hyun Kim, Min Young Kim, Soohwan Song, Daekyum Kim
Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predic
world modelaction-conditioned - arxiv:2609.35308 · cs.CLEpistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient InvestigationFahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah
Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct fail
benchmark - arxiv:2609.35303 · cs.CVPIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM AgentsJiazhou Zhou, Hu Zhou, Yucheng Chen, Jinyuan Qu +2
Reinforcement learning with verifiable rewards (RLVR) via Group-Relative Policy Optimization (GRPO) is widely used for multi-turn VLM agent training, yet it suffers from zero-gradient silence on uniform failures and coarse episode-level credit assignment. While On-Policy Distillation (OPD) and On-Po
embodiedagentagent benchmarkbenchmark - arxiv:2609.35302 · cs.AINarrowing the Horizon: Quantifying Topic Saliency Shifts in Generative MonocultureOriane Peter, Elena Simperl, Kate Devlin
As Large Language Models (LLMs) become central to how we access and share information, they play an increasingly powerful role in shaping global knowledge. However, as these models evolve, their outputs risk converging into a \textit{generative monoculture}, where the diversity of perspectives they
post-training - arxiv:2609.35298 · cs.AITraining-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph FrameworkSurajit Das
Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Inform
knowledge graph - arxiv:2609.35291 · cs.LGNarrow Multimodal Fine-Tuning Can Induce Emergent MisalignmentShunchang Liu, Lukas Fluri, Xin Chen, Francesco Croce
Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in
agenticpost-training - arxiv:2609.35290 · cs.AIEvoIn: Bridging Evolution and Internalization for Agent Fine-TuningShihan Dou, Shaofan Liu, Zhonghang Lu, Jiahang Lin +7
Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, w
agentbenchmark - arxiv:2609.35288 · cs.LG$λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised LearningBerker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero +2
Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the pr
v-jepavjepabenchmark - arxiv:2609.35286 · cs.AIThe Argument and the Letterhead: Source-Position Coherence in AI EvaluationMichele Loi
An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake an
evaluator - arxiv:2609.35285 · cs.AITextual User Taste: Natural-Language User Context for Foundation-Model Recommender System at ScaleGhazal Fazelnia, Paul Gigioli, Eliza Klyce, Sharon Zheng +15
Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embedding vectors remain highly effective for retrieval and ranking, but they are opaque to users and not n
evaluation framework - arxiv:2609.35279 · cs.CLMeasuring Collapse and Correction in Homogeneous-Panel LLM DebateXin Li, Mengbing Liu, Chau Yuen
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing
multi-agent - arxiv:2609.35269 · cs.LGeval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion ModelsMansi, Nikhil Raghavan, Zixia Huang, Kai Sheng Ong +3
The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Py
benchmarkleaderboard - arxiv:2609.35267 · cs.ROGuardPIBT: Counterfactually Gated Neural Guidance for Ultra-Large-Scale 3D Multi-Agent Path FindingYuan Zhou, Zhenyu Hou, Guangtong Xu, Xiaoqiang Ji +3
Large-scale 3D multi-agent path finding becomes increasingly difficult under dense traffic. Priority Inheritance with Backtracking (PIBT) scales well, but its one-step goal-directed ordering may become insufficient under dense interactions and large-scale congestion. We present GuardPIBT, which augm
multi-agent - arxiv:2609.35262 · cs.CLRubric-Aware On-Policy Self-Distillation for LLM PersonalizationYilun Qiu, Xiaoyan Zhao, Chengbing Wang, Cilin Yan +5
LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only
graspbenchmark - arxiv:2609.35255 · cs.AITowards Reliable AI Data Scientists: Data Agents with Workflow HarnessesHuachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai +6
Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science li
benchmark - arxiv:2609.35249 · cs.ROSpatial Grafting: Grounding 3D Features for Flow-Matching Robot PoliciesDingsheng Liu, Yangzheng Wu, Mahboubeh Asadi, Zhiyuan Li +4
Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape wi
vision-language-actionmanipulationrobotwinbehavior-1kbenchmark - arxiv:2609.35246 · eess.SYHard-Constrained Probabilistic Factor Graph Neural Network for Distribution System State Estimation under Non-Gaussian UncertaintyM. Furqan Azam, Marta Vanin, Chris Hermans, Geert Deconinck
Robust and accurate state estimation is fundamental for the reliable operation and monitoring of active distribution networks. Conventional numerical estimators, such as weighted least squares, are computationally slower and often suffer from convergence issues in the presence of sparse measurements
benchmark - arxiv:2609.35236 · cs.LGLong-Horizon Scaling: How Model Capabilities Shape the Returns to ComputationHaoyu Zheng, Zhengyu Chen, Huaisheng Zhu, Ruishan Fang +4
Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab an
benchmark - arxiv:2609.35233 · cs.AIEP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM AgentsFengzhou Sun, Yuan Zhang, Xintong Yu, Jinyao Yan
Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users' social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent m
memorymemory architectureagent memoryagentllm agentbenchmark - arxiv:2609.35231 · cs.ROZero-Shot Reactive Obstacle Avoidance for Generative Robot PoliciesWeihang Guo, Lydia E. Kavraki
We propose NUDGE (Nudge Update via Differentiable GEometry), a training-free obstacle-avoidance procedure that can be incorporated in any robot policy based on diffusion or flow matching, including diffusion policies and vision-language-action models. Our work injects gradients from a signed distanc
vision-language-actionrobot policy - arxiv:2609.35228 · cs.CVToken-Disentangled Latent Test-Time Scaling for Vision-Language ReasoningHao-Xuan Ma, Yihao Liu, Yutao Sun, Yanting Miao +7
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sen
benchmark - arxiv:2609.35226 · cs.CVGenerative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and BenchmarkMarco Parola, Mario G. C. A. Cimino, Sabrina Senatore
Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-
benchmark - arxiv:2609.35225 · cs.CVSignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at ScaleZhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai +1
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propo
benchmark - arxiv:2609.35215 · cs.AIASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement LearningYang Li, Jinhan Yang, hai liu, Di Wan +7
Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tre
retrieval-augmentedagentic - arxiv:2609.35212 · cs.LGAdversarial Consistency-Guided Representation Learning for Multi-view ClusteringYuchen Lin, Kunpeng Xu, Ying Fang, Lifei Chen
Multi-view clustering aims to capture cross-view consistency while exploiting view-specific information. However, shared representations learned to capture cross-view consistency may still retain view-identifying information, potentially compromising the consistency of cross-view clustering structur
benchmark - arxiv:2609.35210 · cs.CLUnderstanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse CrosscodersZichao Yu, Qianshuo Ye, Xu Wang, Difan Zou
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoder
post-training - arxiv:2609.35208 · physics.opticsGenerating Vector-Vortex $γ$ Photons by Nonlinear Compton ScatteringYong-Zheng Ren, Mamutjan Ababekri, Jun-Lin Zhou, Feng Wan +7
Vector-vortex photons, characterized by a nonseparable coupling between polarization and orbital angular momentum (OAM), offer opportunities for optical manipulation, quantum communication, nuclear photonics, etc. However, their generation in the $γ$-ray regime remains challenging. Here, we put forw
manipulation - arxiv:2609.35200 · cs.ROReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot ManipulationPankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj +3
Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned poli
manipulationliberomemory - arxiv:2609.35201 · cs.AIFrom Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference DataHusrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan +5
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounde
post-trainingbenchmark - arxiv:2609.35196 · cs.AIGAC-PINN: Geometry-Adaptive and Constraint-Enhanced Physics-Informed Neural NetworksYanxin Zhang, Yong Zhang, Houbiao Li
For systems with steep gradients, sharp interfaces, or severe spatio-temporal coupling, Physics-informed neural networks (PINNs) suffer from spectral bias, geometric inflexibility, and boundary constraint conflicts, which undermine accuracy and convergence. To overcome these issues, we propose a geo
benchmark - arxiv:2609.35195 · cs.CVCarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis SegmentationMd Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans, Walia Farzana +3
Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically di
benchmarkevaluation protocol - arxiv:2609.35193 · cs.LGConRAG: Lightweight inference of multi-hop relationsKilian Bänziger, Sonia Laguna, Markus Kreft, Robert Jakob +3
Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and other knowledge-intensive analyses. We forma
ragknowledge graph - arxiv:2609.35189 · cs.LGG$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRAJia Song, Wenhow Li, Lichen Bai, Bada Ye +1
Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-tra
post-trainingevaluator - arxiv:2609.35188 · cs.AIBeneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM InferenceSuwesh Prasad Sah
Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled
memorybenchmark - arxiv:2609.35182 · cs.AIResearch-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific AgentsDi Wang, Yu Liu, Bing Cui, Chaoqun Ji +4
We describe AfS (Agent for Science), a platform built for long-horizon scientific work, where a project runs for tens of hours across dozens of agent runs with a human present only occasionally. Most agents for science are general coding agents with a skills folder attached, and they inherit that li
memoryagentbenchmark - arxiv:2609.35177 · cs.LGSubgroup Rank-1 Lattice for Practical High-dimensional Black-box Integral ApproximationYueming Lyu
Estimating integrals of black-box, high-dimensional functions, from expectations and kernel mean embeddings to the softmax kernel in self-attention, is a basic subroutine in machine learning. Rank-1 lattice rules suit this setting: they query the integrand only at a fixed point set and need no gradi
memory - arxiv:2609.35160 · cs.AIFONDANT: Strong and Best-Effort Planning via AntichainsBenjamin Aminof, Tuan Khai Nguyen, Sasha Rubin
A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available
agentbenchmark - arxiv:2609.35158 · cs.AIPEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement LearningJiaan Zhu, Wei Gao, Youhui Bai, Zewen Jin +5
Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use als
agentic - arxiv:2609.35152 · cs.LGCoordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning ApproachGuodong Ma, Baofeng Sun, Wenyu Yang, Zhihong Yao
Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transfera
multi-agent - arxiv:2609.35149 · cs.AIFrom Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and ScaleYaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan +1
Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retention or target-contra
agent - arxiv:2609.35143 · cs.CVTimeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final CutGunin Gupta, Nirmit Arora, Pavan Kalyan Tankala
AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tas
agentai agentbenchmark - arxiv:2609.35142 · cs.MAStrategically Robust Game-Theoretic Multi-Agent Trajectory OptimizationVictor L. Qin, Nicolas Lanzetti, Saverio Bolognani, Hamsa Balakrishnan
Aviation authorities worldwide expect Advanced Air Mobility (AAM) traffic management to be decentralized among service providers, requiring AAM flights to autonomously plan trajectories by predicting other flights' control inputs rather than relying on centralized coordination. Game-theoretic approa
agentmulti-agent - arxiv:2609.35139 · cs.LGCacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache FusionGenglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu +2
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the asse
retrieval-augmentedrag - arxiv:2609.35138 · cs.LGFlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time ScalesShidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun +3
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a
world modelbenchmark - arxiv:2609.35134 · cs.CVVideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene ReconstructionConghan Yue, Yuanjie Chen, Yue Han, Ya Gao +3
Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored.
benchmark - arxiv:2609.35117 · cs.AITool Mediation Alters Refusal Mechanisms in Large Language ModelsAbel Rodríguez, Giuseppe Garofalo, Lieven Desmet, Vera Rimmer
Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms
llm agent - arxiv:2609.35115 · cs.AIDuplexCadence: Exact State and Execution from a Speech Model's Declared TimelinesHaixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives.
memory - arxiv:2609.35113 · cs.LGSymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic RegressionZiwen Zhang, Xiju Wu, Yuheng Jing, Runxiang Wang +8
Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack syst
benchmark - arxiv:2609.35110 · cs.LGSol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and EdgeYitong Li, Jincheng Yu, Junsong Chen, Haopeng Li +5
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substant
memoryself-improvement - arxiv:2609.35107 · cs.AIDoAtlas-2: A Foundation for Self-Evolving Causal Biomedical DiscoveryYulong Li, Rong Xia, Yuxuan Zhang, Jianxu Chen +11
We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, fro
self-evolving - arxiv:2609.35106 · cs.LGDRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation PredictionMustapha Bounoua, Giulio Franzese, Pietro Michiardi
Predicting cellular responses to perturbations is a central problem in cellular biology, with broad applications in systems biology and drug discovery. This task is challenging because cellular responses can be complex and cell-state dependent, intrinsic cell-to-cell variability can be confounded wi
benchmark - arxiv:2609.35099 · cs.LGE3J: An Efficient and Open-Source Backend for Euclidean Equivariant Operations on GPU and TPUOlivier Peltre, Armand Picard, Adrien Pichard, Miguel Bragança +5
We present e3j, a fast Euclid-equivariance backend for geometric deep learning applications with JAX bindings for GPU and TPU. Leveraging both optimized CUDA and Pallas kernels and algorithmic improvements, the library achieves state-of-the-art throughput and runtime on both forward and backward pat
memorybenchmark - arxiv:2609.35097 · cs.LGSpikeLite: Lightweight Spiking Neural Networks for Time-Series ForecastingBang Hu, Changze Lv, Mingjie Li, Xiaoqing Zheng +2
Spiking neural networks (SNNs) offer an energy-efficient paradigm for time-series forecasting through spike-driven computation. However, recent SNN forecasters often pursue higher accuracy through increasingly complex attention mechanisms, or specialized neuronal dynamics, weakening the lightweight
benchmark - arxiv:2609.35096 · cs.LGDF-CBM: Region-Aware Concept Bottleneck Models for Deepfake DetectionGeorgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou, Panos K. Papadopoulos +3
Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only pa
manipulation - arxiv:2609.35094 · cs.ROQuadHand: A Compact Quadrotor Aerial Manipulator with MRC-SDF-Based Whole-Body Motion PlanningRui Jin, Ruiyang Liu, Xinhang Xu, Haotian Jin +3
Uncrewed aerial manipulators (UAMs) integrate robotic arms with aerial platforms for three-dimensional physical interaction. However, enlarging the workspace increases arm-induced disturbances, while existing geometric representations face a trade-off between geometric fidelity and computational eff
manipulationmanipulatorgripper - arxiv:2609.35090 · cs.CVAdvancing Video-Text Pretraining with Multi-View CaptionsFida M. Thoker, Renaud Vandeghen, Karen Sanchez, Marc Van Droogenbroeck +1
Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, whi
benchmark - arxiv:2609.35089 · cs.AICan Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping ResearchZehao Lu, Xingguo Xiong, Wopke van der Werf, Thijs L. van der Plas +1
Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-inte
multi-agentagent system - arxiv:2609.35088 · cs.AIWhen Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM AgentsGeonwoo Kim, Brent ByungHoon Kang
Tool-enabled agents form calls from model-visible interfaces, while hosts later select their implementation. Standard dispatch omits the descriptor-handler relation. An unchanged and schema-valid call can therefore acquire a different security effect during rollout, reconnect, or delayed approval. W
llm agent - arxiv:2609.35086 · cs.LGRetrieval-Augmented Diffusion Modeling for Stochastic Discount Factor PortfoliosKelvin J. L. Koa, Xinyang Li, Ke-Wei Huang
In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial markets exhibit non-stationary dynamics wit
retrieval-augmented - arxiv:2609.35084 · cs.LGGraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM AgentsHaodong Zhu, Yangyang Ren, Changbai Li, Sheng Xu +3
Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not expl
llm agentagentic - arxiv:2609.35083 · cs.LGContinuous Variational SynthesisAlan N. Amin, Mattia G. Gollub, Andrei Slabodkin, Elizabeth B. Wood +1
Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models is challenging: constraints on chemical
post-training - arxiv:2609.35082 · cs.LGCross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement LearningYangyang Ren, Haodong Zhu, Linlin Yang, Sheng Xu +2
Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should inc
llm agentagenticbenchmark - arxiv:2609.35078 · cs.CVRefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action DrivingZhe Sun, Ziyi Luo, Yehao Lu, Lei Zhou +1
Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive
vision-language-actionpost-training - arxiv:2609.35072 · cs.LGNot All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement LearningXuesong Jia, Ziao Yang, Zhanhe Huang, Hongfu Liu
We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated tra
post-training - arxiv:2609.35065 · cs.LGTempoKV: Timely Staging of LLM KV Caches for Memory-Semantic FlashJay H. Park, Hyungjun Kim, Dong Kim
Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas imme
memory - arxiv:2609.35060 · cs.LGBeyond Gradient Flow: Identifiability and Recovery from Distribution SnapshotsNam D. Nguyen, Valeriya Malysheva
Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift $F$ only through its score-weighted divergence $\nabla\cdot F+F\cdot\nabla\logρ$, leaving a $ρ$-solenoidal gauge invisible to any single-time constraint. Time
benchmark - arxiv:2609.35058 · cs.LGTIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RLYibin Huang, Xinming Xu, Conghui Zhu
Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent metho
agenticbenchmark - arxiv:2609.35055 · cs.LGUniversality and Generalization of Causal Transformers Across Context LengthsTakashi Furuya, Maarten V. de Hoop, Gabriel Peyré
Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalize
long context - arxiv:2609.35052 · cs.CVOPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World ModelsHao Wang, Tao Yu, Liuzhou Zhang, HeXin Wang +12
Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an
embodiedworld modelmemorybenchmarkevaluatorevaluation protocol - arxiv:2609.35047 · cs.ROEMPIRIC: Experiment-Driven Learning of Residual World Models for Robot PlanningYichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist +12
A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We prese
manipulationworld modelagent - arxiv:2609.35046 · cs.CVLEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose EstimationHongli Xu, Zhaowei Lu, Junwen Huang, Jiaqi Hu +4
Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS co
benchmark - arxiv:2609.35043 · cs.CVMixed-Prior Decision Risk for Open-Set RecognitionL. A. Erlygin, P. D. Proskura, A. A. Zaytsev
In open-set recognition (OSR), a probe must either be identified as one of the known gallery classes or rejected as unknown, so three error types coexist: false acceptance, false rejection, and misidentification. An uncertainty score for selective recognition should rank probes by the risk of the de
benchmark - arxiv:2609.35039 · cs.RODo Not Cut When Uncertain: Rejectable and Calibrated Decision Heads for VLA Policies in Robotic HarvestingHeng Zhang
Vision-Language-Action (VLA) policies trained with behavior cloning or flow matching are optimized to output an action trajectory, but they cannot express "I don't know" or "I should not act." In robotic harvesting, occlusion makes single-frame decisions fundamentally ambiguous: identical pixels can
vision-language-actionvlamanipulation - arxiv:2609.35038 · cs.LGGUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian OptimizationJintao Wei, Chenxi Li, Songhao Wang
Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and task heterogeneity. We propose GUIDE-FBO,
agentbenchmark - arxiv:2609.35035 · cs.LGTHEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of LayoutGiuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni
The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDS
benchmark - arxiv:2609.35034 · cs.LGRecommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance DecompositionRiki Okamura, Toshiharu Sugawara
Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We pr
policy evaluation - arxiv:2609.35032 · cs.ROJRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World EnvironmentsZhixi Cai, Fucai Ke, Sukai Huang, Maria Garcia de la Banda +3
In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object,
embodiedworld modelagentagenticembodied agentbenchmark - arxiv:2609.35028 · cs.LGVEX-Bench: Benchmarking Verification Complexity of LLM-Generated MisinformationHanxun Huang, Yutao Wu, Qizhou Wang, Silvia Montaña-Niño +8
Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content
benchmarkscalable evaluationscalable evalllm-as-judge - arxiv:2609.35026 · cs.AIWebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web AgentsAnton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev +1
We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emi
agentjudge modelleaderboard - arxiv:2609.35025 · cs.AIAutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?Haotian Luo, Haoyu Wang, Zeyu Qin, Huanjin Yao +4
Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human
agentagenticself-improvementhuman-in-the-loopbenchmark - arxiv:2609.35023 · cs.CVProxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing ThemHongli Xu, Weilong Yan, Anbang Wang, Chunyu Zou +2
Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2Wor
world model - arxiv:2609.35022 · cs.CVDetection of Adversarial Attacks on Super-Resolvers Using Spectral FeaturesEmma J. Reid, Haley Duba-Sullivan, Tony G. Allen
The integration of deep learning models into image preprocessing pipelines such as super-resolution introduces a largely unexplored attack vector for adversaries targeting downstream tasks. To ensure trustworthiness of critical imaging pipelines, we must be able to detect adversarial behavior within
benchmark - arxiv:2609.35021 · cs.LGAddressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided MaskingGuangyu Wang, Jiawei Tong
Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent
benchmark - arxiv:2609.35017 · cs.AITermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation EvaluationNicolas Dahan, Fran{\cc}ois Yvon, Rachel Bawden
Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminolog
llm-as-judge - arxiv:2609.35005 · cs.LGSub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on DevicePaweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański +1
Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must m
memory - arxiv:2609.35003 · cs.ROLearning to Act under Visual Interruptions with Vision-Language-Action ModelsMingle Jiang, Rui Xu, Yunke Wang, Chang Xu
Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting
vision-language-actionvlavla modelmanipulationgr00tworld model - arxiv:2609.35000 · cs.ROGraph-Based Simultaneous Path and Foothold Planning for Multi-Limbed Intra-Vehicular Robots in Space StationsMasazumi Imai, Kentaro Uno, Toshinori Kuwahara, Kazuya Yoshida
Robot-aided operations in space stations are essential for reducing the workload of astronauts and improving the efficiency of on-orbit activities. Multi-limbed intra-vehicular robots (MLIVRs) equipped with grappling end-effectors have emerged as a promising solution, as they can securely grasp pre-
manipulationgrasp - arxiv:2609.34988 · cs.CLThe Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving AgentsYunhe Su, ZiYi Dong, Tong Yu, Weijian Deng +3
Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experie
memoryagenttool-useself-evolvingbenchmark - arxiv:2609.34982 · cs.ROActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuningDi Zhu, Ziheng Yan, Fang Wan
Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited gener
vision-language-actionvlavla modelmanipulationopenvlalibero - arxiv:2609.34979 · physics.app-phFar-field excitation of symmetry-protected bound states in the continuum by nonlinear virtual sourcesMarc Martí-Sabaté, Yijie Zhang, Shengming Sun, Ruxin Li +3
Bound states in the continuum (BICs) enable complete wave confinement within open systems through destructive interference or symmetry protection, giving rise to ideally infinite quality factors and extreme field enhancement. Their practical exploitation, however, is hindered by a fundamental parado
manipulation - arxiv:2609.34978 · cs.CVOne Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMUZhilin Guo, Boqiao Zhang, Oszkár Urbán, Josef Bengtson +7
Consumer earbuds already stream inertial motion data from the head, one of the most widely worn sensor locations on the body. We ask how much of the 3D body pose a single such head IMU can recover, and whether adding more consumer sensors actually helps. We build a multimodal capture pipeline that r
benchmark - arxiv:2609.34975 · cs.LGTeach to Learn: Hint Annealing for Self-improving LLM ReasoningZile Wang, Zijian Li, Haodong Wang, Jian Liu +4
Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based
self-improvingself-improvementbenchmark - arxiv:2609.34974 · cs.AIBefore Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive InterfacesRuozhao Yang, Mingfei Cheng, Xiaofei Xie
LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid
agentbenchmark - arxiv:2609.34973 · cs.AIAPEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex InteractionPuneet Mathur, Dinesh Manocha
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes su
benchmark - arxiv:2609.34972 · cs.CVJust MLPs: Efficient Visual State Reconstruction for Multimodal Language ModelsJingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang +2
Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers
benchmark - arxiv:2609.34971 · cs.AIAction-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema BiasYinhong Liu, Zhili Tan, Zilin Wang, Zhijiang Guo
Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally
agentllm agentagentictool-use - arxiv:2609.34969 · cs.RONavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic MemoryKai Sheng, Liuyi Wang, Jinlong Li, Haojie Dai +2
Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigat
embodiedmemorysemantic memoryembodied agent - arxiv:2609.34968 · cs.RORoboFL: Federated Expert Assembly for World Action ModelsRongyu Zhang, Ruizhi Fan, Yunfan Lou, Hengyu Fang +7
Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient
vision-language-actionvlarobotwinfranka - arxiv:2609.34963 · cs.AIJevVibe: Efficient Classification-Guided Secure Code GenerationArshak Rezvani, Sasha Behrouzi, Ahmad-Reza Sadeghi
Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking
agentbenchmark - arxiv:2609.34962 · cs.LGALICE: In-context, Zero-shot, Mutual Information EstimationGiulio Franzese, Simone Rossi, Pietro Michiardi
Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are more
benchmark - arxiv:2609.34960 · cs.AIProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic OptimizationFeiming Wang, Daibo Li, Kun Yuan
Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system
agent system - arxiv:2609.34951 · cs.AIAX is the New AEOIdo Finder, Assaf Elovic, Gad Shalev
In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter b
agentagentic - arxiv:2609.34949 · cs.AIVD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly DetectionMengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang +8
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abst
benchmark - arxiv:2609.34944 · cs.ROAdjoint Guidance Flow: Amortized Critic Guidance for VLA PoliciesJeongsol Kim, Youngjun Jun, Kyumin Choi, Youngmin Kim +5
Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic through a one-step surrogate of the sampl
vision-language-actionvlavla policyliberomemory - arxiv:2609.34934 · cs.CVTaoTex: Boosting Texture Detail Fidelity for Native 3D Material GenerationXiuchao Wu, Shuichang Lai, Jiangjing Lyu, Chengfei Lyu
Recent 3D generation models can produce accurate geometries while still struggling to reconstruct detailed textures. We propose a diffusion-based native 3D material generation model TaoTex, which faithfully recovers intricate textures through tailored strategies and improvements. First, we develop a
agent - arxiv:2609.34930 · cs.AIPDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM AgentsHuayi Lai, Shichao Song, Qingchen Yu, Simin Niu +3
Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive ex
memoryllm agenttool usebenchmark - arxiv:2609.34924 · cs.LGAudit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic CodingSebastian Bobadilla-Suarez, Bob Suh, Ryan Fortin
An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffo
agentagenticself-improvement - arxiv:2609.34913 · cs.AIDGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance BoardsJeremy Canale
Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in wh
agentmulti-agentbenchmark - arxiv:2609.34911 · cs.RODon't Throw Away the Tail: Action Upcycling for Policy AccelerationTaesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim +4
Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive
vision-language-actionmanipulation - arxiv:2609.34905 · cs.CVReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual ScoutsYaowen Zhang, Xiangyu Qiu, Junyi Hu, Zhi Lu +4
Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We tr
post-trainingbenchmark - arxiv:2609.34900 · cs.CVFILIGREE3D: Scaling Sparse Latent Flow Matching for Ultra-High-Resolution Image-to-3D GenerationHongjie Li, Xinran Yang, Xiuchao Wu, Jiangjing Lyu +1
Scaling image-to-3D generation to ultra-high resolutions requires controlling rapidly growing computational costs without sacrificing fine geometric detail. We present \textbf{Filigree3D}, a sparse latent flow-matching framework that generates 3D geometry from a single image at voxel resolutions up
memory - arxiv:2609.34895 · cs.CVLVMT: Video Mask Transformer for Long-term Video SegmentationNarges Norouzi, Niccol`o Cavagnero, Idil Esen Zulfikar, Bastian Leibe +2
Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across tim
memorybenchmark - arxiv:2609.34893 · cs.ROECHO: Event-Augmented Context with Hindsight and Outlook for Wrist-Only ManipulationXinyue Wang, Yicheng Jiang, Zesen Gan, Junhao He +7
Learning-based manipulation policies relying on RGB cameras often suffer from degraded observations under extreme exposure. Event cameras mitigate this degradation by asynchronously detecting pixel-level intensity changes to offer a high dynamic range. However, their observations heavily depend on c
manipulationgripperevent camera - arxiv:2609.34886 · cs.AIFewer Assumptions by Design: A Reusable Skill for LLM-Assisted Verus VerificationAndrada-Livia Antoneac, Dorel Lucanu, Dragoş Teodor Gavriluţ
LLM-assisted Verus verification is a less tedious method to verify Rust implementations, but paired with self-referential structures, e.g., Doubly Linked Lists (DLLs)ânotoriously difficult to formalise for verificationâit becomes a substantially more demanding verification task. Moreover, a spec
agentllm agent - arxiv:2609.34884 · cs.CVSubRot: Signed Gradient Subspace Calibration for VLM Rotation QuantizationZhenhao Shang, Haizhao Jing, Haokui Zhang, Guoting Wei +3
Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-
post-trainingbenchmark - arxiv:2609.34879 · cs.AIOne Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent RepairXiang Xia, Cheng Yan, Fan Xu, Zhijun Fan +2
Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both oper
benchmark - arxiv:2609.34867 · cs.CVP4Q: Co-designing Token Pruning and Quantization for Vision-Language Model AccelerationHaizhao Jing, Zhenhao Shang, Haokui Zhang, Rong Xiao +1
Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions,
memorypost-training - arxiv:2609.34866 · cs.LGFrom Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of TransformersNafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei, Milad Hosseini +1
Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objectiv
post-training - arxiv:2609.34864 · cs.AIOn the Limits of Metacognitive Monitoring in LLMsDongqi Han, Yifan Yang, Dongsheng Li
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier m
benchmarkevaluator - arxiv:2609.34863 · cs.CVRevisit to Segment: Working Memory Distillation for Reasoning SegmentationCilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu +3
Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our
memorybenchmark - arxiv:2609.34861 · cs.LGWhen Text Matters: Design Principles for Visual Token Pruning in Vision-Language ModelMinchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim +1
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection c
benchmark - arxiv:2609.34860 · cs.LGAttention-based Hierarchical Variational Information Bottleneck for Robust Multi-Agent Communication under Variable BandwidthLukas Koch Vindbjerg, Qi Zhang, Yury Brodskiy, Lukas Esterle
Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first part of each message is received. To add
multi-agent - arxiv:2609.34853 · cs.CVEviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary SegmentationSungho Moon, Kota Shimomura, Junwoo Park, Wonhyeok Choi +3
Open-vocabulary 3D scene understanding enables object localization and segmentation from free-form text queries without a fixed category vocabulary. Many recent methods build on 3D Gaussian Splatting and consolidate multi-view observations, such as masked crops from individual views, into language f
evaluation protocol - arxiv:2609.34851 · cs.ROLearning High-Risk High-Precision Motion ControlNam Hee Kim, Markus Kirjonen, Perttu Hämäläinen
Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of hig
benchmark - arxiv:2609.34849 · cs.LGWhen Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy DistillationXinke Jiang, Tao Feng, Zhibang Yang, Zhixin Zhang +3
Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family o
post-training - arxiv:2609.34848 · cs.AICan We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-DistillationYugu Li, Zehong Cao, Peizhen Li, Yang Zhang +2
RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contributio
benchmark - arxiv:2609.34843 · cs.CVORAV: Benchmarking Audio-Video Generation from Multimodal ContextsJiacheng Hua, Xiaokun Feng, Jiaqi Hua, Chang Liu +2
Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 tas
benchmark - arxiv:2609.34842 · cs.LGQiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous ModalitiesHanyin Cheng, Linfeng Wang, Zhengbo Qu, Yang Shu +6
Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two
benchmark - arxiv:2609.34840 · cs.AINociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One BodyWolfgang Maass
An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never upd
memoryagent - arxiv:2609.34839 · cs.CLOpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin VocalizationsFaadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero +6
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute
benchmarkevaluation protocol - arxiv:2609.34838 · cs.LGDivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn AgentsHanyang Wang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan +6
On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stal
benchmark - arxiv:2609.34834 · cs.CVTransform-Aligned Learned Features for Lossy Point Cloud Attribute CompressionYueru Chen, Pengpeng Yu, Dingquan Li, Wei Gao +2
Transform-based methods provide an effective framework for point cloud attribute compression by representing attributes as transform coefficients. Introducing learned spatial context into this framework requires mapping spatial representations to the transform domain, but this known basis change is
benchmark - arxiv:2609.34833 · cs.CVMulti-Scale Semantic Mapping in Urban Environments via Observation Calibration and Policy Dependence RegularizationRunling Long, Junhao Feng, Jia Wan
Semantic mapping is fundamental to embodied navigation, yet existing methods are developed for indoor environments, where objects exhibit relatively limited scale variation and are observed from a restricted range of viewpoints. Urban environments pose substantially greater challenges: agents must m
embodied - arxiv:2609.34832 · cs.AIBV Loss: Block Verification-Aware Loss for Block Diffusion Speculative DecodingSuyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang +4
Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatc
benchmark - arxiv:2609.34831 · cs.LGStructured Neural SDEs for Functional CalibrationFrancesco Piatti, Andrea Iannucci, Thomas Cass
Neural Stochastic Differential Equations (Neural SDEs) provide flexible continuous-time generative models, but generic neural drift and diffusion networks are costly to simulate on long horizons and can give unstable gradients when the training signal is a path functional rather than a pointwise obs
benchmark - arxiv:2609.34829 · cs.CLFrom Weak Task Specifications to Scientific Extraction Agents: Optimizing Task ConstructionZixiao Dong, Wei Yang, Zihao Liu, Chenshu Li +3
Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually
agent - arxiv:2609.34826 · cs.CVWM-VLM: Probing Internal World Models for Interleaved Visual-Textual ReasoningYuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen +4
Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, w
world model - arxiv:2609.34823 · cs.ROAGRO-SUVIDE: Agentic Robotics for Surgical Viscoelastic DebridementShutong Jin, Ziyang Chen, Preethi Satish, Medow Shen +3
Augmented dexterity has the potential to reduce the fatigue experienced by surgeons during repetitive surgical tasks. In this paper, we propose the first AGentic RObotics framework for SUrgical VIscoelastic DEbridement (AGRO-SUVIDE), the repeated removal of small fragments attached to a viscoelastic
agentagenticself-improving - arxiv:2609.34821 · cs.MACEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive MarketsAn Yan, Yu Huo, Zhiwei Shang, Yiran Peng +1
Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Ea
agentmulti-agentbenchmarkarena - arxiv:2609.34817 · cs.CVESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the WildHongyu Ma, Hairong Qu, Shiqi Zhao, Yongsong Yang +1
Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still ha
manipulationbenchmark - arxiv:2609.34810 · cs.AIUniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement LearningZenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Xiaofeng Han +7
Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by ev
agentic - arxiv:2609.34807 · cs.CVControlTrace: Recovering Control Fields for Hidden-Content RecognitionZijian Liu, Yaoguang Chen, Liwei Liu, Weixi Wu +3
Spatially conditioned diffusion models can embed words and contours in natural-looking images, but vision-language models (VLMs) may fail to recognize the hidden content. Transformation-based recovery depends on parameter and view selection. To evaluate hidden-content recovery and recognition, we co
benchmark - arxiv:2609.34805 · cs.AISIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RLZenghuang Fu, Ningqi Chen, Mingda Jia, Xiaofeng Han +9
Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh sibl
agenticbenchmark - arxiv:2609.34804 · cs.LGPhysics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water InfrastructureJeff Nijsse, Shu Su, Benjamin Oholeguy, Sreenivas Sremath Tirumala
Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model updates derived from fabricated telemetry r
benchmark - arxiv:2609.34800 · cs.CLPass or Fail? Evaluating LLMs on Two Greek Examination BenchmarksPanagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis
The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in th
benchmark - arxiv:2609.34799 · cs.AISTRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse ContextsWanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun +3
Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every s
evaluation frameworkscalable evaluationscalable evalevaluation protocol - arxiv:2609.34798 · cs.CVInfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware SupervisionGuanghao Zhu, Zeyu Liu, Zhitian Hou, Pengkai Wang +8
Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity,
post-trainingbenchmark - arxiv:2609.34792 · cs.CVD$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic ManipulationZijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng +8
Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM)
vision-language-actionmanipulationliberorobotwinmemorybenchmark - arxiv:2609.34790 · cs.AICoSec: Benchmarking Agent Security in CommunitiesHao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao +11
LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthor
persistent memoryagentllm agentagent systembenchmark - arxiv:2609.34785 · cs.LGBEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and VerificationYuheng Wu, Berk Gokmen, Sujeeth Jinesh, Lauren McLane +4
Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE,
agentagenticself-improvingself-improvementevaluator - arxiv:2609.34783 · cs.CLTQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series DatabasesFei Lyu, Zhiyi Peng, Jiaming Liu, Yixuan Yang +4
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domain
benchmark - arxiv:2609.34782 · cs.ROCoHuB: A Simulation Benchmark for Multi-Humanoid CollaborationHyunjin Park, Jebeom Chae, Minwoo Park, Sunghyun Park +10
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduc
humanoidteleoperationbenchmark - arxiv:2609.34781 · cs.CVWhen VLMs Trust Context: Evaluating Scene Text Recognition under Misleading ContextYuxing Cheng, Yuan Wu, Yi Chang
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of
benchmark - arxiv:2609.34776 · cs.AIPage-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFsAbdelhak kelious
We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the c
retrieval-augmented - arxiv:2609.34772 · cs.AIBefore the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMsYadong Wang, Siping Yue, Yu Tian, Chuanxing Geng +1
Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot deter
benchmark - arxiv:2609.34769 · cs.AILongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual PuzzlesBingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li +2
GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains
memorybenchmark - arxiv:2609.34768 · cs.CVPrivacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model SupervisionShuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin
Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-mod
benchmark - arxiv:2609.34765 · cs.LGBeyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language ModelsMinchan Kang, Kyeonghye Park, Seungyeon Sa, Seoyoung Cho +2
Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize r
post-training - arxiv:2609.34764 · cs.AIWeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM PlansMin Jia, Shihao Zhou, Jun-Peng Zhu, Peng Cai +8
Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. How
knowledge graphself-evolving - arxiv:2609.34759 · cs.ROPanoVLN: Towards Effective Panoramic Vision-and-Language NavigationZhen Wang, Changpeng Wang, Zhe Liu, Zhangyang Qi +4
Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward
quadruped - arxiv:2609.34754 · cs.CLDraft-KV: Learning Useful Latent Communication Between Language ModelsLinquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang +6
Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most
memory - arxiv:2609.34749 · cs.CVCoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative DrivingYu Meng, Baining Zhao, Junta Wu, Tengfei Wang +9
Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-
world modelmulti-agentbenchmark - arxiv:2609.34745 · cs.LGNo Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy DistillationSeonghyeon Kim, Chaeyun Jang, Noah Lee, Boseop Kim +1
Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOP
post-trainingbenchmark - arxiv:2609.34743 · cs.ROSimulation for Planetary Robotic Perception and Autonomy: A Concise Survey of Recent Capabilities and GapsHoyun Kim, Giseop Kim
Planetary robotics is an important enabler of scientific exploration in environments where direct human-in-the-loop operation is costly, hazardous, or infeasible. However, developing and validating planetary robotic systems remains difficult because representative field testing is expensive, limited
human-in-the-loop - arxiv:2609.34730 · cs.ROAction Sequence Transfer via LLMs for Heterogeneous EnvironmentsChoongho Chung, DongHwan Shin, Sung-Hee Lee
We present an action sequence transfer system that adaptively transfers user action sequences across different target spaces. Given an input action sequence from a source space and scene graph representations of both the source and target environments, our system predicts a corresponding action sequ
scene graph - arxiv:2609.34727 · cs.AIDynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUsZhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun +5
On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing design
memorylong-context - arxiv:2609.34724 · cs.RODexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human DemonstrationsNaichuan Sun, Haotian Shen, Yizhang Zhang, Luying Feng +4
Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist p
manipulationdexteroushumanoid - arxiv:2609.34722 · cs.CVGeometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video GenerationZesong Yang, Weikai Chen, Liyuan Cui, Lutao Jiang +9
Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric
memory - arxiv:2609.34720 · cs.CVDBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake DetectionFengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb +1
As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and
manipulationbenchmark - arxiv:2609.34718 · cs.LGQuality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement LearningZijun Weng, Zhongan Bi, Xuanang Gao, Xiaohui Hu +4
Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii)
benchmark - arxiv:2609.34717 · cs.CLReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code GenerationHuifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan
Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search fr
memory - arxiv:2609.34715 · cs.AIPDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEsZhentao Tan, Jianrong Zhang, Ruijie Quan, Yi Yang
Physical trajectories contain more than snapshots of a system: they also reveal how its states evolve under governing conditions. However, representation learning for parametric partial differential equations (PDEs) has largely relied on reconstruction-based objectives that emphasize recovering obse
latent dynamicsbenchmark - arxiv:2609.34712 · cs.AIRSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient AgentsHao Li, Hangfan Zhang, Zhiyao Cui, Chunjiang Mu +4
Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance-cost trade-off can be improved through within-task large-small model collaboration, as smaller models can handle some stages even
agenticself-improvementbenchmark - arxiv:2609.34711 · cs.LGLearning Propagation Geometry from Message-Passing FeedbackYingxu Wang, Kunyu Zhang, Xinwang Liu, Mengzhu Wang +3
Learning local geometry enables graph neural networks (GNNs) to adapt how they compare and integrate neighborhood information. However, estimating geometry from aggregated representations can overlook variation among individual messages and dependencies across feature dimensions. We propose GeoF, a
benchmark - arxiv:2609.34710 · cs.AIFromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football ManagementPeiyu Zang
Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark
agentllm agentagent benchmarkbenchmark - arxiv:2609.34707 · cs.ROSOR-Nav: Search or Relocate? Context-Gated Exploration and Cross-Region Relocation for Object NavigationYuan Ji, Zirui Li, Yuxin Cai, Shuge Wu +2
Object navigation requires an embodied agent to find an object in an unseen environment under partial observability and a limited motion budget. Existing methods primarily optimize where the robot should go next by ranking candidate destinations. In contrast to these methods, we present SOR-Nav, a h
embodiedagentembodied agentbenchmark - arxiv:2609.34702 · cs.ROMarsLab: A Martian Rover Simulator for Planetary Rover Autonomous NavigationHoyun Kim, Beomsu Kim, Giseop Kim
Future Mars missions will require rover autonomy that can operate across unstructured terrain, changing illumination, atmospheric dust, and limited communication. Simulation is a practical way to study these conditions before deployment, but existing Mars-relevant resources differ in scope, includin
benchmark - arxiv:2609.34701 · cs.LGResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent SystemsTarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta +2
Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that
agentmulti-agentagent systembenchmark - arxiv:2609.34697 · cs.CVTriangular Resampling for Long-Horizon Motion GenerationKunhang Li, Yiyi Cai, Xiangyue Zhang, Fangyuan Tu +4
We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference st
post-training - arxiv:2609.34691 · cs.AIUsing LLMs to Detect LLM-Generated Texts: A Cross-Generation AnalysisHaiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang +2
Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self
benchmark - arxiv:2609.34688 · cs.CVUnified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA GeneralizationZhiyuan Ma, Jiaming Li, Lingzhen Li, Yu Liu +6
Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-rewa
vision-language-actionvlapost-training - arxiv:2609.34687 · cs.AIVCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual ExperienceSiqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu +5
Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning
embodiedagentembodied agentbenchmark - arxiv:2609.34686 · cs.AIJailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool AgentsXi Wang, Songlei Jian, Yiming Zhang, Bin Ji +4
As large language models increasingly operate as tool-using agents, post-jailbreak safety feedback is often assumed to serve as a reliable safeguard; however, how lingering jailbreak context shapes subsequent agent behavior remains largely unexplored. To systematically examine this dynamic, we intro
agentautonomous agent - arxiv:2609.34684 · cs.RONatural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA ReadoutsHyungjoon Kim, Wonbin Son, Mi Young Lee, Jun Young Lee +1
Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary
vision-language-actionvlaevaluation framework - arxiv:2609.34683 · cs.LGAgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMsCheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin, Yao Lai +5
The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks prima
memoryagentictool-usetool callingbenchmark - arxiv:2609.34682 · cs.CVV-Gym: Enhancing Agentic Visual Reasoning via Skill-Data Co-EvolutionBei Yan, Yuecong Min, Jie Zhang, Junqi Yang +2
Advances in multimodal understanding, reasoning, and tool use enable agents to tackle increasingly complex visual reasoning tasks. By distilling past execution experience into reusable skills, agents can transfer lessons from both successes and failures into future reasoning, reducing repeated error
agentictool useself-improvementbenchmark - arxiv:2609.34681 · cs.LGSOLAR: A State-Driven Online Learning Rate Scheduler for LLM PretrainingQiulin Shang, Binyu Wang, Yongqi Qiao, Songde Rao +2
Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimiza
online learning - arxiv:2609.34680 · cs.LGQuantForge: Discovering Residual Decompositions for MXFP4 Post-Training QuantizationQiulin Shang, Zhoutong Wu, Jie Hu, Kun Yuan
Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the netw
memorypost-trainingevaluator - arxiv:2609.34679 · cs.LGFrom Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based ModelingWangxuan Fan, Xiaoyu Nie, Zhoutian Shi, Xiangcheng Meng +4
Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under lim
llm agentonline learning - arxiv:2609.34678 · cs.AIFair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoShMuhammad Ahmad, Fatemeh Seyedin, Adrian Weller, Dongwon Lee +1
Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We t
benchmark - arxiv:2609.34677 · cs.LGLearning What to Recall: Adaptive Multi-Cue Episodic Memory for World ModelsBeomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar +2
World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for
world modelmemoryepisodic memory - arxiv:2609.34674 · cs.ROHOI-Retarget: Contact-Centric Retargeting for Human-Object InteractionJihwan Shin, Adrià López Escoriza, Junzhe He, Matthias Heyrman +1
Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method t
humanoid - arxiv:2609.34666 · cs.ROOn the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material ManipulationXintong Yang, Minglun Wei, Yu-Kun Lai, Ze Ji
Differentiable physics is increasingly used in robotic material manipulation for system identification, trajectory or skill optimization, demonstration generation, and robot or end-effector design. These applications depend on gradients propagated through long, contact-rich simulation rollouts. We s
manipulationbenchmark - arxiv:2609.34661 · cs.AICodoku: Renewable Program-Reasoning Challenges for Frontier Coding AgentsCong Li, Hao Sun, Zenan Li, Zhendong Su
Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input. Coding agents break two assumptions on which these benchmarks rest: an agent can recover the answer by executing the program instead of reasoning about it, and fixed task sets drawn from
agenttool usebenchmark - arxiv:2609.34660 · cs.CLRewarding Novel Deductions: Solver-guided Process Rewards for Logical ReasoningMuhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza
Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing i
benchmark - arxiv:2609.34658 · cs.CVCapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow ModelsPengyang Ling, Yujie Zhou, Jiazi Bu, Yibin Wang +4
Reward-specialized post-training produces strong experts for flow-based generative models, while multi-teacher on-policy distillation (OPD) consolidates their capabilities into a single student. Existing methods, however, route each prompt to a single teacher according to its semantic category, impl
post-training - arxiv:2609.34656 · cs.LGMinimax Last-Iterate Convergence in Matrix Games with Observed ActionsYuheng Zhang
We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneous
memory - arxiv:2609.34654 · cs.AIA General Harness for Protein Foundation Model Fitness PredictionYang Tan, Qijia Tian, Gangyu Sun, Bozitao Zhong +4
Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their
benchmark - arxiv:2609.34653 · cs.LGOmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM StreamingZongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu +1
On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention
memorybenchmark - arxiv:2609.34652 · cs.CVRevitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language PerspectiveGuoqi Yu, Juncheng Wang, Shujun Wang
Medical time series (MedTS) underpin many clinical classification tasks, yet existing methods usually represent them only as numerical sequences and underuse the morphology that is explicit in waveform inspection. To bridge this gap, we introduce Vision-Informed Retrieval (ViRe), which uses a frozen
benchmark - arxiv:2609.34649 · cs.AIBeyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent HarnessesWeiyuan Li, Jinghan Xu, Aili Chen, Xintao Wang +3
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottl
agentllm agentself-evolvingbenchmark - arxiv:2609.34648 · cs.AISEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech FlowsTianxin Xie, Pengfei Zhang, Kai Jiang, Zelin Zhao +1
Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion ed
benchmark - arxiv:2609.34645 · cs.AINereus: Adaptive Parallelism for LLM Post-TrainingSonglin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang +2
Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a c
memorypost-training - arxiv:2609.34642 · cs.LGTilted Schrödinger Bridge MatchingSergei Kholkin, Evgeny Burnaev, Alexander Korotin
Schrödinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion model
post-training - arxiv:2609.34641 · cs.CVBackdoor as Probe: Test-Time Adversarial Defense for CLIPZhongqi Wang, Jie Zhang, Nie Sen, Zhiyu Chen +2
Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposin
benchmark - arxiv:2609.34636 · cs.AIMechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative PhysicsDanilo Gusicuma, André Freitas
This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying stru
benchmark - arxiv:2609.34634 · cs.AIA Persistent State for Auditable Mixture-of-Experts RoutingAbdurrahman Javat, Allan Kazakov
Mixture-of-Experts (MoE) models repeatedly route tokens to sparse subsets of experts, but conventional routers expose no routing-specific record of how cross-layer influences accumulate. We introduce Scratchpad-Augmented Mixture-of-Experts (SA-MoE), which gives each router access to a low-dimensiona
persistent state - arxiv:2609.34633 · cs.LGGenMem: Generative Symbolic Memory for Self-Evolving HarnessXinke Jiang, Tao Feng, Weixuan Xu, Zhixin Zhang +6
Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address
memoryagentllm agentmulti-agentself-evolving - arxiv:2609.34630 · cs.CVLong Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric VideosFangzhou Ma, Ivo Alexander Ban, Eren Homburg, Gabriele Goletto +4
Real-world AI systems must reason about objects that are no longer visible: an AR assistant guiding a user back to an object used earlier, a household robot retrieving an item someone put away. This requires not just recalling where an object was last seen, but updating its state when it is moved an
benchmark - arxiv:2609.34629 · cs.LGDisKO: Deep Koopman Learning in Distribution Space from Unpaired SnapshotsHe Ma, Xiaochen Liu, Wanfeng Lu, Ying Wang +2
Many complex systems are observed only through temporally unpaired distribution snapshots, making trajectory-based dynamical learning difficult without additional assumptions. We therefore formulate the problem directly in distribution space, treating the distribution itself as the dynamical state.
benchmark - arxiv:2609.34619 · cs.AILLMs for Executable Multi-Agent System Specification GenerationAndreas Kouvaras, Periklis Mantenoglou, Alexander Artikis
MAS specifications express the effects of the actions of the agents and their environment, as well as other temporal phenomena, such as the intervals during which an agent may perform an action. The specification of a MAS should also be executable in order to allow for run-time monitoring. Construct
agentmulti-agentagent system - arxiv:2609.34613 · cs.LGProbabilistic Geodesic Flow Matching on Location-Scale FamiliesZeyuan Yu, Zhi Chang, Shiwei Lan
Flow matching (FM) has recently emerged as a promising framework for generative modeling due to its conceptual simplicity and strong empirical performance. In FM, samples are transported along a vector field parameterized by a neural network, inducing a probability path that evolves from a simple no
benchmark - arxiv:2609.34608 · cs.ROEfficient World Action Model Inference with Adaptive Intermediate StatesZhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma +7
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, obser
liberorobotwin - arxiv:2609.34606 · cs.CVWorldAttention: An Efficient Attention Architecture for Interactive Video World ModelsZeyu Zhang, Jinyuan Mao, Dakai An, Wangbo Zhao +5
Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current
embodiedworld modelmemory - arxiv:2609.34605 · cs.LGPMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy DistillationYouzhi Liu, Ruobing Zheng, Boyuan Tong, Tianqi Li +4
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teach
post-training - arxiv:2609.34604 · cs.LGShaping Persistent Representations from Independent InteractionsJi Dai, Quan Fang, Junyu Gao, Rongfeng Guo +3
World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone, without organizing persistent informat
tactileworld model - arxiv:2609.34603 · cs.AIAfter the Fix: How Corrected Agent Histories Transfer to Related TasksYanfei Zhang, Xu Lin
Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-stat
memoryagent - arxiv:2609.34599 · cs.LGThe Low-Rank Structure of VLA Reinforcement LearningMinjae Oh, Yoonah Park, Jongwon Lim, Yohan Jo
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $π_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CA
vision-language-actionvlavla modelgr00tliberopost-training - arxiv:2609.34598 · cs.CVSummarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal GroundingNanxing Hu, Xiaoyue Duan, Qiwei Yan, Kailin Lyu +2
Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs
memory - arxiv:2609.34596 · cs.CVTemporal Modelling for Burn Scars on Sentinel-3Luca Barco, Edoardo Arnaudo, Andrea Bragagnolo, Claudio Rossi +1
Rapid and accurate burn scar delineation from satellite imagery is essential for post-fire damage assessment. Sentinel-3 OLCI, with daily revisit and 21 spectral bands, suits rapid mapping, yet most pipelines treat acquisitions independently, leaving the pre/post-fire change signal unexploited. We p
benchmark - arxiv:2609.34594 · cs.ROAerial GRIPPER: A Gradient-based Real-time Inverse-game Predictor and PlannerZeshuai Chen, Meng Wang, Jindou Jia, Xiang Yu +1
Accurate capture of non-cooperative targets is critical. In an attempt to tackle this intractable challenge, an aerial gripper system integrated with a Gradient-based Real-time Inverse-game Predictor and PlannER (GRIPPER) framework is proposed. The interaction is formulated as a general-sum pursuit-
grippergrasp - arxiv:2609.34593 · cs.LGDeep Weighted Bellman Residual Minimization for $Q^*$ EstimationLican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong
Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample u
policy evaluation - arxiv:2609.34590 · cs.LGThe Model Knows When to Stop: Training-Free Early Stopping for Long-Context ReadingMuath Alyobi, Mohamed Eltahir, Almoayyad Abuljdail, Riyadh Almutawa +2
Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model w
long-contextbenchmark - arxiv:2609.34587 · cs.CVReinforcement Learning from Intermediate Renders for Image-to-Code GenerationOmri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel
Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that
post-training - arxiv:2609.34586 · cs.LGAgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud ContinuumMichalis Kasioulis, Moysis Symeonides, George Pallis, Marios D. Dikaiakos
Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent development, observability, or benchmarking s
agentagenticbenchmark - arxiv:2609.34585 · cs.LGCompute Time Scaling with Recursive Models for Combinatorial OptimizationZhengxi Zhang, Paul Swoboda
We propose Tiny Recursive Models for Combinatorial Optimization (\ours{}), a general neural method for combinatorial optimization that scales both depth (how often we recursively invoke our network) and width (how much we sample in parallel). Both are fundamental for combinatorial optimization: hard
benchmark - arxiv:2609.34581 · cs.CVCounterfactual Attention Policy Distillation for Temporal Video GroundingShaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu +4
Temporal video grounding is a key capability of advanced \emph{Multimodal Large Language Models} (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspe
benchmark - arxiv:2609.34575 · cs.AIDiffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement LearningHengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang
Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoa
manipulationbenchmark - arxiv:2609.34572 · cs.LGNudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own CompetenceRohit Saxena, Utkarsh Upadhyay
Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), ho
post-training - arxiv:2609.34571 · cs.AIPersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona RepresentationsRui Xu, Yinghui Xu, Libo Wu
Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations---addition, scaling, and linear interpolation---u
benchmark - arxiv:2609.34567 · cs.CVRecent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent IntelligenceMuhammad Owais, Ehtesham Iqbal, Samee Ullah Khan, Muhammad Umraiz +2
This review examines the evolution of plant phenotyping from conventional manual trait measurement to high-throughput, robotic, and artificial intelligence-driven crop monitoring. Despite significant advances in imaging, autonomous platforms, multimodal sensing, and deep learning, current phenotypin
agenticagent frameworkbenchmark - arxiv:2609.34565 · cs.AIFlowState: Execution State as Memory for Long-Horizon LLM AgentsMinghao Li, Bangyan Li, Zifan Wang, Yulong Li +4
Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task prog
memoryllm agent - arxiv:2609.34564 · cs.CVGLF-Q: Global-Local Feature-based Quantization for Vision TransformersPeilin Sun, Guang Liang, Jin Tong, Jianxin Wu
Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervisi
post-training - arxiv:2609.34561 · cs.LGBrain-Conditioned Action Policies for Neural Motor DecodingLuyao Jin, Running Zhao, Huan Zhao, Vincent C. K. Cheung +1
Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarce paired neural-action data. We propose
vision-language-actionvlaopenvla - arxiv:2609.34556 · cs.LGRoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional EncodingsJarod Lévy, Mathurin Videau, Jad Yehya, Jean-Rémi King +2
Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearb
benchmark - arxiv:2609.34557 · cs.AISkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal AgentsBingqing Jiang, Guoxi Zhang, Jasper Wang, Auric Wang +6
Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervisi
agentagent benchmarktool usebenchmarkevaluator - arxiv:2609.34555 · cs.LGPulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM DecodingQiuyang Zhang, Kai Zhou, Kai Lu, Haocheng Lu +5
Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that ex
long-context - arxiv:2609.34554 · cs.ROWhere Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAsTanguy Dieudonné, Jack B. Jedlicki, Heng Yang
Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks sho
vision-language-actionmanipulationmemorybenchmark - arxiv:2609.34550 · cs.ROGaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-TuningYihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang +3
Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleopera
vision-language-actionvlamanipulationteleoperation - arxiv:2609.34548 · cs.AISGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon ReasoningJaeho Jung, Sung Hoon Jung
Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLAC
llm agent - arxiv:2609.34547 · cs.CVActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language ModelsGueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy +2
Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targe
benchmark - arxiv:2609.34545 · cs.AIRemember Before You're Asked: MemDream for Self-Probing Memory EvolutionMingfei Lu, Mengjia Wu, Runsong Jia, Zhe Luo +1
Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This re
memoryllm agent - arxiv:2609.34540 · cs.AIAPOLO: Automatic Prompt Optimization for Ontology LearningHuu Tan Mai, Roman Kochnev, Cuong Xuan Chu, Lukas Lange +2
Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimiza
multi-agentagent system - arxiv:2609.34539 · cs.CVTSGate: Timestep-Aware Gated Attention for Diffusion TransformersBoyu Zhang, Yifan Liu, Shuxia Lin, Qingjian Ni +2
Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-
benchmark - arxiv:2609.34537 · cs.AIThe Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn InteractionsXiaoting Lyu, Xinbo Ma, Yufei Han, Hangwei Qian +4
Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \te
benchmark - arxiv:2609.34528 · cs.CVPreference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt DisagreementHyun-Kurl Jang, Jihun Kim, Kuk-Jin Yoon
Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial ins
benchmark - arxiv:2609.34526 · cs.AIPairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference UseMingfei Lu, Mengjia Wu, Yi Zhang
Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less a
memorybenchmark - arxiv:2609.34510 · cs.AICan AI Make Money in Crypto? Measuring the Gap from Backtests to Real MarketsXingtong Yu, Jiarun Zhou, Guanlin Ding, Wenkang Wei +11
AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can g
benchmark - arxiv:2609.34506 · cs.AIDoes Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision BenchmarksManya Singh, Arjun Pakrashi
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on
benchmark - arxiv:2609.34503 · cs.LGDistribution-Conditioned Task Routing for Class-Incremental LearningLonghuan Xu, Zhipeng Zhou, Wei Ji, Chunyan Miao +2
Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all classes seen so far without access to its t
benchmark - arxiv:2609.34502 · cs.CVSubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot StorytellingXinyu Wang, Huafeng Shi, Zian Li, Yan Zhou +3
We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retai
memory - arxiv:2609.34499 · cs.LGScalable GNN-based Knowledge Graph Representation Learning with Efficient Message PassingHuu Tan Mai, Cuong Xuan Chu, Heiko Paulheim, Daria Stepanova
Graph neural networks (GNNs) excel at representation learning on Knowledge Graphs (KGs), achieving stateof-the-art performance on tasks like link prediction or entity classification. However, their high computational complexity, inherent to their user-defined message passing (MP) algorithm, still pr
knowledge graph - arxiv:2609.34497 · cs.LGQAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement LearningYuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou +3
Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We
humanoid - arxiv:2609.34496 · cs.MAMASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent SystemsYapeng Li, Songze Li, Shuang Yu, Jing Yu +3
LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost
agentmulti-agentagent systembenchmark - arxiv:2609.34492 · cs.AIPowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power SystemsXijing Wang, Yinsheng Yao, Jinru Ding, Yidong Jiang +4
Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chain
llm agentagentictool usebenchmark - arxiv:2609.34491 · cs.LGM3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular OptimizationJunjie Wang, Yaowei Jin, Ruohui Tang, Guonan Cui +6
Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they must recover candidate identities, prior
multi-agentbenchmark - arxiv:2609.34488 · cs.LGFlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RLXun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li +1
Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to natur
post-training - arxiv:2609.34486 · cs.ROModel-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step PredictionVictor Paredes, Ayonga Hereid
Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches.
humanoidwhole-body control - arxiv:2609.34484 · cs.ROARS: Agentic Reward System for Robot LearningSheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng +4
Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework fo
manipulationagentagenticbenchmark - arxiv:2609.34481 · cs.CLCARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASRBashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh +35
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed
benchmark - arxiv:2609.34480 · cs.CVWhen Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and ScenesSungguk Cha, Mintae Kim, Youngsub Han, Byoung-Ki Jeon +1
Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each ques
benchmark - arxiv:2609.34470 · cs.CVPrecise Editing and Flexible Referencing for Interactable WorldsXinyao Liao, Xianfang Zeng, Zhu Liang, Zhoujie Fu +4
We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends
world model - arxiv:2609.34469 · cs.CLThe Last Mile Is the File: OfficeEditBench for Preservation-Aware Office EditingZhiwen Wu, Chengxu Wu
A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, a 170-task benchmark for change-scoped m
benchmark - arxiv:2609.34467 · cs.LGAlignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy LearningShengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng +4
Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, whic
vision-language-actionvlavla modelmanipulationbenchmark - arxiv:2609.34463 · cs.AICoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based AgentsXiao Yang, Yangchen Ou, Yuhan Gao, Le Wang +2
Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based def
agentbenchmark - arxiv:2609.34460 · cs.LGWhen Does Structured Knowledge Help Neural Theorem Proving?Sareh Nabi, Roland Vogl, Marzieh Nabi
Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analog
knowledge graph - arxiv:2609.34459 · cs.AIEscaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement LearningYijie Sun, Sanquan Sun, Yanda Zhu, Yuanyang Zhu +2
Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To ad
agentmulti-agent - arxiv:2609.34457 · cs.LGZonoGPT: Towards An Abstract Domain for Verifying Large GPT ModelsHai Duong, Thanh Le, ThanhVu Nguyen
Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and provide auditable guarantees before deploym
agentic - arxiv:2609.34455 · cs.CLRGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their JustificationsJianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian +6
We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require
benchmarkevaluator - arxiv:2609.34453 · cs.LGSPACE-LoRA: Allocating Activation-Subspace Protection for Continual LearningSeunghyun Yoo, Kiseok Kim, Hyeontae Joo, Junyeop Bang +1
This study addresses the catastrophic forgetting problem that occurs when sequentially learning successive tasks using Low-Rank Adaptation (LoRA) from a lifelong learning perspective. While existing approaches have primarily constrained parameter updates or learning subspaces to reduce interference
lifelong learning - arxiv:2609.34447 · cs.LGUnbiased Top-$k$ Estimation for On-Policy DistillationLinjian Meng, Siyuan Gan, YuHan Li, Xiran Wang +4
On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student vi
post-training - arxiv:2609.34446 · cs.LGLivin' on a Prior: Likelihood Score Approximation for Inverse ProblemsRostislav Makarov, Tal Peer, Danilo de Oliveira, Timo Gerkmann
Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly from paired data. We target a setting th
post-trainingbenchmark - arxiv:2609.34444 · cs.AISocial Circuits behind Multi-agent Echo ChambersChuiyang Meng, Wenlu Yu, Ming Tang, Cheng Li
Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent's internal activations and affects its decision. In this work,
multi-agent - arxiv:2609.34442 · cs.LGMaking LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge PathsJialu Wang, Peizhi Niu, Haoteng Yin, Hans Hao-Hsun Hsu +2
While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact
knowledge graph - arxiv:2609.34440 · cs.ROWhen the Score Becomes the Target: Rethinking Metric Validity in Autonomous DrivingMorui Zhu, Deyuan Qu, Qi Chen, Kentaro Oguchi +1
Driving benchmark scores are increasingly used not only for evaluation but also as optimization targets. This raises a fundamental question: do score gains remain reliable evidence of driving improvement once the score itself is optimized? We address this question by examining how the scoring proces
benchmark - arxiv:2609.34438 · cs.AIRemember by Asking: Retrieval-Induced Memory Evolution for LLM AgentsWanqi Zhou, Jiawei Lu, Yang Wang, Zhaolong Xing +3
Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future
memoryllm agent - arxiv:2609.34428 · cs.AIAgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question AnsweringChanhee Park, Jeongho Yoon, Sungbin Han, Hyeonseok Moon +1
Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagn
agentagenticagent benchmarkbenchmarkleaderboard - arxiv:2609.34427 · cs.LGLLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale OptimizationShihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen +3
Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle
benchmark - arxiv:2609.34426 · cs.LGQ-learning Penalized Transformer for Safe Offline Reinforcement LearningShengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou +4
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety co
benchmark - arxiv:2609.34425 · cs.AIZero-Shot Cue-Grounded Topic Segmentation of Spoken DocumentsSuhwan Choi, Myeongho Jeon, Myungjoo Kang
Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to ada
benchmark - arxiv:2609.34423 · cs.LGOn the Relation Between Interval Regret and Dynamic RegretYi-Han Wang, Peng Zhao, Zhi-Hua Zhou
Non-stationary online learning has attracted much attention in recent years, as static regret is insufficient to guide algorithm design in changing environments. To address this limitation, interval regret and dynamic regret have been introduced as two representative performance metrics that strengt
online learning - arxiv:2609.34422 · cs.LGCoding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement LearningLirui Luo, Kelong Mao, Heming Xia, Rongqing Li +9
Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environm
memoryagent memoryagentagenticpost-training - arxiv:2609.34420 · cs.LGMegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid ParallelismTong Qiao, Ao Zhou, Yingjie Qi, Chunming Hu +1
Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated
memory - arxiv:2609.34419 · cs.AIBeyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction GenerationHanzhong Zhang, Jindong Wang, Siyang Song
Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG
agentagent framework - arxiv:2609.34415 · cs.LGPROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time SafetyDing Jia, Wei Liu, Xianglong Du, Yingjie Li +4
The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PRO
long-contextbenchmark - arxiv:2609.34414 · cs.ROFrom World Models to World Action Models: Rethinking Next-State PredictionTingyu Yuan, Ziming Ji, Biaoliang Guan, Wen Ye +8
Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of
liberoworld modelaction-conditioned - arxiv:2609.34412 · cs.ROFrom Language to Task Maps: Compiling Semantic Relations While Preserving Task-Relevant FreedomJaegyun Park, Jingwang Lee, Jungsoo Lee, Soonwoong Hwang +1
Natural-language manipulation instructions specify qualitative relations, whereas continuous controllers require state-evaluable task quantities, differentials, and completion conditions. Because a qualitative relation generally leaves part of the relative configuration unspecified, expanding it int
manipulationfranka - arxiv:2609.34408 · cs.CVDistilling Visual Reasoning into Text SpaceWenhan Yang, Nilay Naharas, Ali Payani, Baharan Mirzasoleiman
Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these represe
benchmark - arxiv:2609.34406 · cs.LGUnlocking Few-Step Diffusion for Faithful PreviewsJing Jia, Sifan Liu, Guanyang Wang
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can cl
benchmark - arxiv:2609.34404 · cs.AIEval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender SystemsCong Wang, Shoujin Wang, Yishuo Li, Qi Zhang +2
Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of di
benchmarkevaluation frameworkevaluation protocol - arxiv:2609.34398 · cs.LGGeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence SamplingMoshe Eliasof, Eldad Haber
Critical mineral discovery is a positive-only problem: deposits are observed as sparse locations, while unlabeled regions are not reliable negatives, and similar geophysical signatures can arise from different subsurface states. We therefore model mineral targeting as learning a conditional spatial
benchmark - arxiv:2609.34397 · cs.AISkillFocus: Evolving Agent Skills via Capability DecompositionNing Wang, Zhiren Gong, Bingdong Li, Peng Yang +1
Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision t
agentbenchmark - arxiv:2609.34396 · cs.CVHyperDAM: Hyperspectral Distractor-Aware Memory with Amodal Expansion for SAM 3 TrackingRyoga Yuzawa, Tasuku Takagi
Hyperspectral video provides material cues that can disambiguate targets with similar false-color appearance, yet foundation-model trackers update memory primarily from spatial and appearance evidence. We present HyperDAM, a DAM4SAM3-based hyperspectral tracker with three principal contributions. Fi
memoryleaderboard - arxiv:2609.34392 · cs.AIOrg-Agent: Beyond Personal Assistants Towards Organizational AgentsLuyao Zhuang, Yujing Zhang, Zijin Hong, Yilin Xiao +1
Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-user memory and knowl
memorytool use - arxiv:2609.34391 · cs.LGP2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response PredictionHaojie Yang, Ran Su
AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never ob
memory - arxiv:2609.34390 · cs.CVVastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal TrackingZhizhen Li, Zan Wang, Huidong Peng, Bohan Tan +3
Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large
benchmark - arxiv:2609.34387 · cs.CVCAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous DrivingXiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li +12
Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for de
vision-language-actionvlavla model - arxiv:2609.34385 · cs.LGJust-In-Time Agent Memory with Runtime Agentic ResearchBingyu Yan, Chaofan Li, Hongjin Qian, Shuqi Lu +2
Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes
memorylong-contextagent memoryagentai agentagentic - arxiv:2609.34384 · cs.RORoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action ModelsAernaer Akelijiang, Jiannan Li, Zhineng Chen, Jingjing Chen +1
Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicit
vision-language-actionmanipulationfrankamemorybenchmark - arxiv:2609.34380 · cs.AIDPS: Dual-Mode Precision LLM Serving with Semi-Unified MemoryXuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung +1
Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this design by allowing a single stored model to support both full-accuracy and lower-precision execution
memory - arxiv:2609.34379 · eess.SYDerivative-Free Generalized Multivariable Super-Twisting Control for Constrained Euler-Lagrange SystemsChidre Shravista Kashyap, Jishnu Keshavan
Robotic manipulators performing payload lift-and-transfer must track prescribed trajectories under abrupt load changes, with limited actuation and inaccurate plant knowledge. Such plants are inherently described by multivariable Euler--Lagrange~(EL) dynamics with cross-coupling established through i
manipulator - arxiv:2609.34378 · cs.CVMarathoner: Ultra-Long-Horizon Autonomous IntelligenceZhang Ruiyang, Ou Jinpeng, Xie Yifan, Zhou Jingang +3
Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long
autonomous agentagenticpost-trainingbenchmark - arxiv:2609.34375 · cs.ROLRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World ModelsLuzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao
Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appear
world modelv-jepaaction-conditioned - arxiv:2609.34374 · cs.LGABC-Align: Prediction-Powered Alignment with Adaptive Bias ControlEric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff
Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases th
post-training - arxiv:2609.34373 · cs.ROEmergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent CommunicationMihir Chauhan, Aniket Bera
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited
multi-agentself-playarena - arxiv:2609.34372 · cs.AIPersMem: Internalizing Personality into Dual-Pathway Memory for LLM AgentsHanzhong Zhang, Ziwei Xiang, Weicheng Xie, Shizhe Liu +1
The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory proc
memoryagentllm agent - arxiv:2609.34370 · cs.LGSpexis: Speculative Lookahead Scheduling for LLM InferenceHyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim +2
Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new para
memory - arxiv:2609.34367 · cs.CVRate-Distortion Adaptive Primitive Selection for Omnidirectional Gaussian SplattingYulong Cheng, Youneng Bao, Junfeng Zhou, Mu Li +1
Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR). Gaussian splatting (GS) codecs render much faster, yet still lag in reconstruction quality and typically decide primitive allocation without considering
benchmark - arxiv:2609.34366 · cs.AIWhen Harness Beats Scale, and When Reading Beats BothIvan Bondarenko, Nikolay O. Nikitin
We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interp
knowledge graphleaderboard - arxiv:2609.34363 · cs.CVSyncRA: Learning Temporal Correspondence in Omni-Modal ModelsZelong Xu, Yan Li, Wenhe Hu, Xiyang Hu
Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible
benchmark - arxiv:2609.34362 · cs.ROFutureDuet: Decoupling Observation Access from Future Supervision in World Action ModelsJie Wu, Yuzhi Huang, Junqi Liu, Weichen Zhang +4
World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-l
liberorobotwin - arxiv:2609.34359 · cs.AIImproving Large Language Models for Code through Runtime Program-State ReasoningHongwei Li, Spandan Garg, Yufan Huang
Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-
agentpost-training - arxiv:2609.34358 · cs.LGFORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM AgentsXi Xiao, Yunbei Zhang, Chen Liu, Lin Zhao +6
In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answ
llm agentagenticbenchmark - arxiv:2609.34353 · cs.AISemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative PerceptionHu Xu, Chun Li, Siyuan Qiu, Zeyan Li +1
Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open:
agent - arxiv:2609.34349 · cs.AIFuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality PreservationMichael Vasilakakis, Dimitris K. Iakovidis
Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic an
benchmark - arxiv:2609.34348 · cs.LGCommutator Memory: Sparse, Path-Local Reading and Steering in Language ModelsJohn Sweeney
Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leav
memorybenchmark - arxiv:2609.34347 · cs.AISAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-FormerZhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou +2
Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features an
embodiedembodied agent - arxiv:2609.34346 · cs.CVE-WAVE: Event-based Continuous Optical Flow via Warping-Aligned Visual EncodingJiale Wu, Xiaoyang Bai, Haoming Yu, Yiwei Chen +2
Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and comp
memoryevent camera - arxiv:2609.34345 · cs.CLCRISP: Cultural Reward Modeling for Implicit Situated ProprietyZekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao +5
As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined respon
multi-agentagent framework - arxiv:2609.34344 · cs.LGLearning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable VectorsYuchen Cai, Ding Cao, Qixiang Yin, Xin Xu +9
Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify
post-training - arxiv:2609.34342 · cs.AISAGE: Structured Strategic Reasoning for Efficient LLM Game PlayingZhiwei Chen, Tianchun Wang, Zhongtao Rao, Haiming Zhu +2
A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents' behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupporte
agentllm agent - arxiv:2609.34335 · cs.CVSkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt EngineeringYanwei Huang, Mingxuan Zhu, Shujie Li, Shiyuan Liu +2
Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. S
benchmark - arxiv:2609.34330 · cs.CVMiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM InferenceTinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang +10
Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristi
benchmark - arxiv:2609.34327 · cs.AIKnowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric KnowledgeChanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo +2
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model f
self-refinementbenchmark - arxiv:2609.34326 · cs.LGRouting Without Embeddings: Fast And Interpretable Routing With Regular ExpressionsYifan Lu, Qiyue Zhang, Haotian Shan, Hanjie Chen +1
Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders
benchmark - arxiv:2609.34325 · cs.CVDORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision TransformersKaixuan He, Song Chen, Yi Kang
Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose
agent - arxiv:2609.34322 · cs.AITest-Time Scaling via Budgeted Multi-Attribute VerificationBo Xue, Ji Cheng, Shen-Huan Lyu, Yuanyu Wan +1
Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated alon
benchmark - arxiv:2609.34321 · cs.LGOne Rollout Is All You Get: Fully Test-Time Adaptation for GUI AgentsZiqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang +5
GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployme
memoryagent - arxiv:2609.34320 · cs.CLCertified Selective Automation of LLM Agent EvaluationChengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni
Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectorie
agentllm agenttool-use - arxiv:2609.34319 · cs.ROText-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action InferenceQianer Li, Chengjie Zhang, Jingwen Chen, Zanjia Tong +2
Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy,
vision-language-actionvlavla modelopenvlabenchmark - arxiv:2609.34314 · cs.CVPlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?Shayekh Bin Islam, Hwanjun Song
Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Th
agenticbenchmarkjudge model - arxiv:2609.34313 · cs.AIControlScope: Workflow Revision and Reliability in LLM AgentsJingjie Ning, Xueqi Li, Yibo Kong, Dongting Li
How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the ac
agentllm agent - arxiv:2609.34309 · cs.CVMaLiang-Harness: A Programmable Path to Image and Video GenerationHaoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu +3
Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to
benchmark - arxiv:2609.34300 · cs.ROWhen World Models Lie: Adaptive Safety Analysis Under Wrong ImaginationsJohn Cao, Somil Bansal
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safe
world model - arxiv:2609.34299 · cs.CVPSM: Dataset Distillation Based on Precise Statistical Matching by DifficultyHongxu Ma, Guang Li, Shijie Wang, Dongzhan Zhou +5
Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distille
memory - arxiv:2609.34297 · cs.ROTLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot ManipulationXianbo Cai, Hideyuki Ichiwara, Zihang Wang, Yijun Lu +1
Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing ex
manipulationlibero - arxiv:2609.34296 · cs.AIDr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research AgentsYingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding +5
Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on
agentbenchmark - arxiv:2609.34294 · cs.CVSemantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired SettingsDuanning Chen, Ke He, Bin Yang, Yongxiang Yao
Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without
memory - arxiv:2609.34287 · cs.LGReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM PretrainingZichun Yu, Jiarui Yan, Shlok Sanghvi, Nihar Atri +1
LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a
multi-agent - arxiv:2609.34286 · cs.RODexterous Tactile World ModelZiyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan +5
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video wo
manipulationdexteroustactileworld model - arxiv:2609.34284 · cs.CLOver-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMsHaeun Jang, Yonghyun Jun, Hwanhee Lee
Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot
benchmark - arxiv:2609.34281 · cs.LGAgentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided SearchZhixuan Gao, Ke Xue, Rongxi Tan, Ming Chen +1
High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on
agentagenticbenchmark - arxiv:2609.34279 · cs.LGDirect Self-Evolving Optimization: Evolving LLMs without Challenger TrainingYuyang Deng, Yu Wang, Jiayun Wang
Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}ire
self-evolving - arxiv:2609.34277 · cs.CVSee, Measure, and Reason: Learning Visually Grounded Reasoning in PathologyChengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li +7
Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that
benchmark - arxiv:2609.34276 · cs.RONavHarness: Towards Lifelong Embodied NavigationXunyi Zhao, Jian Zhou, Sihao Lin, Gengze Zhou +5
Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, an agent must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new obser
embodiedmemoryagentagentic - arxiv:2609.34274 · cs.AIBIABench: Evaluating AI agents on real-world bioimage analysis tasksZixuan Pan, Davide Panzeri, Lukas Johanns, Marilin Moor +4
Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so
agentai agentbenchmark - arxiv:2609.34271 · cs.CVScaling Versatile 3D Assets Editing with a Million-Scale DatasetBadi Li, Tianxin Huang, Yu Zhou, Wei-Shi Zheng +2
Although recent 3D generative models produce increasingly realistic assets, controllable 3D asset editing remains challenging. Existing methods are limited by scarce training data, insufficient source-aware modeling, and a lack of practical evaluation protocols. To address these limitations, we pres
benchmarkevaluation protocol - arxiv:2609.34268 · cs.ROSAGE: Symbolic Action-Gating and Editing for LLM Task PlannersTrung Minh Bui, JongSul Moon, YoungOuk Kim, Quang-Ngoc Phung +2
Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be s
embodiedmemorybenchmark - arxiv:2609.34262 · cs.LGMaintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned PassesWeijun Luo, Kelvin Luu, Xinyi Liu, Guangze Luo +5
Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless ca
agentagenticbenchmark - arxiv:2609.34261 · cs.RORoboICL: Embodied In-Context Learning with GPT-6 AstraFangcheng Liu, Yeqing Shen, Anda Cheng, Weishi Mi +6
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-contro
embodiedmanipulationmemoryleaderboard - arxiv:2609.34258 · cs.AIInvestigating Human--AI Discrepancies via Multiple-Solution ProblemsZihao Wang, Francesco Insulla, Andrea Montanari
Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and m
benchmark - arxiv:2609.34256 · cs.ROUMR: Universal Manipulation RepresentationSong Liu, Linyi Li, Yanshun Zhao, Rxuan Li +10
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new
vlaembodiedmanipulationliberobenchmark - arxiv:2609.34250 · cs.ROWAM-OPD: Sharpening World Action Models via On-Policy DistillationPanjun Liu, Xiaohan Lei, Shiqi Zhang, Yikun Wang +7
Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. W
manipulation - arxiv:2609.34246 · cs.LGCasEm: A Cascade Architecture for Long-Horizon Neural EmulationZhaoyi Li, Jingtao Ding, Shihua Li
Autoregressive neural emulators can drift or diverge over long rollouts despite accurate short-term predictions. We introduce Cascaded Emulation (CasEm), a one-way rollout architecture that augments an existing full-state backbone with an independently evolving model of physically specified aggregat
benchmark - arxiv:2609.34247 · cs.CLSALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice AgentsWenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen +4
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conve
agentllm agent - arxiv:2609.34245 · cs.LGCertified Multi-Source Integrity for Structured Agent ActionsAnmol Pandey, Aditya Jain, Liang Chen, Carsten Maple +1
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-c
agentllm agent - arxiv:2609.34242 · cs.LGStashbird: Efficient Speaker-Indexed Memory for Conversational AgentsChidera Biringa, Lucas Yannul, Xiaowen Wang, Marco Ayala +4
AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episode
memoryagent memoryagentai agentbenchmark - arxiv:2609.34241 · cs.AIAdaGuard: An Adaptive Guard Model with User-defined PoliciesYunhao Feng, Yifan Ding, Yuxiang Xie, Zheng Li +3
Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's beha
agent - arxiv:2609.34237 · cs.CVDecFlowEdit: Self-Localized Flow-based Image Editing via Guidance DecouplingZheyuan Zhan, Can Wang, Jiawei Chen, Chun Chen +3
Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial backgr
manipulation - arxiv:2609.34235 · cs.CVSegBanana: Steering Unified Multimodal Models into Medical SegmentersXiaoye Liang, Ye Yan, Mingze Yin, Shikun Feng +4
Medical image segmentation remains challenging in practical deployment, as models often struggle to generalize beyond the distributions covered by their training data and high-quality pixel-level annotations are typically unavailable for adaptation. Inspired by the cross-task transferability of larg
agenticpost-training - arxiv:2609.34234 · cs.CLMAS-OPD: On-Policy Distillation for Multi-agent SystemsQiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang +4
Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competenc
agentmulti-agentagent systempost-trainingbenchmark - arxiv:2609.34233 · cs.ROGAE: General Action Expert for Real-Time Humanoid TeleoperationYuefan Wang, Huaicheng Zhou, Xiao He, Zhijie He +5
Humanoid avatars extend human physical presence beyond the body, enabling people to participate in social, service, and labor activities through remotely operated robots. This requires teleoperation systems capable of realizing diverse and dynamic whole-body behaviors while maintaining responsive hu
humanoidteleoperation - arxiv:2609.34232 · cs.CVTrustworthy synthetic visual media: Evidence across the media lifecycleZexi Jia, Zhiqiang Yuan, Jie Zhou, Jinchao Zhang
Images and videos have long helped people understand what happened and how a work came into being. Generative systems complicate that role. Realistic media can now be produced and revised without leaving a stable history, so appearance no longer reveals whether a scene was captured, synthesized, or
manipulation - arxiv:2609.34231 · cs.CVReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel GeometryWangzhi Zhan, Jianpeng Chen, Dongqi Fu, Dawei Zhou
Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porou
benchmark - arxiv:2609.34228 · cs.LGSleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden SignalsJingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez, Joshua Shunk +2
Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addres
agentbenchmark - arxiv:2609.34227 · cs.AIWhen Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision ModelRishabh Sharma, Rishika Lall
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran
memoryagent memoryagent - arxiv:2609.34225 · cs.CLUSA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language AgentsQiyong Zhong, Mao Zheng, Mingyang Song, Huwei Ji +5
On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them,
agentic - arxiv:2609.34223 · cs.CVUncovering Ordinal-Matching Bias in Audio-Visual LLMsJihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo +1
This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic da
benchmark - arxiv:2609.34222 · cs.ROProprioceptive Force Estimation for Quadruped Locomotion and Human-Robot InteractionRun Wang, Xu Yang, Alapati Tuerxun, Yilin Mo
Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy
quadruped - arxiv:2609.34221 · cs.CVWorldWeave: Growing Persistent Geometric Worlds for Video GenerationYifan Huang, Lifan Jiang, Qingyue Hao, Cheng Chen +4
Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples
world modelmemoryagent - arxiv:2609.34220 · cs.ROmmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave RadarJunqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu +6
Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures
vision-language-actionmanipulation - arxiv:2609.34215 · cs.LGSame Winners, Different Success Rates: Evaluating How LLM Agents Recover from FailuresDong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang +4
Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purel
llm agentagenticbenchmark - arxiv:2609.34214 · cs.AIGlyphBench: A Playground for Language-Model Reinforcement LearningRoger Creus Castanyer, Marc-Alexandre Côté, Matthew James Sargent, Augustine N. Mavor-Parker +2
We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through
agentpost-training - arxiv:2609.34211 · cs.AIBehavior-Grounded Semantic Enrichment for Financial Fraud Modeling and ReasoningLinbo Shao, Huilin He, Yating Lou, Dawei Cheng
In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semanti
multi-agent - arxiv:2609.34210 · cs.RORLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning EngineersHaitong Ma, Chenxiao Gao, Rushi Qiang, Na Li +1
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agent
agentbenchmark - arxiv:2609.34207 · cs.LGFunctional Autoencoders for Amplitude-Phase Representation LearningPeida Wu, Xinyang Xiong, Pengcheng Zeng
Functional data are intrinsically infinite-dimensional, and often exhibit phase variation, where corresponding events occur at different times across observations. Existing linear dimension reduction methods struggle with nonlinear amplitude variation, while functional autoencoders without an explic
benchmark - arxiv:2609.34206 · cs.CVWorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action PoliciesLin Liu, Lu Zhang, Ziying Song, Wu Yang +5
Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed i
vision-language-actionvlaliberoworld model - arxiv:2609.34205 · cs.LGLearning to Optimize through Solver-Grounded Self-PlayXia Jiang, Yaoxin Wu, Chenyu Zhou, Mengzhu Xu +2
Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This d
self-play - arxiv:2609.34199 · cs.ROWB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-ManipulationChuan Qin, Shaoting Zhu, Siyuan Luo, Siqiao Huang +2
Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A
manipulationdexteroushumanoid - arxiv:2609.34198 · cs.LGFrozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent EvaluationJiapeng Li
Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 ex
agent - arxiv:2609.34196 · cs.CVConvCue: Complementary Visual Inductive Biases for Vision-Language ModelsZixuan Lan, Shichu Sun
Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations
benchmark - arxiv:2609.34195 · cs.AIPainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language ModelsShane K. A. Dalumura Hettige, Jonas Oppenlaender
Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas t
agentagentictool usebenchmark - arxiv:2609.34190 · cs.CVMotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion SpaceQing Yu, Kent Fujiwara
Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct ma
manipulation - arxiv:2609.34188 · cs.LGAlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement LearningYingbo Zhao, Zeyu Yang, Zhoufan Zhu
Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha p
agent - arxiv:2609.34185 · cs.LGEntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary BitratesHong Zhang, Zhongjie Duan, Yingda Chen
Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility req
memory - arxiv:2609.34184 · cs.AICASS: Contribution-Aware Structured Sparsity for Model MergingYan Li, Guiping Cao, Meng Xu, Tao Jiang +4
Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, trea
benchmark - arxiv:2609.34182 · cs.ROUnified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous ManipulationWenqiao Li, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu +4
Dexterous manipulation requires tactile feedback.However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile intera
manipulationdexterousteleoperationtactile - arxiv:2609.34181 · cs.AIEfficient Reasoning via Constrained Optimization in Latent SpaceZhinan Hou, XingChen Li, Keyou You
Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt
benchmark - arxiv:2609.34179 · cs.LGRAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk ConstraintsRichard Krueger, Lucas Krause, Zach Pocquette
Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a
retrieval-augmentedragbenchmarkevaluatorleaderboard - arxiv:2609.34178 · cs.CVEnhanced Video Text Editing with Trajectory-Aligned Glyph RenderingShulian Zhang, Xiangyu Shu, Wenbo Li, Jian Chen +1
Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stro
benchmark - arxiv:2609.34177 · cs.LGReplayLens: Auditing Agents' Use of OutcomesDong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang +4
When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship
memoryagent - arxiv:2609.34176 · cs.ROAGILE-GS: Anchor-Guided Fast Next-Best-View Selection for Active 3D Gaussian SplattingAmirhossein Mollaei Khass, Nader Motee
Radiance fields need hundreds of views, and their placement matters as much as their number. Next-best-view (NBV) selection for 3D Gaussian Splatting (3DGS) usually scores every candidate in the pool and keeps one. Searching for information and choosing a camera, however, are separable problems. We
embodiedbenchmark - arxiv:2609.34175 · cs.ROFailPatch: Failure Residual Patching for Vision-Language-Action ModelsPeng Yu, Jiacheng Wang, Ziheng Zhang, Xuchong Zhang +5
Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose
vision-language-actionvlavla policyrobotwin - arxiv:2609.34171 · physics.opticsMonolithically integrated photonic neural network at driven-dissipative criticalityYuming Zhang, Ruoran Wang, Jingcheng Li, Hailong Zhou +3
Artificial intelligence increasingly demands computing architectures capable of adaptive representation, temporal information processing, and robust operation under uncertainty. However, conventional photonic neural architectures typically rely on predefined computational operations and separate fun
silicon photonic - arxiv:2609.34170 · cs.RORAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action ModelsYuhan Chen, Ke Yu, Pengfei Liu, Shuxun Wang +2
Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult
vision-language-actionvlamanipulation - arxiv:2609.34167 · cs.CVNatural Image Autoencoder-Based fMRI Representations for Trait and State PredictionJuhyeon Park, Yeonwoo Kim, Peter Yongho Kim, Yansen Wang +4
Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a
benchmark - arxiv:2609.34163 · cs.ROReliability-Aware Sparse Route Memory for Round-Trip Vision-Language NavigationBojun Long, Lingfan Bao, Tianhu Peng, Jingcheng Sun +1
Vision-language navigation (VLN) is typically evaluated as a one-way task, although deployed robots may need to return after reaching a goal. We study continuous round-trip VLN and diagnose failures in directional observability, deviation recovery, and termination stability. We propose a reliability
memory - arxiv:2609.34160 · cs.AIRoutePrism: Tracing Construction Order Effects in Agent MemoryDong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang +4
Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces whic
memoryagent memoryagent - arxiv:2609.34159 · cs.LGWorldGraph: Graph-Native World ModelingZezhong Ding, Yipeng Li, Xike Xie
World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their properties change over time. Prior graph-related
world modelbenchmark - arxiv:2609.34157 · cs.AITableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table CorporaJiaming Tian, Liyao Li, Wentao Ye, Haobo Wang +4
Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depe
agentllm agentagenticbenchmark - arxiv:2609.34151 · cs.AISelf-Evolving Agents via Likelihood-Guided Tool-Space OptimizationXuanqi Zhang, Ruinan Jin, Running Yang, Yuxuan Zhang +3
Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-sp
tool-useself-evolvingbenchmark - arxiv:2609.34149 · cs.CVFunctional Hand Type Prior for 3D Hand Pose Estimation and Action Recognition from Egocentric View Monocular VideosWonseok Roh, Seung Hyun Lee, Won Jeong Ryoo, Jakyung Lee +4
Current methods for egocentric view action recognition often face challenges in perceiving dynamic hand movements relying solely on geometrical or physical information. In this work, we effectively address this problem by gaining insights into the correlation between functional hand configurations a
benchmark - arxiv:2609.34145 · cs.ROBeyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language DrivingJiaxin Liu, Ruilin Yu, Liang Peng, Jingkai Wang +7
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, an
retrieval-augmentedknowledge graph - arxiv:2609.34144 · cs.CVCAST: Reconstruction-Coupled Acceleration of Interactive World ModelsLeyang Chen, Junyi Wu, Fanqing Kong, Shaoqiu Zhang +1
Interactive world models must respond quickly to controls while preserving scene consistency. Existing acceleration methods can miss heterogeneous control responses and spatial transport when recovering skipped features. We observe that interaction-induced feature changes correlate with approximatio
world model - arxiv:2609.34143 · cs.CVBeyond Geometry: Benchmarking and Consistency Reasoning for 3D Logical Anomaly DetectionZhiqiang Qin, He Xie, Junfei Yi, Yang Yang +4
Existing 3D industrial anomaly detection mainly targets local geometric deviations. In contrast, many industrial anomalies violate object-level design or assembly rules, which we define as 3D logical anomalies. To address these challenges, we introduce the Industrial Logical Anomaly Detection Datase
benchmark - arxiv:2609.34139 · cs.AISame Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?Tien Tran, Namho Koh, Daiki E. Matsunaga, Ayush Jain +1
Mobile GUI agents deployed in real settings must work across different applications that support the same functionality. Most existing benchmarks test each task in only one app, so a high score can mean the agent understands the task, or only that it knows that particular app. We introduce AnyAppBen
agentbenchmarkleaderboard - arxiv:2609.34136 · cs.AIWaggle: Learning One Anonymous Local Law for Self-Organizing LLM SwarmsMingxi Zou, Wei Zhu, Zhuo Wang, Langzhang Liang +3
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable
llm agentmulti-agentagent system - arxiv:2609.34135 · cs.AIEvo2Team: When Do Evolved Skills Transfer? From Selection to DeploymentRenxiang Wang, Jiaming Cui
A skill bank that helps one multi-agent system may leave another's behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4--32 agents and GPT and Qwen model
multi-agentagent system - arxiv:2609.34134 · cs.AIStateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data AgentsWenle Liao, Zhao Wang, Jingchao Zhang, Jiajie Jin +2
LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, makin
memorybenchmark - arxiv:2609.34133 · cs.CVPrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise PreferencesChuanzhi Xu, Langyi Chen, Chengkun Yue, Xuanhua Yin +4
Photographic color editing is inherently personal: the same image can appear too warm, too muted, or already satisfactory to different users. Most lookup table (LUT) and reference-guided methods target a specified appearance rather than model persistent preferences from repeated user choices. To add
evaluation protocol - arxiv:2609.34132 · cs.AIFrom Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM AgentsMingxi Zou, Langzhang Liang, Zhuo Wang, Yiyang Zhao +2
As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typic
memorypersistent statepersistent memoryagent memoryagentllm agent - arxiv:2609.34127 · cs.AISTITCH-RAG: Spatio-Temporal Influence Tracing over Topic Hypergraphs for Multi-Hop Retrieval-Augmented GenerationHaodong Yang, Mengzhu Chen, Jia Cai
Multi-hop retrieval-augmented generation requires a retriever to connect evidence distributed across documents while preserving a concise, faithful generation context. Existing indexes leave two complementary gaps: chunk-based RAG can break cross-passage evidence chains, whereas an unlabeled pairwis
retrieval-augmentedragbenchmark - arxiv:2609.34126 · cs.AIJET: Judge-Guided Evolution at Test Time for Agent ProgramsYao Long Teng, Jiayi Cai, Bo An
An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior,
agentevaluator - arxiv:2609.34125 · cs.CLUnderstanding Clinical Cognitive Dialogues Using Large Language ModelsVishalakshi Arumugam, Dan Schumacher, Veronica Rammouz, Enrique Gonzalez Guerrero +2
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study
benchmarkevaluation framework - arxiv:2609.34124 · cs.CVSpatialSkill: Self-Evolving Skills for Cross-View Spatial ReasoningRuifan Zuo, Guocheng Hu, Wanshui Gan, Junyi Wang +2
Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which k
self-evolving - arxiv:2609.34123 · cs.LGWhat Does a Stream Model Buy You in Flow Matching?Jian Xu
Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We ask what such a stream model actually cont
benchmark - arxiv:2609.34117 · cs.LGSlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE ServingGunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin +2
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality
benchmark - arxiv:2609.34115 · cs.LGForecast-Necessary Causal Discovery for Nonlinear Political Panel Data: Feedback, Functional Form, and the Dynamics of DemocratizationMichael Coppedge, Dmitry Zaytsev, Valentina Kuskova
A non-significant coefficient in a dynamic panel model need not imply the absence of a relationship. It may instead reflect heterogeneous effects averaged toward zero, reciprocal dynamics overlooked by a recursive specification, or relationships masked by the omission of correlated covariates. Stand
benchmark - arxiv:2609.34112 · cs.AIUnknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scoresNicolás Vera Zúñiga
Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LL
agentllm agent - arxiv:2609.34105 · physics.opticsStimulated Electro-optic ScatteringViolet Workman, Gaurav Bahl
Stimulated Brillouin Scattering (SBS) couples light to acoustic waves and underpins applications ranging from sensing and signal processing to quantum photonics. In solids, this interaction is generally attributed to photoelasticity and the motion of dielectric boundaries. Here we show that piezoele
quantum photonic - arxiv:2609.34098 · cs.LGBeyond Correctness: Evaluating Semantic Knowledge in Cross-Table TransferSeokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim
Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We
evaluation framework - arxiv:2609.34095 · physics.opticsTensor-Engineered Van der Waals NbOCl2 Resonant Metasurface for Polarization-entangled Bell State GenerationXin Zeng, Wenna Du, Yun-Kun Wu, Yuyang Zhang +19
Polarization-entangled photon pairs are essential resources for quantum information technologies, yet realizing compact sources with intrinsically controllable entanglement remains challenging. Van der Waals (vdW) nonlinear materials such as NbOCl2 provide atomically thin platforms for quantum light
manipulationquantum photonic - arxiv:2609.34088 · cs.LGTRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac CareLovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold +10
TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-style training, which often struggles with
benchmark - arxiv:2609.34085 · cs.ROAD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous DrivingHaoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska
Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM
world modelaction-conditionedbenchmark - arxiv:2609.34083 · cs.LGBeyond One Epoch: Uncertainty-Weighted Sensitivity Regularization for Recommendation ModelsRichard Lettich, Shagun Gupta
Recommendation models with sparse embeddings and a shared consumer often exhibit the one-epoch phenomenon: a second epoch lowers training loss while sharply degrading generalization. We present a view based on the violation of the prequential principle. On the first epoch, an example's label has not
benchmark - arxiv:2609.34082 · cs.CVK-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC DrawingsYunfei Bai, Enrico Chionna, Akash Amol, Kawaljit Singh KC +1
Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillatio
self-improvingpost-training - arxiv:2609.34079 · cs.AIGenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent ComputationTanmoy Kanti Halder, Akash Ghosh, Arijit Roy, Sriparna Saha
Large language models (LLMs) have demonstrated strong capabilities in biological reasoning; however, genomic disease inference remains largely dependent on memorized gene-disease associations rather than understanding biological pathways. This shortcut learning undermines robustness and generalizati
benchmark - arxiv:2609.34078 · cs.CVWhiteCon: Semi-Supervised Domain Adaptation Regression Through Whitening Transform and Dual ConsistencySe Jin Sim, Seoung Bum Kim
Domain adaptation is crucial for addressing distributional shifts that degrade model performance across domains. While most existing research has centered on classification, semi-supervised domain adaptation regression (SSDAR) for continuous-output tasks remains largely unexplored, particularly in p
benchmark - arxiv:2609.34077 · cs.LGMaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE InferenceJunfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui +2
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing all
memorybenchmark - arxiv:2609.34072 · cs.AIPhysFieldBench: Can Multimodal Models Understand Physical Fields?Yuezhou Ma, Huikun Weng, Jialong Wu, Chenyi Zhao +4
Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoni
post-trainingbenchmark - arxiv:2609.34069 · cs.AITowards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program OptimizationPiyush Jha, Aishik Ghosh, Vijay Ganesh
The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. W
agenticself-improving - arxiv:2609.34064 · cs.LGLearning Perturbation Robust Policies for LLM Agents with Stable OptimizationPengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang +1
Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study
llm agentpost-training - arxiv:2609.34061 · cs.ROQuantile Head for Vision-Language-Action ModelsXuan Wang, Yinan Wu, Haoran Duan, Jungong Han
Vision-Language-Action (VLA) models integrate pretrained Vision-Language Models (VLMs) with action heads for robot control. Common action heads have distinct limitations: point regression provides only a point estimate of the action distribution, while standard flow-matching samplers require costly
vision-language-actionaction headlibero - arxiv:2609.34060 · cs.LGKVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent SystemsHyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang +1
Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and con
memoryagentmulti-agentagent system - arxiv:2609.34058 · cs.LGDo World Models Learn Global Understanding?Alexander Detkov, Matt Thomson
AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundame
embodiedworld model - arxiv:2609.34054 · cs.LGPReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral ReconstructionHyesung Jeon, Hyeongju Ha, Jae-Joon Kim
Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache
memoryagentagent systemagent benchmarkbenchmark - arxiv:2609.34049 · cs.AIThinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV InferenceTimothy DeLise, Seth Cromelin
Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens
memory - arxiv:2609.34047 · cs.CVARCH-B: Architectural Representation, Comprehension and Hierarchy BenchmarkKieran Sagar Parikh, Jose Luis Garcia del Castillo y Lopez
Multimodal models increasingly interpret visual environments, but their ability to recognize the same building across photographs, floor plans, elevations, sections, and renderings remains poorly characterized. We introduce ARCH-B, a benchmark of 354 four-choice questions across 11 cross-representat
benchmark - arxiv:2609.34045 · cs.LGKafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity MachinesMurtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya
Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding mem
memory - arxiv:2609.34044 · cs.CVSCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language ModelsAhmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer +11
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual i
benchmark - arxiv:2609.34041 · cs.LGVision--Language Signals in Constrained RL: Safety Gains Without AnticipationSamuel Tetteh, Cody Fleming
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can pro
benchmark - arxiv:2609.34039 · cs.AILarge Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and ValidationErfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten +2
Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL
agentagent frameworkbenchmark - arxiv:2609.34036 · cs.LGUOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn AgentsWenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang +1
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncert
agentic - arxiv:2609.34035 · cs.RO3D Point Tracking with State Space ModelsMasahiro Ogawa, Qi An, Atsushi Yamashita
Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and
memorybenchmarkevaluator - arxiv:2609.34034 · cs.LGADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence ModellingMatei-Ioan Stan, Oliver Rhodes
A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realisti
memory - arxiv:2609.34024 · cs.LGJev in Medicine: A Benchmark Evaluation. Preliminary ResultsAlfredo Madrid-García, Beatriz Merino-Barbancho
Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaM
benchmark - arxiv:2609.34019 · cs.LGSR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-MakingShyam Sundar Murali Krishnan, Dean Frederick Hougen
In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alter
benchmark - arxiv:2609.34018 · cs.ROEstimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor ControlDenis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz, Paul Mattes +2
Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is train
manipulationsim-to-real - arxiv:2609.34017 · cs.AIMaat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM WorkflowsUliana Elina
Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose verdicts are themsel
agentmulti-agentagent systembenchmark - arxiv:2609.34010 · cs.ROZeroBot: Learning from Scratch in Minutes with Generative Real2SimIvan Kapelyukh, Xiaohan Zhang, Stephen James, Laura Herlant +1
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses
manipulationgrasp - arxiv:2609.34006 · cs.ROTacGooseBumps (TacGB): Retrofitting Normal-Only Tactile Sensors with Shear Encoding for Learning Contact-Rich ManipulationWenjie Li, Binyu Yang, Yuxin Chen, Ambrose Wang +1
Contact-rich policies often fail because distinct physical states look alike yet require different actions. Cameras may not reveal whether a connector is aligned or fully seated, while many normal-only tactile sensors can miss the tangential interactions perpendicular to the grasping direction that
manipulationtactilegrasp - arxiv:2609.34004 · cs.LGRICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock ForecastingTong Liu, Lanmiao Liu, Xiang Hu
Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, information availability, and transition reliability are modeled. Existing LLM-based financial agents incorporate historical evidence, yet they provide limited
memoryagent - arxiv:2609.33998 · cs.CVMetaSampling: Making Frame Samplers Efficient for Long-Video Question AnsweringAshim Dahal, Bikramjit Banerjee
Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We i
benchmark - arxiv:2609.33994 · cs.AIA packet-level digital hardware twin for commissioning megahertz diagnostic edge AI and plasma control system integration in tokamaksSemin Joung, Abhilasha Dave, Luca Scomparin, Filipp Khabanov +5
High-bandwidth plasma diagnostics increasingly provide inputs to machine learning and signal-processing algorithms intended for real-time tokamak control, but the complete path from diagnostic sampling to control-system handoff is difficult to commission because of their sampling rates. We develop a
memory - arxiv:2609.33993 · cs.LGASTRA: ADMM-Accelerated Topology Reconfiguration for Dynamic Satellite ConstellationsJoão Norberto, Ricardo Ferreira, Cláudia Soares
Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Satellite Topology via Regret-Aware learni
online learning - arxiv:2609.33991 · cs.CVA Multi-Dataset Benchmark of YOLO-Based Weed Detection in Precision AgricultureHristina Zdraveska, Vlatko Spasev, Ivica Dimitrovski, Ivan Kitanovski +1
Weed detection is an important component of precision agriculture, enabling site-specific weed management and reducing unnecessary herbicide use. Although deep learning methods have achieved strong results for crop and weed detection, many studies rely on single-dataset evaluation, making it difficu
benchmark - arxiv:2609.33989 · cs.LGRewardExplainer: Learning Reward Model Explanations from Counterfactual Preference FeedbackJingyi He, Nier Wu, Shuang Liu, Xin Wang +2
Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their sc
post-trainingbenchmark - arxiv:2609.33987 · cs.CLOpera: A Verbal Critic Framework for Long-horizon Coding AgentsKai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan +6
Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is de
benchmark - arxiv:2609.33986 · cs.LGICMAPE: In-Context Multiagent Pure ExplorationXinyi Hu, Alessio Russo, Aldo Pacchiano
In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent
multi-agentagent systembenchmark - arxiv:2609.33984 · cs.LGFrom HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series ForecastingChaoqi Zhang, Yu Wang, Haixu Tang
Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingul
benchmark - arxiv:2609.33982 · cs.ROTest-Time Spatial Reasoning for Robot Manipulation Using Generative Real-to-SimIvan Kapelyukh, Yafei Hu, Ran Gong, Brandon May +5
Spatial reasoning is fundamental to general robot intelligence, as it enables robots to complete long-horizon tasks involving multi-object interaction. We introduce Simify, a training-free, test-time framework that performs explicit spatial reasoning via massively parallel physics simulation. From a
manipulationsim-to-real - arxiv:2609.33981 · cs.LGFuture Information-Directed Sampling for Bayesian Nonstationary BanditsYichen Song, Alessio Russo, Aldo Pacchiano
Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal
benchmark - arxiv:2609.33980 · cs.LGDynGraphAgentBench: A Benchmark for Agentic Lifecycle Control in Dynamic Graph Anomaly DetectionYuwei Han, Lingwei Wei, Wooseong Yang, Liangjie Huang +3
Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedb
agenticbenchmark - arxiv:2609.33974 · cs.CLBeyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive PruningRuosong Ye, Caiqi Zhang, Jiahao Li, Haijun Wu +8
Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. However, existing MAD frameworks fail to be
agentmulti-agentbenchmark - arxiv:2609.33973 · cs.ROFINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube SolvingYutong Liang, Quanquan Peng, Matthew Kim, Xiaolong Wang
Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and re
dexterousgrasp - arxiv:2609.33969 · cs.CVGaussian Splatting-based Volumetric Video Compression with Sparse 4D AnchorsGe Gao, Siyue Teng, Chanqgi Wang, Fan Zhang +4
Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formu
memory - arxiv:2609.33967 · cs.LGThinkNet: Compact Architecture Selection and Validation-Gated Ensembles for Subject-Independent MI-EEG DecodingAbdul Basit, Saim Rehman, Muhammad Shafique
Practical assistive and rehabilitative brain--computer interfaces require subject-independent motor-imagery EEG (MI-EEG) decoders that generalize to new users under limited target-user data and constrained compute. However, held-out-subject performance can be overstated when test-subject information
benchmark - arxiv:2609.33965 · cs.LGGreenpixie's AI Token Methodology: Assessing the Energy, Water and $\mathrm{CO_2\text{-}eq}$ Impact of AI Tokens for Open and Closed Weight ModelsJoshua Horswill, Ross Hunter, Matt Clifford, James Hall
We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a
embodiedbenchmark - arxiv:2609.33955 · cs.AIDesigning Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business AgentsKaiwen Luo, Ming Gao
Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to ei
agenthuman-in-the-loop - arxiv:2609.33947 · cs.LGSimple Diffusion Language Models Are More Effective Few-Step Generators Than ReportedHasan Amin, Ming Yin, Rajiv Khanna
Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We sho
benchmark - arxiv:2609.33944 · cs.ROReSync: Re-Aligning the Two Clocks of Asynchronous World-Action ModelsXi Lin, Feihong Zhang, Yulong Shi, Yanghong Mei +5
Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. Th
benchmark - arxiv:2609.33940 · cs.LGBehavioral Monitoring of JEPA World Models with Jacobian CentroidsThomas Walker, Randall Balestriero, Richard Baraniuk
Detecting failures in World Model (WM)-based planning requires monitoring whether the model is behaviorally aligned with the current task, which in turn requires studying its internal representations. Here, we show that centroids---sub-component Jacobian row-sums---effectively identify the behaviora
world model - arxiv:2609.33937 · cs.CVTest-Time Generalized Category DiscoveryShambhavi Mishra, Omprakash Chakraborty, Julio Silva-Rodriguez, Ismail Ben Ayed +2
Test-Time Adaptation (TTA) and Generalized Category Discovery (GCD) are traditionally treated as disjoint problems: the former adapts models to domain shift assuming all test classes are known, while the latter discovers novel categories assuming labeled training data for known classes. However, rea
benchmark - arxiv:2609.33935 · cs.CVVideo, Ergo Genero: Unifying Video Tasks via Spatiotemporal AnalogyChia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia
Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks a
manipulationevent camera - arxiv:2609.33927 · cs.LGOptimizing the Phi-2 Small Language Model for Real-time Chatbot Applications Using Parameter-Efficient Fine-Tuning (PEFT) with QLoRA QuantizationPhanTan Khanh Nguyen, Ashfaq Ali Shafin, Khandaker Mamun Ahmed
This study explores the optimization of the Phi-2 Small Language Models (SLMs) for real-time chatbot applications through Parameter-Efficient Fine-Tuning (PEFT) and Quantized Low-Rank Adaptation (QLoRA). QLoRA specifically refers to the integration of PEFT with LoRA alongside a 4-bit quantization pr
memory - arxiv:2609.33920 · cs.AIHyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM AgentsTingsong Xiao, Nithish Balachandar Moudhgalya, Chandrayee Basu, Lichao Wang +4
Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment int
llm agent - arxiv:2609.33915 · cs.LGTraining Witnesses: Trusting the Training without Trusting the TrainerHoujun Liu, Pratyusha Sharma
Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on
leaderboard - arxiv:2609.33910 · cs.AIWhen Consent Outlives Context: Residual Authority Replay in Long-Lived AgentsZhihao Zhang, Chao Wang, Rujia Li, Qingze Wang +2
LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can outl
llm agent - arxiv:2609.33906 · cs.LGJIVE: Jacobian-Informed Volume Expansion for Diverse Generative SamplingGuangxun Zhang, Brian Cai, Boxuan Zhang, Chao Chen +1
Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of the resulting endpoints. We introduce JIV
benchmark - arxiv:2609.33905 · cs.CLSlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI WritingDhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh
SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on
benchmark - arxiv:2609.33901 · cs.LGFinite Probes Suffice: Identifiability and Universality for Weight-Space LearningSoutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron +1
Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical founda
benchmark - arxiv:2609.33899 · cs.CLNSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech ModelsZiwei Chen
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final t
benchmark - arxiv:2609.33893 · cs.LGMISHAP-Bench: A Hallucination Benchmark for Large Audio-Language ModelsZhi Wen Soi, Giulio Segalini, Jian-Jia Chen, Lydia Chen
Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the aud
benchmark - arxiv:2609.33889 · cs.LGWhere Activation Sparsity and KV-Cache Sparsity Cross in LLM DecodingJungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo +1
At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, yet their reported speedups are hard to c
long context - arxiv:2609.33887 · cs.LGFaster Block-Diffusion Serving with Distribution-Free Risk GuaranteesJungseob Lee, Dongyub Jude Lee, Chanjun Park, Sugyeong Eo +1
Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on pr
benchmark - arxiv:2609.33885 · cs.LGProspective Interpretation Risk: Principled Communication Control Between LLMsWanrong Yang, Rehan Deen, Julian Ma, Yuheng Fan +6
Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers
multi-agentagentic - arxiv:2609.33882 · cs.RODexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous ManipulationHan Yang, Yian Wang, Yunlong Song, Zhenjia Xu +1
Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving
manipulationdexteroustactilegrasptool-use - arxiv:2609.33872 · cs.RORobot-GST: geometry-aware spatial-temporal robot policy representation and evaluationSichao Liu, Zekun Wang, Lixuan Tang, Yiming Li +5
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions
manipulationrobot policyevaluation framework - arxiv:2609.33871 · cs.CLPopulation Physics, Population Problems: Safety and Emergence in LLM SocietiesAdrian de Wynter
The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic syst
autonomous agentmulti-agentagenticagent system - arxiv:2609.33870 · cs.AIWhen Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal AgentsJanvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur, Ece Kamar +1
LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable
agentllm agentpost-trainingbenchmark - arxiv:2609.33867 · cs.AIR$^2$ Flow: Recursive Self-Improvement via Recursive Skill EvolutionMingda Zhang, Qiang Huang, Yanjin Li, Zijia Wang +3
LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obst
self-improvement - arxiv:2609.33857 · cs.LGPI-NOMT: Physics-Informed Neural Optimal Mass Transport for Brain Fluid DynamicsMehmet Emin Acar, Vahit Bugra Yesilkaynak, Helene Benveniste, Gozde Unal
Recovering hidden transport mechanisms from sparse spatiotemporal observations is a fundamental inverse problem in scientific machine learning. In brain tracer imaging, dynamic contrast-enhanced MRI (DCE-MRI) provides time-resolved measurements of tracer concentration, while the underlying velocity
post-trainingbenchmark - arxiv:2609.33855 · cs.LGProgram-Verified Self-Evolution for Vision-Language ModelsAhmed Heakl, Sungik Choi, Moontae Lee, Salman Khan
Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of mo
scene graphself-evolvingbenchmark - arxiv:2609.33854 · cs.CVReDrive: Shaping Representations with World Modeling for End-to-End DrivingYueting Zhu, Shaoyu Chen, Yuehao Song, Hui Sun +3
Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system archi
world model - arxiv:2609.33845 · cs.AIHow code helps different tasks? A decompositional lens on LLM post-trainingZheng Yu, Yiwei Li, Yishen Chen, Xiang Li +3
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decomposi
post-trainingpost training - arxiv:2609.33844 · cs.LGViBR-WM: Visual Bayesian Regression for World ModelingJifan Li, Ning Ning
Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle
world model - arxiv:2609.33836 · cs.RODeltaSeek: Toward Active Perception in Evolving Construction EnvironmentsSanjay Acharjee, Md Nazmus Sakib
Construction environments evolve continuously, causing large geometric changes that degrade static mapping and registration performance. This necessitates active perception, where robots deliberately select sensing configurations to resolve the environment's current state. We present DeltaSeek, an i
benchmark - arxiv:2609.33832 · cs.ROAchieve What You Imagined: Learning to Align Actions with Visual PlansYuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo +4
World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before produc
manipulationaction headworld modelaction-conditioned - arxiv:2609.33822 · cs.AIVestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and MemoryJayant Parashar, Eugene F. Douglass, William C. Bastian, Suchendra M. Bhandarkar
An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces in
memoryagentbenchmark - arxiv:2609.33818 · cs.CVAugmenting Visual Anomaly Detection with Automated InterpretabilityAntonio De Santis, Arsenio Leo, Marco Brambilla
Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different compone
benchmark - arxiv:2609.33814 · cs.LGAnnealed Sinkhorn with Momentum: Certified Unregularized Optimal Transport in Linear MemorySamuel J. K. Chin, Maximilian Schiffer
We characterize Bregman Douglas-Rachford splitting (BDRS) for unregularized discrete optimal transport and develop an anytime primal-dual certificate in linear memory. We first establish that BDRS coincides with warm-started Inexact Proximal point method for exact Optimal Transport (IPOT) using a si
memorybenchmark - arxiv:2609.33812 · cs.LGIdentical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight ModelsEduardo Ariño de la Rubia, Szilard Pafka
Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an XGBoost classifier for airline delays, a
agentbenchmark - arxiv:2609.33810 · cs.LGControlling Speaking Rate in Autoregressive TTS via Activation SteeringFrancesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi +3
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-
benchmark - arxiv:2609.33807 · cs.ROCodeActionBench: Evaluating Agentic Code-as-Policy for Embodied ManipulationYiheng Lyu, Xueying Jiang, Wenhao Li, Shijian Lu +1
How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning,
embodiedmanipulationgraspagenticcode-as-policybenchmark - arxiv:2609.33803 · cs.LGDiffusion Reward ModelsXiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze WangZiqing Qiao +11
Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonabl
rlhfbenchmark - arxiv:2609.33802 · cs.LGAutonomous phase discoveryShiyu Zhou, Yuxuan Zhang, Sebastian Wetzel, Roger Melko +1
Understanding quantum phases of matter has long relied on physicists' intuition and mathematical tools such as symmetry and topology. Remarkably successful as these approaches have been, they provide no universal way to explore a Hamiltonian space whose organizing principle is not known in advance.
benchmark - arxiv:2609.33791 · cs.LGDo We Really Need KL Divergence for On-Policy Distillation of Large Language Models?Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi Wang +10
Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in t
post-trainingbenchmark - arxiv:2609.33780 · cs.LGSelecting Diverse SFT Traces Improves Post-RL GeneralizationDylan Zhang, Mingyuan Wu, Jinning Li
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to
benchmark - arxiv:2609.33778 · cs.AIEvidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes WrongMegan Diehl, Ser-Nam Lim
Modern multi-hop LLM agents are equipped with built-in mechanisms to detect errors in intermediate reasoning steps. Such errors trigger corrective actions from these agents, which mostly follow the paradigm of retrying the steps or the reasoning trajectories. Not only are these retries expensive, we
llm agentagentic - arxiv:2609.33773 · cs.AILearning Strategies to Break JudgesGuruprerana Shabadi, Aaditya Naik, Rajeev Alur, Mayur Naik
As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important
agentai agentagentic - arxiv:2609.33772 · cs.AISkill2Env: Capability-Oriented Environment Synthesis from Skills for General AgentsWeiyi Xu, Xiaowen Yang, Wen Da, Hang Xu +6
Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instr
agentagent benchmarktool usetool-usepost-trainingbenchmark - arxiv:2609.33769 · cs.CVM3-Score: Fidelity, Memorization and Coverage as Separate Axes for Evaluating Generative Radiology Image ModelsSathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera
Quantitative evaluation of generative models for radiology remains challenging. Clinically relevant structures are often small and infrequent, feature spaces learned from natural images may represent them poorly, and a single summary score cannot distinguish limited fidelity from limited diversity.
evaluation framework - arxiv:2609.33768 · cs.LGDEALS: Decentralized Expertise-Aware Load Serving for Multi-Agent LLM SystemsJingjuan Huang, Wenbin Wang, Yanchuan Yin, Alvaro Velasquez +1
Multi-agent systems (MAS) have recently emerged as an effective approach for coordinating large language model (LLM)-based agents to solve complex tasks through structured interactions. In practice, MASs often handle a stream of heterogeneous and complex tasks, requiring agents to decompose each tas
agentmulti-agentagent system - arxiv:2609.33765 · cs.ROPrincipal Steering Subspaces for Online Adaptation of Frozen Generative Robot PoliciesJialeng Ni, Nathan Zhao, Kunpeng Song
Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise
vision-language-actionvlavla policyhumanoid - arxiv:2609.33764 · cs.LGBeyond Fixed Features: Architecture-Dependent Sensitivity to Node Representations under HeterophilyPriyanath Maji, Sidharth Gaur, Rajavinoth Paul Durai
Graph Neural Networks (GNNs) perform well on homophilic graphs but struggle in heterophilic settings, where connected nodes often carry dissimilar labels. Existing evaluations typically compare architectures under a fixed node-feature representation, leaving unclear whether conclusions about heterop
benchmark - arxiv:2609.33763 · cs.AISecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity VulnerabilitiesXiaonan Luo, Yue Huang, Kehan Guo, Ping He +8
Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at sc
agentbenchmark - arxiv:2609.33762 · cs.LGEfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li +6
LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsist
memoryagentllm agent - arxiv:2609.33759 · cs.LGPositions Are Not Facts: The Mismatch Between KV Caches and MemoryChanghai Zhou, Yuhua Zhou, Shiyang Zhang, Jun Gao +4
When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and
memory - arxiv:2609.33758 · cs.CVENet-GP: Unified Document Image RestorationSujal Burad, Aakanksha, A. N. Rajagopalan, Sumit Shekar
Reliable document digitization in uncontrolled capture settings is challenging because real images exhibit multiple interacting degradations rather than a single isolated distortion. Documents thus captured are affected simultaneously by geometric distortions, like page warping, as well as photometr
benchmark - arxiv:2609.33757 · cs.LGYuE2: Unifying Symbolic and Audio Music Generation at Frontier QualityRuibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu +31
Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning.
agenticbenchmark - arxiv:2609.33754 · cs.LGCollaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational SilosSimeon Allmendinger, Domenique Zipperling, Burhanettin Bahadir Kibar, Niklas K{ü}hl
Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collabor
benchmark - arxiv:2609.33748 · cs.ROAnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action ModelsRui Wang, Xiangyu Wang, Donglin Yang, Yibo Li +3
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less
manipulationrobotwin - arxiv:2609.33746 · cs.LGPQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate AttentionKunming Shao, Jierun Chen, Yanli Wang, Ruoyu Wang +3
At each decoding step a language model attends over the key-value (KV) cache of every earlier token, so at long context the attention call is bounded by memory bandwidth. Sparse attention reads only a subset of keys chosen by a cheap score estimate, and most methods give the unread tokens zero weigh
memorylong context - arxiv:2609.33738 · cs.CLThe Effects of Incremental Instruction Delivery on Language-Model Creative WritingAnshuman Singh, Abrar Eyasir, Haseeb Yaqoob, John Manavalan
Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with obj
benchmark - arxiv:2609.33737 · cs.ROMomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous DrivingZiying Song, Shengkai Zhang, Lei Yang, Haozhuang Chi +5
Long-horizon planning enables autonomous vehicles to anticipate scene evolution and potential risks, supporting safe and stable decisions in complex interactions. However, existing methods struggle to propagate motion trends from observed history into the future. Long rollouts based on a single late
world model - arxiv:2609.33729 · physics.app-phWave-Domain Semantic Equalization Using a Practical Dynamic Metasurface Antenna with Strong Mutual CouplingYassine Sghaier, Emilio Calvanese Strinati, Philipp del Hougne
Semantic mismatch between independently trained AI-native agents in heterogeneous networks can impair semantic communications. Hybrid analog-digital semantic equalization can align the incompatible latent representations without retraining the semantic transceivers. We study a practical realization
benchmark - arxiv:2609.33728 · cs.LGALDER: Discovering the Laws of a World by Acting in ItTeng Cao, Yu Deng, Quentin Delfosse, Kristian Kersting
Reliable world models should not only predict future states but express how actions change the world in an explicit, transparent and testable form, such as equations. Yet methods that rely on a fixed set of trajectories cannot distinguish equally good competing hypotheses, while searches over a fixe
world modelbenchmark - arxiv:2609.33725 · cs.LGReliable Replay through Spatial Coherence in Online Continual LearningHaixiang Sun, Jiefu Zhang, Yinghao He, Yang Xu +3
Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated re
memory - arxiv:2609.33717 · cs.AISelf-Designed Evaluators and Warm Memory for Long-Horizon AgentsSaeid Asgari, Emre Kiciman, Leonardo de Oliveira Nunes, Ranveer Chandra
A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materi
memoryagentagenticbenchmarkevaluator - arxiv:2609.33716 · cs.CVRevisiting Diffusion Fine-Tuning for Unsupervised Domain AdaptationXuan Qi, Yi Wei, Daniele Berardini, Vito Paolo Pastore +1
Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multipl
benchmark - arxiv:2609.33715 · cs.LGTheory Guided and Interpretable Neural Operator Design for Partial Differential Equation LearningZeyuan Song, Zheyu Jiang
Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the
benchmark - arxiv:2609.33713 · cs.AIBIRD: Distilling Decision Boundaries into Rationales for MLLM AdaptationAnglin Liu, Yanlin Wu, Ruichao Chen, Yuting Zhang +5
Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional ob
self-improving - arxiv:2609.33707 · cs.RODoes Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View CollapseFuta Waseda, Shuhei Kurita, Isao Echizen
Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial
vision-language-actionvlalibero - arxiv:2609.33702 · cs.AIUnderstanding Confabulation and Rethinking Reconstruction in Activation ExplanationsGert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen
Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for pred
evaluation framework - arxiv:2609.33699 · cs.AISpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware SpecificationsFeilian Huang
Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification a
benchmark - arxiv:2609.33691 · cs.CLWhen Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their TranslationsMáté Metzger
Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Tibetan, comprising 9,833 human-aligned un
agentautonomous agentagentic - arxiv:2609.33688 · cs.LGTopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology OptimizationBin Lou, Yuxuan Cheng, Huaizhi Zong, Junhui Zhang +1
Deep learning has emerged as an efficient alternative for predicting high-performance material distributions in topology optimization. Existing methods struggle to accurately capture load-transfer information, limiting out-of-distribution generalization, while their model architectures often incur h
benchmark - arxiv:2609.33683 · cs.CVMAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic TasksHao Chen
When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary
manipulationbenchmark - arxiv:2609.33678 · cs.AISWE-Game: Can Coding Agents Build the Games We Want?Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou +7
We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-
agentbenchmarkevaluator - arxiv:2609.33676 · cs.AIAuditing Agent Actions through Query-Conditioned AttributionYifan Liu, Praveen Venkateswaran, Abdulhamid Adebayo, Dong Wang
LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for
agentllm agentbenchmark - arxiv:2609.33672 · cs.CLReset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy HysteresisAdi Shnaidman
Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocat
benchmark - arxiv:2609.33668 · cs.CVWhen Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-LabelingPing Guo, Zhiqi Huang, Xinran Li
Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degra
benchmark - arxiv:2609.33667 · cs.LGScalable Attribution and Control of Model Behavior During TrainingSleem Abdelghafar
Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguis
evaluator - arxiv:2609.33666 · cs.CLOne Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment GenerationNafis Irtiza Tripto, Delvin Ce Zhang, Mahjabin Nahar, Dongwon Lee
Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human dis
ai agent - arxiv:2609.33665 · cs.AICompoWorld: Compositional Environment Scaling for General AgentsXiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu +8
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We in
world modelbenchmark - arxiv:2609.33662 · cs.LGAudit-First VAPO: Risk-Certified Selective Updates under Imperfect VerificationMiaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang +4
Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite seconda
benchmark - arxiv:2609.33658 · cs.AIAgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM AgentsTianzhuo Yang, Zirui Mi, Yantao Huang, Guoxi Zhang +3
Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct chal
agentllm agentagenticevaluation framework - arxiv:2609.33655 · cs.LGTowards Eliminating Catastrophic Forgetting in the Curriculum Learning of Math Reasoning TasksZengyan Yang, Yangyang Wu, Kai Huang, Pengfei Lyu +2
Curriculum learning has found broad application across numerous domains. Nevertheless, its effectiveness is intrinsically curtailed by catastrophic forgetting, driven by the shifts in model parameter distributions between curriculum tasks. In this paper, we investigate the phenomenon of catastrophic
curriculum learningbenchmark - arxiv:2609.33653 · cs.RODemonstration-Free Success-Probability Reward Learning for Generalist Robot PoliciesDuo Wu, Haifeng Wang, Rongwei Lu, Jinghe Wang +5
Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonst
manipulationlibero - arxiv:2609.33650 · cs.LGFuseAlign: Forced Alignment in the WildMithilesh Vaidya, Stephen Bailey, Sumukh Badam, Matthew Bendel +1
Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect
benchmark - arxiv:2609.33647 · cs.ROInfraVLA: Extending Vision-Language-Action Navigation with Infrastructure CamerasLukas Vierling, Benjamin Ramtoula, Luke Robinson, Ronald Clark +1
Many indoor environments in which robots operate, such as warehouses, offices, and hospitals, already have cameras installed. They observe parts of the building that the robot cannot see from where it stands, yet navigation policies, including recent vision-language-action (VLA) models, do not use t
vision-language-actionvlaquadruped - arxiv:2609.33646 · cs.CVProbe to Act: Elevating Browser-Use Agent via Active Visual ProbingKeliang Li, Heng Wang, Chen Hu, Daxin Jiang +2
Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve
agentbenchmark - arxiv:2609.33642 · cs.AILearning to Learn from Context: Synthetic Training from Perturbed Public DocumentsHaoyi Wu, Yang Xiao, Yusong Sun, Wenyang Hui +3
Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality docu
long-context - arxiv:2609.33640 · cs.LGSafeMol: Dual-Modality Safety Alignment for Molecular Multimodal ModelsXinmiao Wang, Ruijie Wang, Menghui Wang, Jiawei Chen +4
Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows tha
benchmark - arxiv:2609.33639 · cs.AITrajectory Unlearning on LLM-based AgentsYingdan Shi, Ren Wang
Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an
agentautonomous agentagenticbenchmark - arxiv:2609.33637 · cs.ROHierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication LossWeihao Sun, Gehui Xu, Andreas A. Malikopoulos
In this paper, we propose a hierarchical multi-agent reinforcement learning framework for coordinating robot teams in warehouse environments under communication loss. We partition the robot team into groups, with centralized coordination within each group and distributed coordination across groups.
multi-agent - arxiv:2609.33634 · cs.AISafety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety GuardrailsGert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen +2
Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are stru
benchmark - arxiv:2609.33630 · cs.LGlapanda: A Matrix-Free Differentiable Solver for Nonconvex Constrained Optimization LayersYuankun Chen, Zifei Nie, Kangyu Lin, Ján Drgoňa +1
Differentiable optimization brings the structural guarantees of mathematical optimization to network pipelines, allowing them to be trained end-to-end. However, its application remains challenging for nonconvex constrained problems, as existing differentiable solvers often suffer from limited modeli
memorybenchmark - arxiv:2609.33628 · cs.LGClimbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement LearningChenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang +1
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such a
curriculum learning - arxiv:2609.33627 · cs.CVStoryEngine: A State-Grounded Agentic Framework for Video StorytellingYingrui Wang, Zeqing Wang, Yeying Jin
Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of
agenticbenchmark - arxiv:2609.33618 · cs.AIParaAgent: Reinforcing Parallel Acting in Open-World Tool EnvironmentsShengbin Yue, Hongru Wang, Siyuan Wang, Xiaoxin Chen +2
Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them withou
multi-agentbenchmark - arxiv:2609.33609 · cs.LGYou Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration RefinementJiarong Wen, Qi Wang, Yun Qu, Yixiu Mao +7
In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely o
benchmark - arxiv:2609.33608 · cs.AILearning Transferable Reaction Mechanisms from Visual Chemical KnowledgeYujian Yuan, Jiaxin Xu, Xin Cai, Yufan Chen +3
Reaction mechanisms describe the step-by-step transformations underlying chemical reactions and are central to reaction analysis and synthesis. Learning-based models have achieved strong performance on established mechanism-prediction benchmarks, but transferring them to unseen chemistry remains cha
benchmark - arxiv:2609.33603 · cs.CVViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable RevisionYujian Yuan, Xin Cai, Yufan Chen, Jiaxin Xu +4
Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection an
benchmark - arxiv:2609.33601 · cs.AIJustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation QuantizationKaicheng Yang, Kaisen Yang, Chunyu Liu, Xianglong Yan +7
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization.
post-training - arxiv:2609.33597 · cs.LGThe Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM TracesGeorgios Feretzakis, Alexandros Papaspyridis
Side-channel evaluators routinely inspect leakage tests while acquisition is still running, and extend or stop the campaign based on what they see. Fixed-horizon screening such as the Welch $t$-test with threshold $|t|>4.5$ gives no error guarantee for this monitored decision rule. We study anytime-
evaluator - arxiv:2609.33595 · cs.ROBeyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual PlanningBoyuan Zhang, Yingjun Du, Xiantong Zhen, Ling Shao
Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not
world modelaction-conditioned - arxiv:2609.33593 · cs.CVLoopLUT: 3D Lookup Tables with Progressive Region Refinement for Real-Time 4K Image EnhancementYang Ye, Jiajun Ma, Chen Wu, Wei Wang +4
Color enhancement of 4K images must meet a quality target under a tight compute budget. Three-dimensional lookup tables (3D LUTs) dominate real-time enhancement because they decide at low resolution and apply a per-pixel lookup at full resolution. A single global LUT, however, is spatially invariant
benchmark - arxiv:2609.33589 · cs.LGTGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMsZihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai +3
Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or lea
benchmark - arxiv:2609.33582 · cs.CVFill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-ResolutionXingfu Yi, Xiaoxue Yu
Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a ma
memorybenchmark - arxiv:2609.33581 · cs.CVForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language NavigationKunhui Wang, Xintong Zhang, Junyu Gao, Changsheng Xu
Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across diff
benchmark - arxiv:2609.33579 · cs.AIOpenFC: Learning Verification Policies towards Open-Search Fact CheckingXinming Wang, Kaixiang Qiu, Yansong Lin, Chunji Lv +4
Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipeline
tool-usebenchmark - arxiv:2609.33575 · cs.ROSLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action ModelsTianfu Li, Haoxuan Xu, Wenbo Chen, Haitian Li +6
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often
vision-language-actionvlavla modelmanipulationworld modelaction-conditioned - arxiv:2609.33570 · cs.LGHiLoRe: What to Store, Compress, or Recompute for Efficient GRPO TrainingXinrui Chen, Mengyang Li, Ou Wu, Ji Zhang
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existi
memory - arxiv:2609.33566 · cs.LGCorrect then Forecast: Observer State-Space Models for Time Series ForecastingAlexis-Raja Brachet, Guillaume Clavier--Frémond, Abdelhakim Ziani, Pierre-Yves Richard +1
Time series forecasting requires extrapolating the dynamics of an observed process beyond the last available measurement. Yet recurrent forecasting models typically treat observations as inputs that directly control their latent dynamics. It leads to a regime change when these observations become un
latent dynamicsbenchmark - arxiv:2609.33565 · cs.AIDr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search AgentsZhipeng Qian, Zihan Liang, Yufei Ma, Jie Ma +9
A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Y
knowledge graphself-evolvingbenchmark - arxiv:2609.33563 · cs.LGMA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement LearningBrandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal f
world modelaction-conditionedmulti-agent - arxiv:2609.33558 · cs.CVPGL-3D: Towards Progressive Geometric Learning for 3D Visual Query LocalizationLiang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang +5
3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object a
memorybenchmark - arxiv:2609.33551 · cs.ROFoLD: Force-Informed Learning for Dexterous Articulated Object ManipulationHaowei Shen, Tingai Li, Yumeng Liu, Wenyuan Guang +5
Transferring human demonstrations to dexterous robots remains challenging because differences in hand morphology and contact dynamics often cause retargeted motions to fail at producing the intended object behavior. We present \textbf{FoLD}, a framework for learning dexterous manipulation of articul
manipulationdexterousbenchmark - arxiv:2609.33550 · cs.LGSparsity by Default: The Theory and Practice of ARD in Gaussian Process Regression for Variable SelectionJia Cai
Automatic relevance determination (ARD) is the standard device for input selection in Gaussian process (GP) regression. By giving the covariance kernel a separate lengthscale for every input and learning those lengthscales by maximizing the marginal likelihood, ARD lets the data decide which coordin
embodied - arxiv:2609.33548 · cs.LGTerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the EdgeHoussem Sifaou, Prabodh Katti, Bipin Rajendran, Osvaldo Simeone
Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family o
memory - arxiv:2609.33546 · cs.ROSteer2Grasp: Inference-Time Embodiment-Aware Steering for Diverse Physically Feasible Grasp DiffusionVignesh Vembar, Ayush Kaura, A Padmaprabhan, Siddharth Sinha +4
Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps throug
grippergrasp - arxiv:2609.33536 · cs.LGDoes Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM ServingJiantong Jiang, Yue Yang, Peiyu Yang, Feng Liu
Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready.
memory - arxiv:2609.33534 · cs.CLManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold PerspectiveRui Liu, Chenheng Zhang, Haoxuan Li, Zhouchen Lin
Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forg
benchmark - arxiv:2609.33530 · cs.LGE-CONAN (Entailment, CONtradition And Neutral) Diagnostics Dataset Investigating Linguistic Phenomena in Arabic Natural Language UnderstandingKhloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
Natural Language Understanding (NLU) plays a crucial role in various applications, yet its performance suffers from weaknesses in handling the complexities of human languages, ranging from lexical ambiguity to high-level reasoning difficulties. Analyzing errors across diverse linguistic phenomena is
benchmark - arxiv:2609.33524 · cs.AIEverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha ResearchSiyuan Li, Jiangfeng Zhang, Rui Yao, Weihua Qiu +2
Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the
self-evolving - arxiv:2609.33521 · cs.LGLLM4Trust: Exploring the Capabilities of Large Language Models for Trust EvaluationJie Wang, Yanbo Sun, Zheng Yan, Jiahe Lan +1
Trust evaluation plays a critical role in cybersecurity by supporting risk mitigation and decision-making. A variety of trust evaluation methods have been proposed, with learning-based approaches offering high accuracy and automation. However, they often require substantial ground truth, suffer from
benchmark - arxiv:2609.33517 · cs.MATRACE: Governing Memory Validity in Evolving Multi-Agent SystemsWenjun Xiong, Shengtao Zhang, Shangding Gu, Bo Tang +4
Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, and faithful to its s
memorypersistent memoryagentmulti-agentagent system - arxiv:2609.33509 · cs.AIWhat Happens During Autonomous Deep Research After the User Steps Away?Yimin Liu, Yijia Zhang, Yanmin Li, Tangwen Luo +3
In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommenda
agentevaluatorevaluation framework - arxiv:2609.33505 · cs.AIWhen Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned RoutingJian Chen, Zixuan Yuan
Modern learned systems increasingly combine learned components with search, repair, or external solvers. Benchmarks often measure the resulting end-to-end system, while scientific claims may concern only one component, creating an attribution problem: evidence can fail to support the requested compo
benchmark - arxiv:2609.33503 · cs.AIRelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache ReuseRuoling Qi, Yirui Liu, Xuaner Wu, Yuxin Jin +4
Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk inter
retrieval-augmented - arxiv:2609.33497 · cs.LGHamiltonian JEPA: Action-Conditioned World Models with an Inherited Control StateTamim Zoabi, Ameen Ali, Lior Wolf
Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing actio
world modelaction-conditionedbenchmark - arxiv:2609.33496 · cs.LGChameleon: Dynamic Format Adapter for Efficient DiffusionArnab Sanyal, Sandeep Chinchali
Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best f
post-training - arxiv:2609.33495 · cs.CLLLMs Trust Their Own: Identity-Dependent Conformity in Multi-Agent SystemsLiron Soffer, Ravid Shwartz-Ziv, Chen Shani
Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs' responses depend on the social identity of other agents, beyond the effect o
manipulationmulti-agentagent system - arxiv:2609.33492 · cs.AIFederated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement LearningDebasmita Dey, Tanmay Sen, Himel Mallick
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences i
agentmulti-agent - arxiv:2609.33485 · cs.CLDISCO: Distributed Long Context Scaling with Grounding-Reasoning DisaggregationGuanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee, Branislav Kveton +4
While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exha
long-contextlong context - arxiv:2609.33484 · cs.ROAMBIT: Anticipatory Multimodal Body Recruitment for Bimanual Tracking on a HumanoidHanlong Li, Sihan Tan, Takeshi Ashizawa, Benjamin Yen +1
A humanoid with 5-DoF arms cannot track generic bimanual end-effector trajectories with its arms alone; pelvis and waist motion must be recruited, but which motion, and when, is not uniquely determined. On a Unitree R1 in fixed double support, the set of dynamically valid recruitment strategies (pel
humanoid - arxiv:2609.33482 · cs.LGHow Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional CoverageQianyi Chen, Bo Li
Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study wh
benchmark - arxiv:2609.33477 · cs.AIJust Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMsYirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin +5
Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbi
long-context - arxiv:2609.33470 · cs.AILiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear PayoffsHaochen Luo, Yifan Li, Binh Minh An, Xiaolong Luo +3
Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamental
llm agentmulti-agentagent systemevaluation framework - arxiv:2609.33467 · cs.LGA Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous RewardsAndreas Plesner, Curtis Northcutt, Francisco Guzmán, Anish Athalye
When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifie
post-training - arxiv:2609.33464 · cs.ROVIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided SimulationJianan Wang, Haoquan Zhai, Siyang Zhang, Bin Li +6
World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action precondi
embodiedworld model - arxiv:2609.33463 · cs.CLRethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned FeaturesCunchun Li, Haonan He, Yifan Gao, Minglei Li +3
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-rewe
benchmark - arxiv:2609.33462 · cs.CVSphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 CameraShriram Damodaran, Soumyaratna Debnath, Cheston Tan, Lin Wang
Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images
embodiedai agentbenchmark - arxiv:2609.33460 · physics.app-phBare-Die Antiferromagnetic ComputingYu Liu, Zhuoting Han, Zexin Feng, Peixin Qin +14
Semiconductor electronic devices are increasingly constrained by fundamental quantum tunneling effects and charge-based mechanisms, which severely limit further miniaturization, write-speed scaling, and environmental robustness of silicon-based technologies. These limitations are particularly prohib
memory - arxiv:2609.33457 · cs.LGA Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based EvaluationMd Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany, Tanzima Hashem
Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration:
benchmark - arxiv:2609.33455 · cs.AIWhat Shared Prefixes Hide: Trajectory Dropout for On-Policy DistillationZzizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo
On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted
benchmark - arxiv:2609.33451 · eess.SYDAOCP: a dual active set solver for optimal control problemsAlberto Zaupa, Samuel Erickson, Mikael Johansson
We present DAOCP, a dual active set solver for linear quadratic optimal control problems with stage-wise equality and inequality constraints. Active set methods are leading Model Predictive Control benchmarks for full-body robotics, but existing solvers operate on dense QPs, while typical problem di
benchmark - arxiv:2609.33446 · cs.AIHESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage AgentsZhuowen Liu, Zhixuan Wang
Security operations centers receive far more alerts than analysts can investigate, and organizations that cannot send their telemetry to hosted models must automate triage with small open-weight LLMs on their own hardware. Current LLM agents leave the investigation procedure to the model, and small
llm agent - arxiv:2609.33443 · cs.CLContext Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM BackendsSeonghyeon Go, Yongwoo Kim, Hyeonjin Cha, Jaeho Shin
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving the
benchmark - arxiv:2609.33441 · cs.CLMIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal MisinformationRuihong Zeng, Jonathan Tonglet, Preslav Nakov, Iryna Gurevych
Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what
benchmark - arxiv:2609.33440 · cs.AIMAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRIMd. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni
Objective cognitive assessment from neural signals supports neurorehabilitation, but individual-level prediction from task-based fMRI (tfMRI) remains difficult because neural features coexist with substantial demographic and scanner-related variation. We present the Multi-task Activation and Contras
benchmark - arxiv:2609.33439 · cs.AIRaven: The Harness of Harnesses for Composable Agentic IntelligenceEverMind AI
As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits t
agentai agentmulti-agentagenticagent system - arxiv:2609.33436 · cs.LGSchemaMem: Schema-Indexed Recurrent Memory for Delayed State RetrievalSungwoo Goo, Hwi-yeol Yun, Sangkeun Jung
Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture comb
memorymemory architecturepersistent state - arxiv:2609.33430 · cs.AIAPEX: An Extensible Model for Agent-Assisted Production SchedulingFelix J. Grumbach, Stefan Görlitz
Production scheduling requires realistic models that reflect operational constraints and efficient methods that balance competing goals. Putting these methods into use also requires data integration, model adaptation and specialist expertise. We present APEX, an extensible production scheduling fram
agentbenchmark - arxiv:2609.33429 · cs.AIGraph-Guided Repository Environment ConstructionJianying Pan, John Zhang, Hongyu Zhang
Coding agents now increasingly rely on execution to validate their solutions, making the construction of reliable execution environments a critical enabling capability. However, repository environment construction is challenging because execution requirements are fragmented across repository artifac
benchmark - arxiv:2609.33426 · cs.CLTeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher AlignmentZhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng +4
Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out cha
benchmark - arxiv:2609.33419 · cs.CVTT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video PretrainingShih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang +2
Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized re
v-jepa - arxiv:2609.33417 · cs.CLGrounding Memory Summarization in Utility IntentZhenyu Lei, Mingjia Shi, Xingbo Fu, Haoyu He +2
Existing summarizers for memory systems are typically optimized for human-facing criteria such as faithfulness, which misaligns with their true objective: preserving the evidence needed to support future queries. We show that conditioning summarization on query-answer pairs substantially improves an
memory - arxiv:2609.33414 · cs.CVTTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language ModelsShuning Wang, Zhiheng Wu, Xun Zhou, Chongyang Cui +7
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anc
benchmark - arxiv:2609.33412 · cs.CVResolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA PlanningJunhao Xiao, Haoxiang Zhao, Menghao Fang, Jinkui Zhang +7
Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pix
vision-language-actionvlaaction-conditioned - arxiv:2609.33411 · cs.AIMetaBench-Harness: Unlocking End-to-End Optimization of Benchmark HarnessesXuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma +2
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving be
benchmarkevaluator - arxiv:2609.33410 · cs.LGFoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic BackwardSriman Achanta
Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization reference as it scans keys. Earlier contributions therefore remain provisional and may require resca
long-contextlong context - arxiv:2609.33409 · cs.CLDense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy DistillationYuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen +4
On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need
agentagenticagent benchmarkbenchmark - arxiv:2609.33408 · cs.LGStarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC NetworksMustafa Bora Çelik, Ceren Çelik, Orhan Gazi
In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90\% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses toward maximum uniform entropy ($H=2.584$ bits) as sparsity increases, failing to capture long-range gai
persistent state - arxiv:2609.33407 · cs.LGLet CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure PriorsEmma Lei Hovmand, Jonas Elsborg, Melih Kandemir, Arghya Bhowmik
De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical prior that should be improved by likeli
evaluator - arxiv:2609.33403 · cs.AIDataMagic: Authoring Data Videos through Declarative Multi-Agent OrchestrationYupeng Xie, Zhenyang Wang, Liangwei Wang, Jiayi Zhu +2
Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack na
multi-agent - arxiv:2609.33402 · cs.CVVaME: Exploring Variational Latent Reasoning for Multimodal EmbeddingsPeixi Wu, Mingzhou Jiang, Feipeng Ma, Biao Yang +10
Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing appr
benchmark - arxiv:2609.33401 · cs.AIEvaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective AutomationYixuan Liu
Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models expose typed decisions with probabilities that software can use to allow, block, or escalate inputs, but whether these probabilities support reliab
agent - arxiv:2609.33399 · cs.CVSciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image GenerationJiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin +5
In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet ve
benchmark - arxiv:2609.33398 · cs.AICOEVO: Co-Evolving Context and Parameters for Recursive Self-ImprovementSiwei Chen, Xinping Bao, Xinyu Cai, Yuan Cao +2
Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the ext
self-improvementonline learning - arxiv:2609.33397 · cs.AICoViST: Visual Token Compression via Composable StatesQi Zhang, Xiandong Meng, Ronggang Wang, Siwei Ma
Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. T
benchmark - arxiv:2609.33392 · cs.LGAutoHGNN: Robust and Efficient Neural Architecture Search for Hypergraph Neural NetworksSirui Li, Pietro Liò b, Xinsheng Li, Baisong Liu +1
Hypergraph neural networks have achieved significant success in recent years. However, manual architecture crafting is labor-intensive and often fails to capture complex, higher-order relations, making the automation of hypergraph neural network structure design crucial. To improve the automation an
benchmark - arxiv:2609.33385 · cs.CLOLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert OffloadingJingyuan Xiao, Jiayue Wang, Yitao Hu, Xinning Wang +6
Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a na
memory - arxiv:2609.33384 · cs.CVPulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion TransformersYutong Wang, Xingtong Ge, Enhuai Liu, Yunke Wang +3
Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with
post-training - arxiv:2609.33379 · cs.AINLPG: Natural-Language Policy Gradients for Self-Evolving Language AgentsXu Liu, WenZhang Wei, Jun Cao, Dehua Peng +3
Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typically rely on scalar rewards or repeatedly m
agenttool useself-evolvingbenchmark - arxiv:2609.33378 · cs.RORecursive Harness Distillation across Agents for Robot ManipulationSeungyeon Kim, Junhoo Lee, Minkyu Kim, Baekseung Kim +1
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventi
vision-language-actionmanipulationgr00tagent - arxiv:2609.33377 · cs.LGOptimal Transport Dropout for Structured Predictive UncertaintyGiacomo Lorenzon, Francesco Regazzoni
Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo dropout offers a computationally conven
benchmark - arxiv:2609.33371 · cs.LGAPI Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution BoundaryPatrick Kenney, Hadi Ahmadi, Denis Lusson, Donald Nguyen +1
Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conv
memoryagentllm agentagenticagent system - arxiv:2609.33370 · cs.LGThe Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly DetectionFarhan Shahriyar Hossain, Taufikur Rahman Fuad, Md Abrar Jahin, Md Rizwan Parvez
Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most
benchmark - arxiv:2609.33362 · cs.CLFrom Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTSKangxiang Xia, Xinfa Zhu, HangRui Hu, Kexin Huang +7
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation,
agenticiterative refinementbenchmark - arxiv:2609.33359 · cs.CVWhen Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language BackbonesYueh-Cheng Huang
Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO
benchmark - arxiv:2609.33357 · cs.AIDISCERN: Can AI Agents Work Like Scientists and Guide Discovery?Nan Huang, Mario Tapia-Pacheco, Kun Zhou, Yiming Huang +3
Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical ta
ai agentbenchmark - arxiv:2609.33356 · cs.AILong-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design TasksAnalog Design Bench Team
Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 c
agentagenticbenchmark - arxiv:2609.33354 · cs.ROTraceable Human-to-Humanoid Sign Language BenchmarkingAo Liu, Shengeng Tang, Lechao Cheng, Yanbin Hao +2
Sign data collection is costly, and teleoperation scales poorly, motivating reuse of large video corpora. Humanoid signing requires converting video-derived human motion into robot trajectories while preserving linguistic motion cues. Errors from fitting, human-motion repair, retargeting, robot geom
humanoidteleoperationbenchmark - arxiv:2609.33352 · cs.LGHow Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation ProcedureVictor F. Lopes de Souza, Sébastien Destercke, Abdelhak Imoussaten
Set-valued classifiers, whether derived from precise probabilities and an adapted cost function, from convex sets with a robust inference mechanism, or from conformal methods, are routine options to obtain more robust, trustworthy predictions. However, there is a lack of operational tools to measure
benchmark - arxiv:2609.33351 · cs.LGQuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAGHyojun Ahn, Emily Jimin Roh, Soohyun Park, Walid Saad +2
Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent ret
ragbenchmark - arxiv:2609.33350 · cs.LGKoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution SnapshotsWanfeng Lu, Yutong Zhang, Keyi Zhou, Chenxin Ge +2
Learning population dynamics from temporally sparse, unpaired distribution snapshots is a fundamental challenge in developmental biology. Recent approaches based on neural differential equations and flow matching can interpolate between observed population snapshots, but may struggle to extrapolate
latent dynamicsmemory - arxiv:2609.33347 · cs.LGMultiEcho: An Experimental Science of Learned WorldsMeng Zhu, Airui Zhang
World models can be studied as experimental systems with response laws of their own. We introduce MultiEcho, a framework for estimating these laws through controlled counterfactual interventions, delimiting their applicability, and separately testing their physical correspondence. Across nine simula
world model - arxiv:2609.33342 · cs.LGCalibHyper: Chance-Corrected Relational Hypergraphs for Few-Shot Molecular Property PredictionLinyu Li, Zhi Jin, Yuanpeng He, Dongming Jin +6
Molecular property prediction is central to drug development and materials discovery, but experiments are costly and labeled data are scarce. Context-aware methods use auxiliary assay labels to support few-shot prediction, and recent work supervises property relations with label agreement. However,
benchmark - arxiv:2609.33338 · cs.CVOPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video SegmentationJingchen Ni, Yuji Wang, Shannan Yan, Haoru Li +2
Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent b
agentbenchmark - arxiv:2609.33337 · cs.ROSafe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement LearningBoyang Li, Matthew Kim, Sylvia Lee Herbert
Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates tha
benchmark - arxiv:2609.33336 · cs.LGBeyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation LearningXuanlin Chen, Ziyue Wang, Xunlan Zhou, Yuan-yih Shang +2
Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mit
manipulationworld model - arxiv:2609.33335 · cs.AIDoes Learning to Predict the World Help Agents Act? Auditing World-Model Post-TrainingXinyu Che, Hang Yan, Yanchen Liu, Haochen Liu +4
Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. E
world modelagentpost-trainingworld-model post-training - arxiv:2609.33326 · cs.AIANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information SpacesJerry Wang, Haibo Jin, Xiaopeng Yuan, Peng Kuang +1
Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what
long-contextmulti-agentagent system - arxiv:2609.33325 · cs.CVVisionHOPE: Visual Backbones as Self-Modifying Learning SystemsSiran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao +7
Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an
memory - arxiv:2609.33323 · cs.LGAgentic Multi-Turn Reasoning: A Fairness ApproachThanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren +2
Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit
memoryagentictool usebenchmark - arxiv:2609.33322 · cs.AIRobust Hierarchical Structures for Agentic Document AnalysisRuiying Ma, Yiming Lin, Aditya G. Parameswaran
Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the fact that they are often organized hier
llm agentagentic - arxiv:2609.33319 · cs.AIPhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics ReasoningKecheng Liang, Haoyang Liu, Zexin Chen, Zirong Liu +4
A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the under
benchmark - arxiv:2609.33314 · cs.LGZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph ModelsZhongjian Zhang, Xiao Wang, Busheng Zhang, Bo Yan +4
Zero-shot graph models (ZGMs), which learn transferable knowledge from source graphs and directly apply to unseen target graphs without any adaptation, have achieved promising performance and attracted considerable attention. Despite their proliferation, existing ZGMs are predominantly evaluated on
manipulationbenchmark - arxiv:2609.33311 · cs.ROSocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion GenerationChengqun Yang, Tengjie Zhu, Liang Xu, Fulong Liu +9
Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-tim
embodiedhumanoidwhole-body control - arxiv:2609.33310 · cs.ROCompliantWBC: Whole-Body Compliance for Heavy Humanoids via Force Latent Estimation and Residual Impedance TargetsTan-Dzung Do, Cuc T. Trinh, Tuan Dat Phuong, Chien Le +3
Whole-body compliant control is essential for deploying heavy humanoids under high payload in human-centric environments. Most prior force-aware learning-based pipelines focus on end-effector resistance, per-link upper-body springs, or end-effector stiffness modulation, leaving arbitrary-site pertur
humanoid - arxiv:2609.33306 · cs.CVLoopTrack: A Simple Baseline for Parameter-Efficient Transformer TrackingLiang Peng, Chenxiao Li, Libo Zhang, Xingping Dong +1
Current Transformer-based tracking methods typically stack multiple Transformer blocks with separate parameters to model interactions between the target template and the search region for target localization. These trackers often incur substantial parameter overhead from stacked blocks, making their
memory - arxiv:2609.33304 · cs.CVRelevance Does Not Imply Applicability: Experience Activation for Personal GUI AgentsFuyao Zhang, Xuan Wang, Zherui Li, Jiaming Zhang +2
Personal Graphical User Interface (GUI) agents rely on interaction history to infer what a user wants from ambiguous instructions and to anticipate recurring routines. Existing approaches retrieve task-relevant history and append it to the policy's context, implicitly assuming that experience releva
agent - arxiv:2609.33303 · cs.LGBITS: Rethinking Fair and Comprehensive Evaluation for Irregular Time Series ForecastingKangjia Yan, Linfeng Wang, Tianen Shen, Xiangfei Qiu +6
Despite recent progress in irregular time series forecasting, the field still lacks a unified benchmark for fair and comprehensive evaluation. Existing evaluations are often conducted on a limited set of datasets with inconsistent experimental protocols and predominantly error-based metrics, renderi
benchmark - arxiv:2609.33301 · cs.AIHesitation-Aware On-Policy Distillation for Diffusion Language ModelsJianguo Huang, Lipeng Wan, Yanchen Deng, Bo An
Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the
benchmark - arxiv:2609.33299 · cs.ROAquaWAM: A Dynamics-aware World Action Model for Underwater Embodied AgentsCunhao Zhu, Yifeng Wang, Dongliang Xu, Yunzhong Hou +2
World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as
embodiedaction-conditionedembodied agentbenchmark - arxiv:2609.33298 · cs.LGDirect Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMsFansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan +2
In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model c
post-training - arxiv:2609.33296 · cs.CLBaatCheet: A Multilingual Corpus for Dialogue Translation in Indian LanguagesPriyanka Dasari, Yuvrajsinh D. Bodana, Vandan Mujadia, Arafat Ahsan +2
Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks foc
benchmarkllm-as-judge - arxiv:2609.33295 · cs.AITraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment TracesDehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang +12
An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-s
agentagent systemtool useself-improvementbenchmark - arxiv:2609.33290 · cs.CLCalibration, Not Answer Selection: Distilling Internal Confidence in Reasoning ModelsYadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo +1
Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer
benchmark - arxiv:2609.33289 · cs.AILearning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product MarketsShuze Daniel Liu, Claire Chen, Jiuqi Wang, Thorsten Joachims
Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, w
agentpost-training - arxiv:2609.33288 · cs.ROInformative Viewpoint Selection for Episodic-Memory Embodied Question Answering using Omnidirectional ImagesKaname Kitamura, Asako Kanezaki
Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omn
embodiedagent - arxiv:2609.33286 · cs.CVInfoEdit: Probing Global Layout Reasoning in Infographic EditingCheng Yang, Chufan Shi, Huijuan Wang, Bo Shui +5
Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be
benchmarkevaluation protocol - arxiv:2609.33279 · cs.LGDomain Generalization under Sampling Pattern Shifts in Irregular Time SeriesChanghun Kim, Joohyung Lee, Kwanhyung Lee, Donghwee Yoon +2
Irregularly sampled multivariate time series (ISMTS) are prevalent in real-world applications, where both observation times and available measurements can vary substantially across domains. While recent models increasingly exploit such sampling information for prediction, its robustness under sampli
benchmark - arxiv:2609.33276 · cs.AIChronoFlow: Hierarchical Flow Matching for Irregular Time Series GenerationChanghun Kim, Sunguk Jang, Jeongjun Lee, Juhwan Choi +4
Recent advances in generative modeling have substantially improved time series generation, yet most existing methods either assume a regular temporal grid or focus on feature dynamics under a given sampling structure. This makes them illsuited for generating irregular time series in their native for
benchmark - arxiv:2609.33270 · cs.AIStructured Sparse Memory for Recurrent ReasoningZixuan Zhao, Samuel Wheeler, Neil Getty, Xiaotian Duan +2
Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a comp
memory - arxiv:2609.33269 · cs.ROQ-WAM: 4-Bit Quantization of World Action Models with Action-Subspace ProtectionArash Akbari, Arman Akbari, Jingwu Luo, Yuhao Lei +6
World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but exist
manipulationhumanoidrobotwinmemorypost-trainingbenchmark - arxiv:2609.33268 · cs.AILSTMem: Hierarchical Long Short-Term Online Memory for Large Language ModelsXianglong Shi, Ruijie Yang, Sirui Zhao, Shukang Yin +3
Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to a
memorypersistent statebenchmark - arxiv:2609.33264 · cs.ROVehDyn: A Driving World Model Benchmark for Vehicle DynamicsTianyi Wang, Wangsheng Du, Jiazhou Chen, Tianyi Zeng +11
Video world models are emerging as data engines, action planners, and generative simulators for autonomous driving, but existing benchmarks primarily assess visual fidelity and coarse physical plausibility, providing limited evidence on whether generated driving futures obey realistic vehicle kinema
world modelbenchmarkevaluation framework - arxiv:2609.33261 · cs.CVEngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention ReasoningJinchang Zhang, Yingda Tao, Jiakai Lin, Guoyu Lu
Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state
benchmark - arxiv:2609.33260 · cs.LGCORTEX: A Verified Experience Layer for Generalist AgentsGarapati Keerthana, Manik Gupta
An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and
agentagent system - arxiv:2609.33258 · cs.ROPORTER: Edge-Cloud Residency for Persistent 3D Scene Graph MemoryYue Chang, Yifan Tian, Jiajing Peng, Dazhi Huang +4
Recent task-driven and just-in-time 3D Scene Graph (3DSG) methods reduce per-task representations by constructing or activating only task-relevant information. Yet sparse per-task working sets do not bound onboard memory usage over a robot's lifetime: as tasks change, payloads accumulated for earlie
memoryscene graph - arxiv:2609.33256 · cs.ROActionGround: Training-Free Runtime Refinement of Frozen VLA PoliciesNamai Chandra, Madhur Thareja, Shriram Damodaran, Addison Lin Wang
Vision-Language-Action (VLA) models map visual observations and language instructions directly to robot actions, but they do not explicitly represent the phase structure of manipulation tasks or the rigid-body dynamics governing execution. We present ActionGround, a neuro-symbolic, training-free run
vision-language-actionvlavla policymanipulationopenvlalibero - arxiv:2609.33254 · cs.LGBERT4DTI : BERT-based Model for Predicting Drug-Protein InteractionsThanina Hamitouch, Khadidja Henni, Abdelkrim Arie, Amina Selma Haichour +2
Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce a
benchmark - arxiv:2609.33248 · cs.LGFeedback-Robust AI for Patient Knowledge GraphsMohammed Sameer Syed
Patient knowledge graphs from bedside monitoring should type their relations and state whether the data support their signs. In anesthesia and intensive care, clinicians titrate drugs and ventilation in response to the physiology, so temporal relations mix the patient's response with the clinician's
knowledge graph - arxiv:2609.33244 · cs.AIActiveMem: Dynamic Latent Memory Trees for Long-Horizon AgentsSong-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan
Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-ste
memoryexternal memoryagentagent benchmarkbenchmark - arxiv:2609.33243 · cs.AICodeSkill: Latent Skill Abstraction for Long-Horizon Code AgentsSong-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan +1
Code agents require long-horizon decision-making over complex interaction trajectories. However, existing reinforcement learning (RL) approaches typically optimize behavior at the token level, creating a mismatch between low-level generation and high-level behavioral reasoning. This limitation leads
agentagenticbenchmark - arxiv:2609.33239 · cs.LGSchrödinger--Föllmer Actor--Critic: Diffusion Policy Improvement with Finite-Sample AnalysisYuling Jiao, Lican Kang, Jerry Zhijian Yang, Jincheng Ying
Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schrödinger--Föllmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (K
diffusion policy - arxiv:2609.33237 · cs.ROSurgFlow: 3D Object-Centric Contact Flow for Surgical Robot ManipulationChangwei Chen, Xiao Liang, Yinuo Yang, Nicole Shen +5
Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should m
manipulationhumanoidgrasp - arxiv:2609.33236 · cs.CVPARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence EscalationJi Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao
Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional histor
benchmark - arxiv:2609.33235 · cs.ROLearning with Object-centric Representations of Tactile Interactive Perception for Robot ManipulationXinyi Yang, Zilin Si, Zhuowei Xu, Zeynep Temel +1
Implicit object properties that are difficult to directly infer from vision, such as material, container contents, or softness, can be revealed through tactile sensing and exploratory interactions. However, because tactile signals are transient and sparse, extracting informative tactile events and e
manipulationtactile - arxiv:2609.33233 · cs.AIInspire: Benchmarking Scientific Literature Search for Open Research ProblemsJianrong Ding, Zhengyan Shi, Jianyuan Zhong, Kai Qiu +4
Scientific literature search often begins with an open research problem rather than a known target paper or a fixed candidate set. We introduce INSPIRE, a benchmark for evaluating agents that search prior literature to make progress on solution-redacted research problems. Each instance pairs a resea
agentbenchmark - arxiv:2609.33232 · cs.LGMTLiquid: Enabling Efficient Multi-Task Learning using Liquid Neural Networks for Lightweight Healthcare Monitoring SystemsRachmad Vidya Wicaksana Putra, Fahad Abdul Rauf, Muhammad Shafique
Continuous-time sensing and monitoring with timely and accurate decision-making are critical for many real-world applications. In healthcare monitoring systems, physiological signals are often available or sampled at irregular time intervals, hence requiring continuous-time processing to provide acc
memory - arxiv:2609.33230 · cs.ROAevaScenes: An FMCW LiDAR Dataset and Benchmark for Long-Range PerceptionGautham Narayan Narasimhan, Heethesh Vhavle, Kumar Bhargav Viswanatha, James Reuther +1
FMCW LiDAR measures per-point radial Doppler velocity alongside range, providing a motion cue unavailable in conventional time-of-flight sensors. Exploiting this signal at long range remains understudied. We present an FMCW LiDAR dataset of 575 sequences (57.5K frames) with over 8 million annotated
benchmark - arxiv:2609.33226 · cs.CLBeyond Memory Construction: Rethinking Memory Access for LLM-based Conversational AgentsDonghua Cai, Yongheng Deng, Yifei Wang, Zijun Shen +1
Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While
memoryragrag pipeline - arxiv:2609.33221 · cs.LGRMB: Reward Model Boosting Mitigates Reward HackingJiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou
Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with resp
rlhf - arxiv:2609.33218 · cs.ROScope-WM: Scoped Computation for Efficient Visual World ModelsChunzheng Li, Zesheng Jia, Hongda Zhang, Jiaying Tang +3
Visual world models enable robotic planning by predicting future observations, but dense latent-state propagation and sample-intensive trajectory optimization incur high inference latency and peak memory usage, limiting real-time deployment on resource-constrained platforms. Existing sparse world-mo
world modelaction-conditionedmemory - arxiv:2609.33217 · cs.CVRepFlow: Reciprocal Supervision Improves Generation and Representation in Flow ModelsWeili Zeng, Feng Tian, Shengqi Liu, Yichao Yan
Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the gener
post-training - arxiv:2609.33212 · cs.CLCoLMbo-SV: A Grounded Language Model for Explainable Speaker VerificationMassa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan +2
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-S
evaluation framework - arxiv:2609.33208 · cs.CVWorldAgent: Verification-Guided Agentic Physical World ConstructionCaoliwen Wang, Mengdi Wang, Yige Chen, Zejia Wu +12
Constructing complex physical worlds from language requires coordinating extensive 3D environments, detailed structures and objects at different spatial scales, and interacting physical processes under both stated goals and implicit physical constraints. We present WorldAgent, an agentic framework f
agentic - arxiv:2609.33207 · cs.LGMorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision TransformersRachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, Muhammad Shafique
Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. However, huge parameter counts and complex multi-head self-attention (MHSA) operations make it challenging to achieve high energy efficiency in SViT infere
memory - arxiv:2609.33205 · cs.LGSMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEsYangyuan Li, Weichao Li, Shaowu Pan
High-fidelity simulations of time-dependent partial differential equations (PDEs) are computationally expensive, motivating data-driven reduced-order surrogates for many-query tasks such as uncertainty quantification, design optimization, data assimilation, and optimal control. However, existing sur
latent dynamics - arxiv:2609.33204 · cs.CLTurning Speech Language Models into Multilingual ListenersTolúlopé Ògúnrèmí, Dan Jurafsky, Chris Manning, Ahmet Üstün +1
Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present MULTISPEECHQA, a large-scale, synthetic
benchmark - arxiv:2609.33200 · cs.LGTeach Yourself Where to Look: On-Policy Attention Self-Distillation for ReasoningSafaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi +2
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding
benchmark - arxiv:2609.33197 · cs.ROTAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated ManipulationYongsheng Zhao, Han Gao, Baoping Cheng, Jingyao Tang +6
Vision-Language-Action (VLA) models provide a unified framework for grounding high-level semantic information into low-level robot actions, enabling scalable robotic manipulation across diverse tasks. However, existing VLA models lack explicit mechanisms to disentangle the states and intents of the
vision-language-actionvlavla modelmanipulation - arxiv:2609.33196 · cs.AIAre Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability BoundariesHaiquan Hu, Yuzhu Liang, Weicheng Tang, Yanzeng Li +2
Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boun
benchmarkleaderboard - arxiv:2609.33186 · cs.LGOffline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic RepresentationsAmitakshar Biswas, Yuhan Li, Ruoqing Zhu
Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing
policy evaluation - arxiv:2609.33183 · cs.LGIdentifying Temporal Features within Transcoders for Time Sensitive Factual RecallSanjay Govindan, Yang Song, Maurice Pagnucco
Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (RAG), the internal mechanisms governing
retrieval augmented - arxiv:2609.33181 · cs.AISeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-ThoughtXiaoshu Chen, Xiangyu Wong, Sihang Zhou, Ke Liang +1
Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations
self-improvementself-evolving - arxiv:2609.33180 · cs.LGWhich Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their BenchmarksXiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang
As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite r
self-improvementbenchmarkevaluation framework - arxiv:2609.33177 · cs.RODeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action ModelTianyun Jiang, Wenrui Bao, Bingxin Xu, Yu Tian +1
World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the
manipulationliberoworld model - arxiv:2609.33172 · cs.RODynamic Manipulation with World-Action Models via Counterfactual PlanningSunwoo Park, Wonbin Lee, Seonghyun Jin, Youngmin Kim +2
World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increasingly biased toward the learned continu
manipulation - arxiv:2609.33171 · cs.LGPerturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse ProblemsAbduragim Shtanchaev, Arip Asadulaev, Luiza Labazanova, Aidar Alimbayev +2
Latent diffusion models serve as powerful priors for solving inverse problems in image restoration, such as deblurring, inpainting, and super-resolution. Current methods have a trade-off between generality and efficiency. Solvers that are restricted to a fixed set of degradation operators are fast a
memory - arxiv:2609.33169 · cs.ROWhen Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and ObservabilityXingjian Li, Yi Han, Jianhua Z. Huang
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matter
memory - arxiv:2609.33165 · cs.ROBeyond Tasks: A Vision for Reproducing an Animal-like Behavioral Substrate Using Modern Robot Learning TechniquesSamiyuru Menik, Hemadri Jayalath
Recent advances in robot learning have produced increasingly capable embodied agents. Yet comparatively less attention has been given to a more basic form of competence that animals exhibit continuously: the ability to remain situated, responsive, and behaviorally coherent as physical, environmental
embodiedembodied agent - arxiv:2609.33161 · physics.app-phUncertainty Quantification for the Fission Matrix Method: A Rigorous Mathematical Framework and Computationally Efficient AlternativesValerio Mascolino
The fission matrix (FM) method recasts neutron transport in matrix form, enabling fast, interpolation-based reactor calculations from a pre-computed database of Monte Carlo-derived coefficients. Propagation of the underlying Monte Carlo statistical uncertainty to the FM eigenvalue and eigenvector is
memorybenchmark - arxiv:2609.33158 · cs.CVFOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMsDavid Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins +3
Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. Thi
benchmarkleaderboard - arxiv:2609.33157 · cs.ROTimelyDAgger: Timing-Aware Expert Querying for VLA Policy ImprovementZhixuan Zhao, Peiyan Li, Enhao Zhang, Yueran Tao +9
DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing a
vision-language-actionvlavla policypost-trainingevaluation framework - arxiv:2609.33155 · cs.CLWhere Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?Yuyang Zhao, Xuan Liu, HaoYang Shangm Haojian Jin
Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used tes
post-training - arxiv:2609.33153 · cs.AIWhat Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical ReviewShuyang Zhang
Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries. Reports that a retriever, router, or skill library "improves" an agent may refer to retrieval recall, the success change from enabling a libra
agentllm agent - arxiv:2609.33150 · cs.CLGeneralization Dynamics of LM Pre-trainingJiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computatio
eval - arxiv:2609.33147 · cs.LGCFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free AggregationYanan Ma, Qiyuan Chen, Zihan Fang, Xianhao Chen +1
Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separately does not equal averaging their pro
benchmark - arxiv:2609.33146 · cs.AILiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen TasksEuntae Choi, Sumin Song, Sungjoo Yoo
An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budg
agentllm agentagenticbenchmark - arxiv:2609.33144 · cs.LGBeyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped TransformersJia Liang, Xi Jin, Liangming Pan
Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we call the Matched-Recurrence Looped Tran
self-correction - arxiv:2609.33143 · cs.LGCAME: Company-Aware Evidence-Memory Experts for Interpretable Quarter-Ahead Revenue ForecastingYa-Wen Wu, Meng-Fen Chiang, Kuang-Da Wang, Wen-Chih Peng
Quarter-ahead revenue forecasting requires company-scale numerical accuracy, strict temporal validity, and company-specific interpretation of narrative disclosures. LLMs can distill textual evidence but can produce scale-misaligned forecasts, whereas history-based anchors are stable but miss forecas
memory - arxiv:2609.33141 · cs.AIOn Device Agentic Operation Caches -- Classifier-Centric NL-to-Action GenerationMoghis Fereidouni, Anthony Arnold, Sumit Gulwani, Mark Marron +1
Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resour
agentic - arxiv:2609.33138 · cs.ROMulti-Modal Non-Prehensile Estimation of Physical Parameters via Press-and-Pull TippingSteven M. Hyland, Jing Xiao, Cagdas D. Onal
Recovering physical properties of unknown objects through non-prehensile interaction is challenging because no single manipulation primitive reveals all relevant parameters. Planar pushing couples mass and friction, while conventional tipping cannot recover friction and may fail entirely when low-fr
manipulationgrasp - arxiv:2609.33131 · cs.LGILP-BO: Integer Linear Programming-Based Black-Box OptimizationHyakka Nakada, Shu Tanaka
Black-box Optimization (BO) is a powerful framework for optimizing expensive objective functions or unknown functions with a limited number of evaluations. A central step of standard BO such as Bayesian optimization is the optimization of a surrogate-based acquisition criterion, which is commonly pe
benchmark - arxiv:2609.33127 · cs.LGPolicy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online AdaptationYuheng Huang, Yunpeng Qing, Yixiao Chi, Yilun Kong +1
Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition through value calibration, while genera
online learning - arxiv:2609.33125 · cs.ROTrain Together or Merge Later? Unifying VLA Experts via a Shared Action InterfaceZhizhen Zhang, Yuxia Fu, Zijian Wang, Helen Huang +1
Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining i
vision-language-actionvlamanipulationgr00tliberopost-training - arxiv:2609.33123 · cs.AICompositional Safety Failures in Harness Evolution: Identification and Runtime MonitoringZhixiang Zhang, Zesen Liu, Wai Ip Lai, Hongxu chen +1
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or
agentself-evolvingbenchmark - arxiv:2609.33120 · cs.LGMycelium: A Generalizable Cross-Grid Multi-Task Model for Electrical Distribution SystemsZhengyang Wei, Shourya Bose, Helgi Hilmarsson, Elena Carnio +1
Electrical distribution grid operations require inference across heterogeneous networks from sparse, noisy, and incomplete time series measurements. In this work, we identify challenges and explore solutions towards a unified model that can perform diverse tasks grounded in the physics of the electr
benchmark - arxiv:2609.33119 · cs.AIMedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based ReasoningLang Cao, Binghang Lu, Yuhao Shen, Yue Guo
Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any si
agenticbenchmark - arxiv:2609.33117 · cs.LGECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory ElectrocardiogramsHaitao Li, Chenglin Li, Zhengyao Ding, Ziyu Li +2
Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream
memoryagentllm agenttool usebenchmark - arxiv:2609.33115 · cs.AIModular Discovery of General Game-Playing Algorithms with Large Language ModelsZun Li, John Schultz, Marc Lanctot, Daniel Hennes
General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to disc
multi-agentbenchmark - arxiv:2609.33114 · cs.LGCan Tabular Foundation Models Amortize Statistical Inference?Kai Ye, Shijin Gong, Hongyi Zhou, Valentina Zangirolami +1
For decades, statistical inference has largely been developed one problem at a time. Given a scientific target, such as a treatment effect or a regression function, statisticians design a problem-specific estimator together with a procedure for quantifying its uncertainty. This paper proposes a diff
post-trainingbenchmark - arxiv:2609.33113 · cs.AIParallelPilot: Supporting Coordination and Monitoring in Parallel AI CodingTao Long, Weili Shi, Hussein Mozannar, Maya Murad +1
As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent work. Through a formative study (N=14), we identified PILOT: five supervisory practices for Planning, I
agent - arxiv:2609.33112 · cs.LGSimulation-Free Learning of GP-SDEs from Irregular ObservationsZhidi Lin, Yuhao Liu, Ying Li, Edwin Fong +1
Gaussian process stochastic differential equations (GP-SDEs) provide a flexible Bayesian model for unknown continuous-time state dynamics with uncertainty quantification, but learning and inference from noisy and irregular observations remain computationally challenging. To address this issue, we pr
benchmark - arxiv:2609.33110 · cs.LGD-JEPA: Design-Recoverable JEPA Representation with Swappable Physics DecodersNitin Nagesh Kulkarni, Aashwin Anand Mishra, Yin Yu, Peter Lyu
Joint-Embedding Predictive Architectures (JEPAs) provide a framework for learning compact representations without directly reconstructing high-dimensional observations. However, in parameterized physical systems, learned representations can entangle geometry with operating conditions and task-specif
benchmark - arxiv:2609.33109 · cs.CVToward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language ModelsTuo Liang, Disheng Liu, Nengbo Wang, Vipin Chaudhary +1
Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orient
benchmark - arxiv:2609.33104 · cs.ROVPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation PlanningZhenghao Xiao, Minting Pan, Nantian He, Dongzhan Zhou +1
While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-
manipulationworld modelaction-conditioned - arxiv:2609.33102 · cs.MAORBIT: A Framework for Multi-Agent Safety and Security EvaluationsBen Hagag, William L. Anderson, Srija Chakraborty, Christian Schroeder de Witt
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, from
agentmulti-agentagenticbenchmarkevaluation framework - arxiv:2609.33101 · cs.ROEvolving Dexterous Robots from ScratchZihan Guo, Shuzhe Zhang, Muhan Li, Peiyang Li +1
Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines
manipulationdexterousmanipulator - arxiv:2609.33097 · cs.ROQuery, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language NavigationZhihao Chen, Yiyuan Ge, Ziyang Wang, Pu Cao +1
Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deploy due to their large parameter counts and computational requirements. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence select
benchmark - arxiv:2609.33096 · cs.ROHumanoids for Robot-Assisted Surgery: Bimanual Base Placement and Tool-Mount Optimization via Capability MapsPeihan Zhang, Zekai Liang, Florian Richter, Nikita Thareja +3
Rapid advances in humanoid robotics have motivated growing interest in the application of humanoids for healthcare and clinical tasks. However, it remains unclear how close contemporary humanoids are to meeting the kinematic demands of robot-assisted laparoscopic surgery. In this work, we address th
humanoid - arxiv:2609.33095 · cs.ROREALM: A Coarse-to-Fine Generative Framework for Embodied Reactive ListeningPeizhen Li, Longbing Cao, Yang Zhang
Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the
embodiedhumanoid - arxiv:2609.33093 · cs.LGHow Linear Attention RemembersKichang Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko
Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memor
memory - arxiv:2609.33090 · cs.CVOneSign: Unifying Sign Language Understanding Tasks with One ModelShiwei Gan, Yafeng Yin, Xiao Liu, Desibieer Tuerdaken +2
SLU encompasses a diverse set of tasks, including ISLR, CSLR, and SLT. Although these tasks share basic semantic and linguistic foundations, they are typically addressed with task-specific architectures and training pipelines, which hinders knowledge sharing and requires costly pretraining and finet
benchmark - arxiv:2609.33077 · cs.LGParameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial DimensionalityAdham M. Alkhadrawi, Mohammed A. B. Mahmoud
Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share t
benchmark - arxiv:2609.33075 · cs.AIQureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning SupervisionJanhavi Prabhu, Sahil, Shivam Ashok Shukla, Manoj Tadepalli
Radiological similarity depends on disease relationships and on fine details such as laterality, lobe, severity, size, and certainty. Broad biomedical similarity can overlook these qualifiers, particularly when several attributes vary together. We introduce QureRadEmbed, a 4B radiology-aware encoder
benchmarkevaluator - arxiv:2609.33066 · cs.LGZero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet ValidationVolkan Dağlı, Zerrin Dağlı, Dağhan Dağlı
Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Virtual Machine (EVM). Verifying terminati
memory - arxiv:2609.33061 · cs.AILLM sequential decision making under uncertainty in biochemical domainsMattias Akke, Soojung Yang, Jurgis Ruža, Sathya Edamadaka +1
Large language models (LLMs) are increasingly used to drive scientific discovery. Understanding how LLMs make decisions from new data and memory of the literature is vital before trusting them to design experiments under tight experimental budgets. However, their decision strategies are invisible in
memorybenchmark - arxiv:2609.33053 · cs.ROSwingRL: Adaptive Observation Reinforcement Learning with World-Model Prediction for Cable-Suspended Hoisting ControlGuangming Wang, Xiaoyu Zhang, Yucheng Xin, Wanli Ma +7
Cable-suspended hoisting is widely used to move heavy or bulky payloads that cannot be handled conveniently by rigid pick-and-place systems, for example in crane-assisted construction. Robotic hoisting using flexible cables is challenging because payload motion is underactuated, external disturbance
world model - arxiv:2609.33051 · cs.LGSketchSSM: Write to the Full State, Read from a Compact SketchOmin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer +2
Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires
benchmark - arxiv:2609.33049 · cs.ROIs Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via PerturbationBumgeun Park, Donghwan Lee
Behavior cloning (BC) is vulnerable to covariate shift during closed-loop execution, where small prediction or execution errors can drive the robot toward states poorly covered by the demonstration data. We focus on action-sequence planning, where a policy predicts a finite-horizon sequence of actio
manipulation - arxiv:2609.33044 · cs.CLPinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge LeaderboardsKrishna Chytanya Ayyagari
Modern LLM evaluation assumes that pinning a judge to a fixed model snapshot and decoding at temperature zero yields reproducible verdicts. We show this assumption fails as a property of how LLM-as-Judge is operationalized on cloud serving infrastructure, not of any particular model family. Across f
benchmarkllm-as-judgeleaderboard - arxiv:2609.33039 · cs.AIAgent Safety From Within: Detecting Harmful Trajectories from LLM Internal StatesDifan Jiao, Ashton Anderson
Language model agents can now perform sophisticated sequences of actions via tools and harnesses, which has increased the scope of the damage they can cause. Guard models, however, are mainly built for content moderation and thus are not well-suited to detecting this agentic risk. To address this, w
agentictool usebenchmark - arxiv:2609.33037 · cs.LGByzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal PredictionPrasanjit Dubey, Xiaoming Huo
Retrieval-augmented generation (RAG) lets language models answer questions more accurately by consulting relevant documents. Many valuable collections, such as medical records, cannot be pooled because of privacy rules. Federated RAG leaves each collection with its owner, or node, which scores candi
retrieval-augmentedrag - arxiv:2609.33030 · cs.ROWhat Must a World Model Distinguish for Planning?Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu, Yifan Li +1
World models simulate the consequences of action candidates, but good planning need not preserve every physical distinction required for accurate prediction. We formalize this gap through a hierarchy of mechanism, response, and decision sufficiency. Given a candidate set, the planning query determin
world modelaction-conditioned - arxiv:2609.33023 · cs.AISRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability AgentsYifang Tian, Yingjian Bai, Yifeng He, Zichun Chong +3
Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuratio
agentbenchmark - arxiv:2609.33020 · cs.CVResidual Diffusion Implicit ModelsJoão Guerreiro, Pedro Tomás, Helena Aidos, Jacinto C. Nascimento
Diffusion models achieve state-of-the-art results across multiple tasks. However, in inverse problems, standard initialization from pure Gaussian noise misaligns the generative process with real-world degradations. More recent methods such as diffusion bridges impose strict endpoint constraints and
benchmark - arxiv:2609.33017 · cs.AITrust and Task Completion in the World of Consumer AI AgentsJeroen Olieslagers, Eduardo Pujol, Gal Zahavi, Lukas Ingemarsson +1
Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold b
ai agent - arxiv:2609.33014 · cs.AITCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner ReferenceTzu-Heng Huang, Jet Lin, Eric Lin
Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rar
benchmark - arxiv:2609.33013 · cs.AIThe Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM AgentsSasank Annapureddy, Anjaneya Prasad Thamatani
Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how m
agent memoryagentllm agentbenchmark - arxiv:2609.33012 · cs.AIWhen Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons EstimateWei-Jung Huang
When an agent benchmark compares every pair of leaderboard entries, the number of comparisons can look much larger than the independent evidence behind them: A versus B and A versus C both reuse A. Whether this reuse affects inference depends on what the analysis is meant to describe. If the board a
agentagent benchmarkbenchmarkleaderboard - arxiv:2609.33010 · cs.AIModel-Aware Data Selection from In-and-Out Information InterplayYifan Wang, Xiaomin Li, Yuexing Hao, Dongwon Jung +6
LLMs are effective representations that assimilate vast amounts of knowledge during pretraining, but post-training is necessary for models to reliably access this knowledge and "know what they know." We observe an interesting rank equilibrium between knowledge stored in the weights and the data stre
post-training - arxiv:2609.33007 · cs.ROCAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive ReasoningShivam Aarya, Zhang Xi-Jia, Chengyue Huang, Junhyun Kim +5
Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general
manipulationdiffusion policyfranka - arxiv:2609.33000 · cs.ROTriDrive: Joint Driver, Vehicle, and Road Modeling for Forecasting and Driver MonitoringYuhang Wang, Jingxin Yang, Chuheng Wei, Yuechen Guo +3
Predicting how drivers, vehicles, and road scenes interact and evolve together is central to driver monitoring. Prior work models in-cabin activity or traffic-conditioned driver motion in isolation, motivating joint driver, vehicle, and road modeling with real-time on-vehicle evaluation. We introduc
v-jepabenchmark - arxiv:2609.32996 · cs.CVOracle Gaps in Reliability Coverage: Sampling Noise or Policy Specialization?Mert Onur Cakiroglu, Mehmet Dalkilic, Hasan Kurban
Policies trained from the same base model can appear to solve different problems. An oracle that chooses the best policy for each problem may therefore appear much stronger than any single policy. Selecting the largest estimated success rate also selects favorable sampling errors. We study this effe
post-training - arxiv:2609.32993 · cs.AIX-Tree: Tokenizing Reusable Experience for Efficient Agent GeneralizationSitao Cheng, Xunjian Yin, Zhiyuan Sun, Yuxuan Li +3
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less f
agent - arxiv:2609.32991 · cs.LGWhat Should Data Teach? Moving Bottlenecks Across Circuit, Store, and UseYixiao Chen, Ke Cheng, Jiangtao Guan, Shuo Huang +4
What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing o
long-context - arxiv:2609.32990 · cs.AICertified Long-Horizon Code Agent Evolution via Validation-Gated Skill OptimizationYifan Wang, Hao Cheng, Xiaomin Li, Yuexing Hao +12
Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite
agentself-improvement - arxiv:2609.32984 · cs.CVReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical PerceptionHo Hin Lee, Yuyin Zhou, Yannan Yu, Shi Gu +1
Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-gen
self-improvement - arxiv:2609.32980 · cs.LGFeasible Flow Matching for Graph Reconstruction via Within-Sampling Primal-Dual GuidanceHaoming Chen, Nicolas Zilberstein, Santiago Paternain, Santiago Segarra
Graph reconstruction from partial observations often comes with structural side information, such as degree bounds, triangle counts, or an edge-density band. Prior-Informed Flow Matching (PIFM) reconstructs graphs by transporting a local prior toward the graph distribution, but it provides no mechan
benchmark - arxiv:2609.32976 · cs.LGAdaptive Ensemble Selection for Noisy Labels on Tabular DataFaizaan Ali, Inwon Kang, Oshani Seneviratne
Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label
benchmark - arxiv:2609.32966 · cs.LGSelf-Confirming Superposition Traps in Reinforcement LearningDai Shi, Andi Han, Feng Chen, Yiqun Duan +2
Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirmin
world modeldreamerv3 - arxiv:2609.32965 · cs.AIRelic: From Multi-Agent Collaboration to Persistent Organizational CapabilityHongyi Du, Tianyi Zhang, Weijia Zhang, Yi Yang +8
Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the l
agentmulti-agentbenchmark - arxiv:2609.32964 · cs.AIThe Commit-Abstain Circuit: Why Language Models Hallucinate Instead of AbstainingVy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He +4
Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally
benchmark - arxiv:2609.32961 · cs.LGBeyond Token Savings: A Systematic Study of Context Compression in LLM AgentsRitul Satish, Prasoon Sinha, Akiho Kawada, Neeraja J. Yadwadkar
As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to co
context compressionagentllm agentagentic - arxiv:2609.32960 · cs.AIEnvironmental Impact of Generative and Agentic AI: An in-Depth Analysis and Green SolutionsAbderaouf Bahi, Amel Ourici, Ibtissem Gasmi
The proliferation of generative and agentic artificial intelligence (AI) systems has introduced computational demands whose environmental consequences are substantial yet underexamined. This paper examines the environmental footprint of modern AI systems across energy consumption, carbon emissions,
embodiedagentic - arxiv:2609.32958 · cs.ROLearning Geometry-Aware Virtual Fixtures From Sparse DemonstrationsMaximilian Mühlbauer, Marcella Piacentino, Raffaella Mancino, Paolino De Risi +9
In many teleoperation applications, collecting a large number of demonstrations as required for traditional probabilistic learning from demonstration (LfD) approaches may not be feasible. To still give operators the ability to intuitively create trajectories as Virtual Fixtures (VFs), we propose to
teleoperation - arxiv:2609.32949 · cs.LGOptimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial InventoryMengxiao Zhang, Yingfei Wang, Haipeng Luo
We study online dynamic pricing with censored demand, where an arbitrary inventory level is revealed before pricing and may adapt to past observations, while demand follows an unknown, price-dependent distribution that is stationary over time. For a horizon of $T$ rounds, Xu et al. [2026] achieved $
benchmark - arxiv:2609.32942 · cs.LGThe Key Handoff: Retrieval in Hybrid Language ModelsKaan Kale, Oguzhan Baser, Sriram Vishwanath
A two-hop question makes a language model retrieve twice: once to produce a bridge entity, and once to retrieve with it. Transformers resolve that entity in their early layers. Hybrid models replace most of the attention with a recurrent state, so where the key becomes usable, and where it is spent
memory - arxiv:2609.32940 · cs.LGTwinS-GCN: Spectral conjugate for Spectral Graph Convolutional NetworksChun Hei Michael Chan, Flavia Petruso, Dimitri Van De Ville
Graph convolutional networks propagate information by repeated local aggregation through a graph shift operator; i.e., a $K$-layer network reaches $K$ hops neighborhood. On the one hand, such spreading can lead to oversmoothing. On the other hand, long-range dependencies demand the depth. Transporti
benchmark - arxiv:2609.32939 · cs.LGTheory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM CoordinationLiangqi Yuan, Wenzhi Fang, Shiqiang Wang, Christopher G. Brinton
Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge on targets they must take together, a d
agentmulti-agentagenticagent systembenchmark - arxiv:2609.32935 · cs.ROCommunication-Aware Heterogeneous Graph Learning for Decentralized Multi-Human Multi-Robot Task AllocationZiqin Yuan, Ruiqi Wang, Baijian Yang, Byung-Cheol Min
Multi-human multi-robot (MH-MR) teams combine robotic autonomy with human expertise, but effective task allocation requires coordinating scarce, dynamically available human support with distributed robot execution. Limited robot-robot and human-robot communication further complicates this coupling b
multi-agentbenchmark - arxiv:2609.32934 · cs.LGThe Impact of Stochasticity on the Rashomon Effect in Machine LearningAndrea Apicella, Francesco Isgrò, Andrea Pollastro, Roberto Prevete
Neural network training is inherently stochastic, with factors such as weight initialization leading to distinct models despite comparable predictive performance. This phenomenon is commonly associated with the Rashomon effect, which describes the existence of multiple near-optimal models for the sa
benchmark - arxiv:2609.32933 · cs.LGLast-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPsNam Phuong Tran, Trinh Ha Mai Huynh, Tuyen Pham Le, Van-Truong Nguyen +2
In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has esta
policy evaluation - arxiv:2609.32929 · cs.LGEfficient Dynamic Algorithms for Graph Neural Networks with Non-Linear PropagationKiarash Banihashem, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Silvio Lattanzi +1
Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures withou
benchmark - arxiv:2609.32927 · cs.LGThe Geometry of Logic: Stratification Induces Semantic Structure and Robust ReasoningCristina V. Lopes, Yuangang Li, Justin Tian Jin Chen, Alberto Krone-Martins +2
Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning bene
manipulation - arxiv:2609.32924 · cs.AIDiagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity EvidenceXiao Yue, Guangzhi Qu
Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Cove
evaluation protocol - arxiv:2609.32922 · cs.AIPrecision As You Need: Stochastic Computing Is a Dense Adaptive QuantizerHaoran Jin, Kangqi Zhang, Jirong Yang, Barry Lyu +3
Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conv
post-training - arxiv:2609.32921 · cs.LGAdaptive Latent Capacity for World ModelsIdan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer +1
We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over
world model - arxiv:2609.32918 · cs.ROVision-Language Agents for Active Perception in Optics LaboratoriesRyan Lopez, Sachin Vaidya, Seou Choi, Serena Landers +1
Vision-language models (VLMs) are increasingly being used in scientific workflows, but their ability as agents to directly control laboratory experiments from visual feedback remains underexplored. This capability is important because many laboratory tasks do not naturally provide dense, pre-defined
agent - arxiv:2609.32917 · cs.AIPlanner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent WorkflowsVivek Kumar Singh, Preeti Priyam, Gautam Bhowmick
Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks th
multi-agentagenticbenchmark - arxiv:2609.32915 · cs.AIAgentTell: Behavioural Side-Channel Leakage in Browser-Use AgentsAsif Shahriar, Md Nafiu Rahman, Sadif Ahmed, Farig Sadeque +1
Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user. For example, an agent may learn a user's affiliation a
memoryagentbenchmark - arxiv:2609.32912 · cs.AIProgression- vs Automata-based Anticipatory Monitoring of LTL over Finite Traces (Extended Version)Sarah Winkler, Toryn Klassen, Sheila McIlraith, Marco Montali
When safety-critical systems are developed from a known internal specification, their correctness can be established by model checking. In the frequent case where such a specification is unknown or inaccessible, runtime verification presents an attractive alternative, e.g., to ascertain that autonom
agentic - arxiv:2609.32902 · cs.CLLinger and Lose: Knowledge Collapse in Low-Bit Language ModelsPrashanna Mani Paudel, Shivanand Venkanna Sheshappanavar
Training language models with ternary weights is commonly judged by loss and downstream accuracy, which record only a modest cost relative to full precision. We show that these metrics can conceal a much larger failure. We instead measure knowledge capacity, the factual bits stored per parameter, on
post-training - arxiv:2609.32900 · cs.AIConstraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language ModelsJianchang Su, Wei Zhang
Diffusion language models (dLLMs) predict masked positions in arbitrary order, but their exact constrained decoders still encode constraints as sequential languages, whose state must track every unresolved dependency between positions. For relational constraints this encoding grows exponentially: fo
benchmark - arxiv:2609.32898 · cs.LGWhen Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier DetectionTianyang Zhou, Leman Akoglu
Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is typically motivated by efficiency, we unc
benchmark - arxiv:2609.32894 · physics.app-phComposition-Driven Metal-to-Semiconductor Transition and Enhanced Phonon Transport in B-C substituted ClathrateGhulam Hussain, Dario Massa, Rajibul Islam, Magdalena Birowska +2
Establishing chemical design rules that simultaneously control the electronic structure and thermal transport is a long-sought goal for heat-management and energy materials. Here, we demonstrate that a single B-to-C substitution changes the electron count and simultaneously reconstructs the bonding
benchmark - arxiv:2609.32890 · cs.LGA Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired BenchmarksMaciej Cichoń, Bartłomiej Dmitruk
Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired test, a model must flag a vulnerable functi
benchmarkevaluation protocol - arxiv:2609.32886 · cs.AIStraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM SkillsZeping Liu, Yan Li, Ni Lao, Gil Wolff +1
Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision form
self-evolvingbenchmark - arxiv:2609.32885 · cs.AICan LLMs Predict the Future? A Brier Score Analysis of Prediction MarketsYuanbo Li, Zekun Li, Xiaoyan cong
We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protoc
benchmark - arxiv:2609.32876 · cs.LGMultimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological SimilarityYishu Zhang, Yun Li, Daiwei Zhang
State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outper
benchmark - arxiv:2609.32874 · cs.LGMachine learning for the LHC physics program: a 2025-2026 stocktakeJesse Thaler
The first sentence of this abstract--and the introduction to these proceedings--was authored by a human, but the bulk of this document was generated by an agentic AI system. In this talk, I take stock of machine learning (ML) for the LHC physics program over the twelve months from May 2025 to May 20
ai agentagentic - arxiv:2609.32870 · cs.LGCounterfactual Self-Evolving Agents for Evidence-Grounded ReasoningXing Han, Yuxin Wang, Chen Chen, Wei Dai +7
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be
memoryself-playself-evolving - arxiv:2609.32868 · cs.AIASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU WorkstationsJ. Paul Liu, Uthpala Herath, Andrew Petersen
Traditional scientific computing requires researchers to translate computational intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler state and application logs. We present ASCEND (Autonomous Scientific Computing Engine and Novel Discov
agentai agentbenchmark - arxiv:2609.32867 · cs.CLAre You Sure You're Sure? Two Confounds in a Sycophancy BenchmarkAtharv Gupta, Akshat Jindal, Lavanya Nigam, Aryan Sood
Sycophancy is a language model's tendency to cave when a user pushes back, abandoning a correct answer for the user's. Several benchmarks now measure it by scripting an objection and recording how often the model caves. Because that objection is a prompt template, whatever else the template varies i
benchmark - arxiv:2609.32863 · cs.CVSV2V-RSim: A Comprehensive Benchmark for Self-Selective V2V Cooperative Perception with Near-Realistic DataYulu Wu, Chao Wei, Jujun Cheng, Zhangkai Ni +5
Vehicle-to-Vehicle (V2V) cooperative perception enhances autonomous driving by enabling vehicles to share information beyond their direct line of sight. However, existing V2V datasets are limited by a small number of participating agents, static collaborator selection strategies, and a significant d
sim-to-realagentbenchmark - arxiv:2609.32862 · cs.RORoboFoundry: System-as-Policy Evolution for Self-Learning Embodied AgentsJingsong Liang, Shuhao Liao, Shizhe Zhang, Diyuan Hou +14
A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alon
embodiedliberomemoryagentagenticembodied agent - arxiv:2609.32856 · cs.CVPolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing ImageryZeping Liu, Ni Lao, Weiwei Sun, Gil Wolff +4
Vector polygon generation converts visual inputs, e.g., remote sensing (RS) images, into vectorized polygonal geometries, supporting applications such as autonomous driving, vector map construction, and remote sensing. Early pipelines predict raster masks and post-process them into polygons, which p
benchmarkevaluation framework - arxiv:2609.32855 · cs.ROFINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language NavigationKhang H. Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen, Khanh Dinh Binh +5
Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In pa
embodiedworld modelagent - arxiv:2609.32852 · cs.CLARSM: Auto-Regressive State Machine for Agentic Reasoning CompressionXiafeng Man, Siyuan Ye, Xiaosong Ma
While Large Language Model (LLM)-based agents demonstrate strong capabilities in long-horizon tasks by interleaving reasoning with external environment interactions, the continuous accumulation of context rapidly creates a critical memory bottleneck. Existing memory compression methods rely on task-
memoryautonomous agentagentic - arxiv:2609.32846 · cs.LGSynCo: Learning Cross-Modal Synergy by Contrasting Interaction ResidualsYavuz Yarici, Ghassan AlRegib
Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal
benchmark - arxiv:2609.32838 · cs.LGHamiFormer: Dual-Expert Diffusion Fields with Affine Symplectic MapsHaoxiang Huang, Xiang Liu, Shuwei Wang, Jingheng Ma +1
Predicting smooth dynamics and collisions requires modeling continuous evolution and abrupt state changes. We introduce HamiFormer, a dual-expert diffusion field combining whole-window denoising with residual-corrected Hamiltonian propagation. Their mixed-state feedback attenuates the direct contrib
iterative refinement - arxiv:2609.32837 · cs.ROScanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound NavigationXuesong Li, Shuai Chen, Feng Li, Zhongliang Jiang +2
Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditi
world modelscene graph - arxiv:2609.32835 · cs.AIFinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World PriorsJerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang +4
As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or
ai agentbenchmark - arxiv:2609.32831 · cs.LGUniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal ModelsWanqi Yang, Yuexiao Ma, Mei Xie, Xiawu Zheng +1
Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods ar
long-context - arxiv:2609.32827 · cs.AIImproving LLM Collaboration via Multi-Agent Preference LearningShuo Liu, Xinzichen Li, Tianle Chen, Christopher Amato
Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comp
agentmulti-agentagent systemtool use - arxiv:2609.32825 · cs.AIThe Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own InterfacesTianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao
A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still se
benchmark - arxiv:2609.32813 · cs.LGUSAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built EnvironmentsDongdong Wang, Qingqi Song, Yuzhou Chen, Deepak Balakrishnan +2
Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Questi
benchmark - arxiv:2609.32810 · cs.LGOpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion TrajectoriesAnqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian +8
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across
benchmark - arxiv:2609.32809 · cs.AIOverwhelmed by Choice: Studying LLM Decision Making at ScaleYu-Chi Lin, Aryan Seth, Anshul Aravind, Eugene Lee +3
Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically
long-contextbenchmark - arxiv:2609.32805 · cs.LGDecision-Sufficient State Representations: Measuring and Reducing Write-Time RegretBingyu Shen, Boyang Li
Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. St
memoryagentllm agent - arxiv:2609.32803 · cs.AINutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutritionUttej Kallakuri, Boxun Hu, Ankur A. Butala, Najim Dehak +1
Generative and Agentic IoT systems offer a promising foundation for digital healthcare applications that combine sensing, personalized reasoning, and autonomous interaction in real-world environments. Nutrition assistance is a natural use case, but existing Large Language Model (LLM)-based systems a
embodiedknowledge graphagentagenticembodied agent - arxiv:2609.32802 · cs.LGRe-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream FaultTianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao
One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-w
agent - arxiv:2609.32801 · cs.ROPlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied AgentsJunchi Chen, Changtao Miao, Yuxiao Xiang, Zhenchao Jin +8
Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors a
embodiedembodied agent - arxiv:2609.32798 · cs.MAAdaptive and Resilient Dual-Layer Resource Slicing for Hovering Aerial Backhaul NetworksChuan-Chi Lai, Jen-Hsiang Li
This paper investigates adaptive and resilient dual-layer resource slicing in hovering aerial agent (HAA)-assisted backhaul networks for heterogeneous 5G/6G services, including enhanced mobile broadband (eMBB), ultra-reliable and low-latency communications (URLLC), and massive machine-type communica
agent - arxiv:2609.32795 · cs.AIAgentHabit: Characterizing Distinct Behaviors of Agents on Everyday TasksWoojung Song, Hoyeol Yang, Jeonghoon Shim, Sungjib Lim +3
Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users' preferences and needs. For example, agents differ in whether they ask clarifying questions or se
agentbenchmark - arxiv:2609.32792 · cs.LGUnderstanding and Exploiting Anisotropy in Post-TrainingSamyak Jha, Harshvardhan Saini, Yizhen Liao, Yiming Tang +1
LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry dispro
post-training - arxiv:2609.32791 · cs.AI$T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-TrainingNan Qiao, Yebin Yang, Weinong Wang, Shuning Wang +9
Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return predict
benchmark - arxiv:2609.32790 · cs.LGFedHV: Low-Overhead Hypervolume Weighting for Federated Multi-Objective OptimizationAmirardalan Dehghanpour, Seyed Mohammad Azimi-Abarghouyi, Christopher G. Brinton
Task-wise federated multi-objective optimization (FedMOO) trains a shared model for competing prediction objectives under heterogeneous data, partial participation, and communication constraints. Existing methods commonly derive task weights from gradient or update geometry. This requires task-speci
benchmark - arxiv:2609.32789 · cs.CVHierarchical Frequency-Domain Compression of Implicit Geometric Representations for Large-Scale Point CloudsManlin Yao, Jiabin Liu, Guan Wang, Haixu Liu +1
Large-scale point cloud representations of complex geome tries incur prohibitive computational and memory costs, necessitating compressed implicit representations. To ad dress this, we propose a unified framework comprising im plicit geometric field representation, hierarchical frequency domain comp
memory - arxiv:2609.32787 · cs.AICLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form CompletionGarapati Keerthana, Manik Gupta
Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for Reasoning and Entry), a hybrid workflow t
agenticbenchmark - arxiv:2609.32785 · cs.LGLearning When to Recur: Token-Adaptive Recursion for Imbalanced Ophthalmic Domain Incremental LearningNanxi Yu, Kang Li, Ye Du, Xiaowei Hu +2
Domain incremental learning is essential for adapting ophthalmic deep learning models to sequential clinical domains while preserving diagnostic expertise. Existing domain incremental learning methods predominantly address the domain shift induced by style variations. However, they often overlook th
benchmark - arxiv:2609.32783 · cs.ROCollisionGAT: Controller-Agnostic One-Step Collision Screening for Multi-Agent MotionAlan Debbas, Edwin Meriaux, Gregory Dudek
Before a team of robots moves, each proposed step must be checked for collisions with other robots and with obstacles. We present CollisionGAT, a graph-attention network that reads the current and proposed states of moving agents together with locally relevant stationary obstacles and returns one co
multi-agent - arxiv:2609.32781 · cs.LGCT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language ModelsLong Qian, Bingke Zhu, Jiaqi Wei, Yu Li +2
Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal
post-trainingbenchmark - arxiv:2609.32780 · cs.LGOmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth RoutingLong Qian, Bingke Zhu, Jiaqi Wei, Yingying Chen +1
Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representati
benchmark - arxiv:2609.32779 · cs.ROCopper-Policy: Focus on the Representation for Robust Robot ManipulationZexin Feng, Yixu Feng, Lingyu Xiao, Shang Su +5
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve c
embodiedmanipulationliberorobotwin - arxiv:2609.32778 · cs.AIAgentic Network Traffic MonitoringManuel Tsoukatos, Hayden Jananthan, Jeremy Kepner
As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way that is aligned with the users intent. Auditing an agent's network traffic provides a clear record of
agentai agentagentic - arxiv:2609.32776 · cs.CLOn the Behavioral Traits of LLM AgentsHaokai Zhao, Jie Gao, Yunze Xiao, Xintao Wang +4
Users increasingly describe different AI agents as distinct colleagues to work with. AI personality research aims to quantify such impressions by attributing human-like "traits" to agents. However, existing measures fall short: models' self-reports (S-data) diverge from their actual behavior, while
agentai agentllm agent - arxiv:2609.32771 · cs.LGContinual Learning via Self-Probe GradientsDongkyu Cho, Rumi Chunara, Sungmin Cha
Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-pro
benchmark - arxiv:2609.32770 · cs.CLC-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated TextQing Yang, Zixiang Luo, Zhenyu Mao, Zezheng Wu +5
Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because huma
benchmark - arxiv:2609.32767 · cs.ROCLAP: Closed-Loop Alignment with Pressure for Precise Suction ManipulationYixian Zou, Chongyang Xu, Yuling Xin, Ziliang Feng +2
Stacking and palletising demand precise placement: error left in one layer is inherited by the next, and a flat pad offers no feature to funnel a wrong pose into the right one. Top-down suction suits such dense arrangements, and suction has already been brought into vision-language-action (VLA) poli
vision-language-actionmanipulationgrasp - arxiv:2609.32766 · cs.LGStructuring Relations Among Learning Paradigms via Protocol--Objective--Resource ReductionsJunwei Su, Changjie Wang, Dongyang Chang
Modern machine learning spans supervised, transfer, continual, meta-learning, and related regimes that often reuse the same hypothesis classes, architectures, and optimizers but differ in information access, objectives, memory, adaptation, and sample accounting. This makes it difficult to determine
memory - arxiv:2609.32763 · cs.AIMandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing ThemYicheng Bao, Zhenkun Gao, Xiahui Guo, Mingqian Yang +6
Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts,
manipulationbenchmark - arxiv:2609.32762 · cs.ROAn Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation LearningMino Nakura, Sriram Krishna, Yufei Wang, Shubham Tulsiani +2
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study o
manipulationaction head - arxiv:2609.32761 · cs.CVFrom Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You ThinkHaoru Wang, Qianfan Shen, Kai Ye, Wenzheng Chen +1
Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive
world model - arxiv:2609.32759 · cs.LGThe Extender: A Log-Structured TransformerJakob Eriksson
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each lay
memorylong-context - arxiv:2609.32757 · cs.LGReadout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language ModelsDrandreb Earl Juanico
VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Q
benchmark - arxiv:2609.32756 · cs.LGReuse or Relearn? A Spectral View of Earth Observation Foundation ModelsMehmet Ozgur Turkoglu, Valerio Marsocci, Dominik J. Mühlematter, Dominik Senti +2
Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reus
benchmark - arxiv:2609.32750 · cs.AICUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement LearningXin Yan, Zhengbo Jiao, Jiaqi Liu, Zhenglin Wan +10
Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications. However, conventional deployments replicate an initialized runtime for each independent rollout, even when trajectories use the same softwar
memoryagentevaluator - arxiv:2609.32749 · cs.LGRetrospective Distillation Attribution via Normalized Response SimilarityMinwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang +1
Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However,
post-training - arxiv:2609.32747 · eess.SYPath Invariance of a Quadrotor System under Cyber Attacks with Theoretical GuaranteesHamza Mahmood, Usman Ali, Adeel Akhtar
This paper presents a path-following controller for a quadrotor system to guarantee safe maneuvers, in terms of forward path invariance, in the presence of cyber-physical attacks. We assume that an adversarial agent can control any one of the rotors through a false data injection (FDI) type of attac
agent - arxiv:2609.32746 · cs.LGSelf-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental AnalysisKelvin J. L. Koa, Filip Orestav, Shengqiong Wu, Michael J. Wooldridge +1
While symbolic regression (SR) has been successfully used in science to discover new equations, its use in financial valuation is hindered by several limitations. Whereas the natural sciences provide objectively correct relationships, financial valuation constitutes a distinct class of symbolic disc
memorymulti-agentagent frameworkself-evolving - arxiv:2609.32745 · cs.ROMORPH: Self-Organising Multi-Robot Task Allocation via Neuroplasticity-Inspired Adaptive TopologyXuezhi Niu, Didem Gürdür Broo
Multi-robot task allocation (MRTA) in dynamic environments faces a fundamental tension: effective coordination requires learned structure, but that structure must adapt when conditions change. Existing methods resolve this by assuming prior task knowledge, a utility function, a cost matrix, or a tra
multi-agentbenchmark - arxiv:2609.32743 · cs.LGBenchmarking EEG Foundation Models at Scale: Lessons from 20,000 EvaluationsZhige Chen, Shu Peng, Chengxuan Qin, Rui Liu +3
Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG
benchmarkevaluation framework - arxiv:2609.32740 · cs.CVAnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-MakingZiwei Huang, Qi Gao, Zhe Ji, Yuanyuan Yao +4
Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesT
benchmarkevaluator - arxiv:2609.32737 · cs.LGGradient-Guided Decoupled Adaptation for Geospatial Vision-Language ModelsDongdong Wang, Deepak Balakrishnan, Ravi Srinivasan, Shenhao Wang
Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristic
benchmark - arxiv:2609.32734 · cs.CVREALIS: A Curated Dataset for Studying the Challenges of AI Image DetectionAleksandr Gushchin, Khaled Abud, Georgii Bychkov, Ekaterina Shumitskaya +4
AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for ev
benchmark - arxiv:2609.32731 · cs.AISkillVine: Agent Skill Evolution via Branching ExplorationKaiwei Liu, Jiqian Dong, Liran Dong, Shuai Mao +6
Agent skills encapsulate reusable procedural knowledge that enables LLM agents to perform tasks, and they can be improved automatically using trajectories from interactions with the environment. This is the classic problem of skill evolution. Existing approaches predominately follow a linear evoluti
agentllm agentbenchmark - arxiv:2609.32720 · cs.LGBiasReducer: Adaptive Bias Mitigation for Reward ModelsShuang Liu, Yongliang Miao, Yanguang Liu, Haoyi Xiong +1
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods ei
benchmark - arxiv:2609.32712 · cs.AIMassAlloc Attention: Let Attention Allocate Its Own ComputeJingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi +5
FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized con
long-contextbenchmark - arxiv:2609.32706 · cs.AILearning to Refer: Client-Resolved Generation for Privacy-Aware Language ModelsJeongho Yoon, Chanhee Park, Yongchan Chun, Duong Tuan Thanh +4
Cloud-based large language models (LLMs) require users to disclose plaintext data to service providers, creating privacy risks in sensitive domains. Existing privacy-preserving approaches often trade utility for protection, incur substantial computational or communication overhead, remain vulnerable
tool calling - arxiv:2609.32704 · cs.AICoWindow Attention: Full Causal Coverage Is a Collective PropertyJingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi +5
FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. A
memorylong-contextbenchmark - arxiv:2609.32700 · cs.AICAIRN: Dynamic Fact-Intent DAGs for Multi-Agent ExplorationZuyao Xu, Yuyang Jia, Junwei Guan, Xiang Li +2
LLM-powered autonomous systems have demonstrated promising capabilities in mathematical reasoning, engineering, and cybersecurity. Yet how to organize these systems for effective, reliable, and sustained performance remains an open question. In this paper, we present CAIRN, a fact-intent-driven mult
multi-agent - arxiv:2609.32698 · cs.ROSEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy AdaptationZiwen Li, Hanlue Zhang, Zhenyang Ren, Tianyu Huang +8
Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate.
vision-language-actionvlavla policyembodiedmanipulationself-evolving - arxiv:2609.32696 · cs.LGFlat-Consensus Diffusion for Robust Data Reshaping under Noisy EvaluatorHongyu Cao, Kunpeng Liu, Fei Xie, Sandip Ray
Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data reshaping: generating feature transformati
evaluator - arxiv:2609.32694 · cs.AIIGSD: Environment-Verified Hindsight Self-Distillation for Search AgentsAngqing Jiang, Gaoming Zhang, Chaoqun Zhang, Jianchun Song +4
On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from
agentbenchmark - arxiv:2609.32692 · cs.AIWorld Agent: Can Language Models Keep a World Running?Weixing Chen, Weipeng Zhang, Nan An, Yang Liu +1
World models are moving from generating realistic frames to generating playable worlds, yet whether a delivered world can keep running is not tested anywhere. Existing evaluations stop at generation, at delivery, or at single-step transitions, and each stops at a different point along the way. Corre
world modelagentbenchmark - arxiv:2609.32691 · cs.LGSilent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt InjectionAnimesh Shaw
LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that evaluation is rarely examined. We audit
agentllm agentagentictool usebenchmark - arxiv:2609.32690 · cs.CVMM-OPD: Towards One More Bottleneck Between Perception and ReasoningJintao Tong, Yujing Lou, Zhanming Shen, Jiaqi Gu +5
Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding
benchmark - arxiv:2609.32689 · cs.LGSelf-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy LearningJunyi Wang, Yilin Wang, Wen Wu, Chao Zhang
LLM-based agents are increasingly used for time-series forecasting because they can organise contextual information, perform multi-step analysis, and guide the sequence of actions required to complete forecasting tasks. Most existing agents focus only on the current forecasting instance. However, in
memoryepisodic memoryagentself-evolvingbenchmark - arxiv:2609.32687 · cs.AIDude, Where's My State? Execution Information Requirements for Stateful AgentsNikita Mehrotra, Ashish Tiwari, Priyanshu Gupta, Sumit Gulwani
Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that gene
semantic graphagent - arxiv:2609.32685 · cs.LGPINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural NetworksXu Yang, Mingyang Yu, Jun Zhang, Keqian Li +1
Physics-informed neural networks (PINNs) provide a learning-based framework for solving partial differential equations (PDEs), yet their training behavior can change substantially throughout optimization. Residual distributions, gradient interactions, regional learning difficulty, and model-capacity
benchmark - arxiv:2609.32683 · physics.opticsReconfigurable Linear Optical Transformations in a Single Integrated Multimode WaveguideManoah van der Worp, Jonas Krimmer, Ivo M. Vellekoop, Pepijn W. H. Pinkse +1
Wavefront shaping enables control over optical fields for applications ranging from imaging to photonic information processing. While conventional wavefront shaping relies on free-space systems comprising bulk optical components, compact and fully integrated approaches are comparatively unexplored.
quantum photonic - arxiv:2609.32681 · cs.CVRCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous DrivingLianqing Zheng, Xiaokai Bai, Yixuan Luo, Runwei Guan +5
4D radar provides geometric and motion cues that complement visual semantics, but integrating it into vision-language-action (VLA) models requires both radar--language alignment for semantic reasoning and explicit use of radar measurements for trajectory refinement and selection. To support these ca
vision-language-actionvla - arxiv:2609.32679 · cs.LGThe GUI Is Not the State: Diagnosing State Aliasing in GUI World ModelsDongsheng Liu, Chao Jin, Wenkui Yang, Hejin Wang +6
GUI World Models (GUI-WMs) are increasingly used to predict future states for agent planning and simulation, yet most existing formulations condition only on the current GUI observation and action. We identify state aliasing, where the vis- ible interface omits transition-relevant environment state,
world modelagentbenchmark - arxiv:2609.32677 · cs.LGWhen Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive WorldsKe Wang, Zijie Zhao, Zhiyi Yuan, Changlun Li
Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. W
self-improvingself-improvement - arxiv:2609.32674 · cs.AIExpected Reasoning-Step Return Unifies On-Policy Learning from Rewards and TeachersQiangqiang He, Jin Li
On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning ste
benchmark - arxiv:2609.32671 · cs.CVConcepts Complement Dense Semantics: Learning Compact Sparse Spaces for Text-Image RetrievalYoonseo Kim, Jungwoo Choi, Cheonyoung Park, Youngwook Kim +2
Cross-modal retrieval has been advanced by vision-language pre-trained models that encode images and texts into a shared dense embedding space. While dense representations effectively capture overall semantic similarity, they often obscure fine-grained visual-textual information needed for precise c
grasp - arxiv:2609.32663 · cs.LGDistance-KV: Exploiting Relative Distance for Efficient Long-Context InferenceXianpeng Shang, Canbin Huang, Jiang Li, Tian Lan +3
The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval c
memorylong-contextlong contextbenchmark - arxiv:2609.32662 · cs.AIAsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software EngineeringKaituo Zhang, Zhen Xiong, Zhimeng Jiang, Mingyu Zhong +7
Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distributed collaboration, mul
multi-agentagent systembenchmark - arxiv:2609.32661 · cs.LGEquivariant Neural Primal-Dual Assignment for Maximum Common Edge SubgraphsJiaqing Xie, Yanchao Li, Zhuo Yang, Yuxin Wang +2
Maximum common edge subgraph (MCES) matching finds a partial vertex correspondence between two labeled graphs that preserves as many labeled edges as possible. Molecular similarity search requires matching many graph pairs, making the cost of repeated queries important. The strongest baseline attain
benchmark - arxiv:2609.32660 · cs.CVInterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table ReasoningHanqian Li, Sirui Huang, Chen Ling, Jungang Li +11
Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-
benchmark - arxiv:2609.32659 · cs.LGQuantization-Aware Pre-Training with Constrained Empirical Weight DistributionNingfeng Yang, Tor M. Aamodt
Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this
memory - arxiv:2609.32658 · cs.AIContract Memory Compiler: Resolve, Then TraverseZhi Song, XiMing Xing, Chunhan Li, Weian Mao +8
External memory lets language-model agents answer questions about histories too long for the answer model's context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We s
memoryexternal memory - arxiv:2609.32657 · cs.AIWorld Models with Predictable Long-Horizon MarginalsYuhao Du, Shunian Chen
Accurate one-step predictions do not ensure that a world model's rollouts retain the data distribution. We make the model's decoded stationary law explicit by learning a decoder of a fixed Gaussian reference and constraining the behaviour-averaged transition to preserve that reference. For controlle
world modeldreamerv3 - arxiv:2609.32656 · cs.LGMixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays OffIbram Abdelmalak, Mischa Putzke, Jungmin Choi, Tom Hanika +2
Multivariate Time Series Forecasting (MTSF) models that mix information across channels assume that the past of one channel carries information about the future of another. Yet they are evaluated on a small fixed set of standard datasets whose cross-channel structure is rarely examined. We ask two q
benchmark - arxiv:2609.32653 · cs.CVFrom Knowing to Abstaining: Bridging the Representation-Action Gap in Vision-Language ModelsJialuo He, Huangxun Chen
The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often con
benchmark - arxiv:2609.32652 · cs.LGPrediction Limits and Koopman Closure of Geometry-Induced Soft State AbstractionsMohit Kumar, Somayeh Kargaran
We study when geometry-induced soft state abstractions admit accurate finite-dimensional linear dynamics. Each state is represented by simplex-valued coordinates obtained from class-specific Kernel Affine Hull Machine (KAHM) reconstruction scores, and a matrix is used to predict the next-state coord
benchmark - arxiv:2609.32645 · cs.AIFrom Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous DrivingYiyao Wang, Pei Liu, Fangzhou Liu, Jun Ma
Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some
scene graph - arxiv:2609.32643 · cs.AIBusiness Compromise Detection with Agentic AI and LLM-driven Knowledge DiscoveryDiego Palma, Kyu Bin Kim, Zhen Han, Allbright Dsouza +1
Detecting compromised business ad accounts is a challenge in digital advertising, as attackers exploit hijacked accounts to launch fraudulent campaigns. Large Language Model (LLM) agents show promise for integrity enforcement, but hallucinated mistakes on hard cases create business friction. In a st
agentautonomous agentagenticbenchmark - arxiv:2609.32641 · cs.LGExtremely Fast and Compact Binary Graph Representations via Randomized Operator SketchingSrajan Agarwal, Megha P, Bikas C Das, Zakaria Laskar +1
Graph neural networks typically rely on dense, floating-point node representations, which can impose substantial memory and computational costs. Binary graph hashing offers an alternative by encoding node information as compact bit strings. However, existing approaches either sacrifice global topolo
memory - arxiv:2609.32638 · cs.AICan Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration StudyGrandee Lee, Wang Yue
Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary m
benchmark - arxiv:2609.32635 · cs.AITrust the Brand, Lose Control: How Identity Hijacks LLM Agent OrchestrationXutao Mao, Rui Qian, Linghan Chen, Yudong Gao +4
LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents dis
agentllm agentagent systembenchmark - arxiv:2609.32634 · cs.ROPF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action ModelsYunpeng Qing, Yilun Kong, Sixu Lin, Ming Zhou +8
Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progr
vision-language-actionvlamanipulationliberorobotwin - arxiv:2609.32631 · cs.AISWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering AgentsChaoqun Cui, Hao Zhou, Meiqi Chen, Fandong Meng +1
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an a
agentevaluator - arxiv:2609.32630 · cs.CLExpVoyager: Direct Experience Navigation for Dynamic Agent Skill SynthesisKwangwook Seo, Dongha Lee
Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable pro
agentllm agentself-evolving - arxiv:2609.32626 · cs.ROWorld SLAM Model: Joint World Modeling for SLAM and NavigationMinghui Qin, Yijun Yuan, Weicheng Zheng, Kenan Li +6
We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent
embodiedworld modelmemorypersistent memory - arxiv:2609.32622 · cs.CLDoes CoT-Pass@k Really Check the CoT? A Multilingual Mathematical AuditTarık Tuna Taşaltı, Burcu Hüdaverdi, David Semedo
Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain before it counts. Its value rests entirel
benchmarkllm-as-judge - arxiv:2609.32619 · cs.LGIntuition vectorsShahar Haim, Daniel C. McNamee
Large self-supervised vision models learn representations that support scene segmentation and the semantic decomposition of physical objects. We ask whether their representational geometry supports transfer to visual reasoning problems without any task-specific fine-tuning. We hypothesized that rela
benchmark - arxiv:2609.32617 · cs.AIThe Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language ModelsQingjia Huang, Yakai Li, Jianguo Wu, Qihang Zhou +4
Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data
post-trainingbenchmark - arxiv:2609.32616 · cs.AI"You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely AccusedXutao Mao, Rui Qian, Longxiang Wang, Xinjian Yi +5
LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, an
agentllm agentagenticbenchmark - arxiv:2609.32613 · cs.LGLearning the Graph and the Embedding Together: Classifier-Independent Rewiring for Heterophilic Node ClassificationHarshit Kumar, Sujan Chakraborty, Priyanka Saha, Pritam Kar +1
Graph neural networks lose much of their advantage on heterophilic graphs, where connected nodes often carry different labels. Graph rewiring is a popular remedy, but rewiring methods are usually evaluated with a single classifier, which makes it hard to tell whether the gains come from the new topo
benchmark - arxiv:2609.32612 · cs.CVLevy-Driven Correspondence Estimation for RegistrationQianliang Wu, Jiaqi Yang, Wankou Yang, Le Hui +3
Finding reliable point correspondences is difficult when point clouds have low overlap or undergo non-rigid deformation. Iterative refinement can correct uncertain matches, but costly network evaluations limit the number of updates. We present LevyMatch, a Lévy-driven method that uses random jumps t
iterative refinement - arxiv:2609.32600 · cs.AICUA-SWE: When Computer-Use Agents Meet Visual Software EngineeringPrince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan +5
Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studi
benchmark - arxiv:2609.32597 · cs.LGNot Every Term Adds New Structure: Sobolev Novelty for Symbolic RegressionBoxiao Wang, Kai Li, Yuheng Jing, Tianyi Liu +3
Symbolic regression (SR) aims to discover compact and meaningful mathematical equations from data, but searching the vast combinatorial space of symbolic structures remains challenging. Existing methods typically guide this process using expression-level objectives, such as fitting error, which asse
benchmark - arxiv:2609.32595 · cs.RORECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot NavigationIncheol Cho, Jintae Park, Jinkyu Kim, Jungbeom Lee +2
Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However,
quadruped - arxiv:2609.32594 · cs.AIMA-FPPO: Multi-Agent Flow-Pretrained Policy OptimizationGuowei Zou, Haonan Chen, Haitao Wang, Beiwen Zhang +2
Multi-agent flow policies learn cooperative behavior from fixed offline datasets, but often struggle to complete tasks in situations not covered by the offline data. In these situations, agents must both adapt to changes in the environment and coordinate with one another, yet action patterns learned
multi-agentonline learningevaluation protocol - arxiv:2609.32592 · cs.CVSPACE: Sparse Predictive Attractor via Counterfactual Eviction for Streaming Video MemoryHongjin Niu, Weizhan Zhang, Shuo Bao, Jiahao Wang +3
Fixed-capacity streaming video memory requires repeated eviction decisions whose effects accumulate over time. Yet existing policies are evaluated primarily in terms of retained information or downstream accuracy, leaving how repeated updates alter the futures supported by memory largely unexamined.
memory - arxiv:2609.32591 · cs.ROThink Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPCYi Xian Goh, Sze Jue Yang, Hao Luan
Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well
world model - arxiv:2609.32590 · cs.CVRetrieved but Not Delivered: Multimodal Memory Delivery for Long-Term AgentsYuhang Jiang, Qingwei Liao, Kaize Yin, Xingling Liu +2
Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery, and a controlled deco
memoryagentbenchmark - arxiv:2609.32584 · cs.AIEMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round RetrievalJinlan Liu, Hongliang Sun, Yong Wang, Bolin Zhang +3
Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks. However, existing memory systems struggle to utilize continuously evolving historical information, as they often rely on static memory representations and single-round retrieval strategie
memoryagentllm agent - arxiv:2609.32581 · cs.LGHERO-MoE: Historical Expert Routing with Scale-Preserving FusionJunxiang Qiu, Zhengsu Chen, Xinting Hu, Shuo Wang +5
Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation a
memory - arxiv:2609.32579 · cs.LGDimPO: Dimensionality Reduction for Attention using Preference OptimizationVojtěch Lanz, Yufei Cui, Prasanna Parthasarathi
A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal tha
long-context - arxiv:2609.32577 · cs.LGGroupwise Agentic Grading and Advantage Redistribution for Code Agent RLJinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma +6
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and
agentagentic - arxiv:2609.32576 · cs.ROAssisting for Open-Ended Tasks: Goal-Oriented Shared Autonomy as a Particle FilterMengxue Fu, Ethan Xu, Sam Iyer-Singh, Yinlong Dai +2
A common approach for shared autonomy blends human inputs with autonomous assistance based on the human's likely goal. However, most existing approaches assume that a static set of possible goals is known a priori, which limits the use of such methods in unstructured assistive settings. We instead i
manipulation - arxiv:2609.32574 · cs.AICUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal ConversationsYulin Hu, Yanyan Zhao, Zimo Long, Xing Fu +5
Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Exis
memorybenchmark - arxiv:2609.32573 · cs.LGTrapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy DistillationJinqian Chen, Jihua Zhu, Chang Liu
On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated collaboration? We study federated OPD and fin
benchmark - arxiv:2609.32570 · cs.LGCAESAR: Clustering via Autonomous Embedding-Space Agglomerative ReorganizationIlan Bacry, Rémi Devaux, Antoine Jardin
Clustering algorithms that operate on nearest-neighbor graphs, such as FINCH (First Integer Neighbor Clustering Hierarchy), depend heavily on the quality of the embedding space they are given. However, pretrained vision and language model embeddings are not optimized for this purpose. We propose CAE
benchmark - arxiv:2609.32569 · cs.CVOmniSmartHome: A Multimodal Reasoning Benchmark for Smart-Home AgentsJihoo Jung, Suho Yoo, Jeongsoo Choi, Hyebin Cho +4
Smart-home assistants are expected to handle diverse, realistic requests that arise in daily life. In such interactions, users often rely on the surrounding multimodal context-pointing at objects or referring to what they see or hear, leaving their requests underspecified in language alone. Existing
memoryagentbenchmark - arxiv:2609.32566 · cs.LGCompositional Objectives: Learning Structure in StructurePranavchandra Vivekananda, Sumukh Bettadapura, Ajan Subramanian
Intelligence is defined in many ways. One of these definitions defines intelligence as the pursuit of learnable novelty. However, learnable novelty can be meaningless without the ability to compose the learned structures to take action and achieve goals. Learnable novelty builds on epiplexity, which
benchmark - arxiv:2609.32564 · cs.AIProTTT: Learning to Learn Semantic User Memory with Test-Time TrainingSejun Park, Hyoungjo Bhang, Hyein Jeong, Yohan Jo
Personalization requires language models to capture user-specific knowledge from a growing user history. Existing context-based approaches incur increasing inference costs as user history accumulates and rely on separate retrieval or summarization stages, while parametric-based approaches often requ
memorybenchmark - arxiv:2609.32560 · cs.CLHow to Reduce Whisper HallucinationHusein Zolkepli
Whisper is still what runs in production: one permissively licensed checkpoint, 99 languages, no per-language tuning. But it writes sentences nobody said. On 42 clips of pure room tone, whisper-large-v3 emits words on 61.9% of them and emits something on 100%. The usual response is to distil a stude
benchmark - arxiv:2609.32550 · cs.ROAre Vision-Language-Action Models Robust to One-Step Observation Perturbations?Shojiro Yamabe, Jun Sakuma
Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation c
vision-language-actionvla - arxiv:2609.32547 · cs.LGPrioritizing Repeated LLM Evaluation for Hidden Failure DiscoveryKeita Broadwater, Akin Broadwater
Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important failures. A prompt that produces no failure
benchmark - arxiv:2609.32544 · cs.AIPorimon: An LLM-Based Pokémon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented GenerationDongyin Zhuo, Fengjunjie Pan, Nenad Petrovic, Alois Knoll
In this paper, we use Pokémon Battles as a case study to investigate how to improve the performance of LLM-based agents in tasks that require opponent-aware planning without additional fine-tuning. We propose Long/Short-Term Knowledge Augmented Generation (LSTKAG), a mechanism that enables LLM-based
agent - arxiv:2609.32537 · cs.CVRIPE-MambaSpike: Resolution-Independent Spiking-State-Space Interfaces for Parameter-Efficient Event-Based VisionMd Muhiminul Islam, Shoaib Ahmed Dipu, Sayeed Shafayet Chowdhury
Spiking-Mamba hybrids reach strong accuracy on event-based vision, but existing designs often require tens of millions of parameters. Much of that cost comes from how the spiking front-end is connected to the state-space backbone rather than from the hybrid architecture itself. In a representative m
benchmark - arxiv:2609.32536 · cs.AIDo Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice AgentsYanjie Zhang, Nanchen Hu, Yushi Sun
Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch
manipulationagentpost-trainingbenchmark - arxiv:2609.32534 · cs.LGDepthBench: Measuring How Residual Connections Enable More Computational DepthKeyu Wang, Yangyi Huang, Jiale Kang, David González-Martínez +2
Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g.}, mHC, AttnRes) to enable better infor
benchmark - arxiv:2609.32533 · cs.AILLMAdBench: A Human Preference Benchmark for Advertising in LLM ResponsesRui Ai, Yuqing Liu, Sitao Qiu, Yun Qiao +13
Inserting advertisements (ads) into consumer-facing LLM output is emerging as a new business model, but there is little shared evidence on how such ad insertion should be evaluated or how it affects user preferences. We introduce LLMAdBench, a human-preference benchmark for studying advertising in L
benchmark - arxiv:2609.32528 · cs.AIFail Loudly: An Auditable Runtime for Agentic Data AnalysisHanxu Yan, Langxuan Deng, Zhengle Wang, Yibo Wang +1
Large language models (LLMs) have enabled data-science agents to automate multi-step analyses over heterogeneous files. However, incorrect choices regarding data sources, scope, or statistical definitions often lead to silent errors: computations execute successfully but produce plausible yet incorr
agentagentic - arxiv:2609.32527 · cs.LGAmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared ResponsesSikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen +1
Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolu
benchmark - arxiv:2609.32525 · cs.LGDoes Transolver really need a Transformer?Shizheng Wen, Siddhartha Mishra
The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empiri
memorybenchmark - arxiv:2609.32522 · cs.AIBeyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken ConversationsWenxu Jia, Xize Cheng, Zihan Zhang, Dongjie Fu +4
Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational conte
memoryagentagentic - arxiv:2609.32521 · cs.AIMemAgent: Learning to Manage Heterogeneous Memory Providers for LLM AgentsYongxian Wei, Yilin Zhao, Runxi Cheng, Xinrui Chen +4
Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., tra
memoryagent memoryagentllm agentagenticbenchmark - arxiv:2609.32520 · cs.AIWhen Users Change Their Minds: Measuring and Repairing Intent Drift in LLM AgentsYanjie Zhang, Bowen Cao, Zixin Chen, Yushi Sun
LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that con
llm agentbenchmark - arxiv:2609.32518 · cs.CVUnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than DistillationYoussef Mansour, Enis Simsar, Fadime Sener, Markos Georgopoulos +3
Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In th
memory - arxiv:2609.32517 · cs.LGLocalProp: Neuro-Localized Memory-Efficient BackpropagationDiana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu
The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly loca
memorybenchmark - arxiv:2609.32516 · cs.LGREFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations CentersHuimin Chen, Quan Long, Yanhao Wang
Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organization-specific, rapidly evolving SOC o
llm agentagent framework - arxiv:2609.32514 · cs.AIFrom Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace DiagnosisShu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen +6
LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally
agenticbenchmark - arxiv:2609.32512 · cs.LGWhat Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local ResponseEnzo Nicolás Spotorno, Josafat Leal Filho, Antônio Augusto Fröhlich
Models of vehicle dynamics learned from logged states and commands complement physics-based models, and latent world models, which predict in a learned representation, are used to plan and train controllers in other domains. Vehicle controllers are usually specified in physical terms: costs, limits,
world modelaction-conditioned - arxiv:2609.32511 · cs.AILearning from Others, Acting for You: Cross-User Memory Sharing for LLM AgentsJinming Hu, Haodong Zhao, Qi Jia, Die Chen +3
Large language model (LLM) agents serving different users often solve related tasks, yet separate user histories can leave reusable experience inaccessible to other agents. Pooling memories expands access but risks transferring preferences that conflict with the receiving user's requirements. We int
memorymemory architecturellm agent - arxiv:2609.32499 · cs.AILearning an Anchored Prompt Space for Continual Adaptation of Large Language ModelsRongguang Ye, Zhan Zhuang, Yichen Wu, Ming Tang +1
Continually adapting large language models requires acquiring new knowledge while preserving previously learned capabilities. Jointly adapting model parameters and task-specific soft prompts offers a promising solution, but faces two key limitations: historical prompts may become less effective as t
benchmark - arxiv:2609.32498 · cs.AIDAAF: From Failure Localization to Editable System Assets in LLM AgentsXiaoyang Yuan, Qi Liu, Yubin Ruan, Xinyi Mou +8
Deployed LLM agents increasingly rely on persistent, versioned system assets such as routing rules, knowledge segments, prompt instructions, and reusable skills. Failure-localization methods can identify where an error manifests in an agent or execution trace, but repair requires a different decisio
agentllm agent - arxiv:2609.32496 · cs.CLLocally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop ReasoningBohao Chu, Hendrik Damm, Qianli Wang, Hui Wang +3
Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound yet globally insufficient. Local soundne
benchmark - arxiv:2609.32495 · cs.AIHearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?Jiahong Dai, Zhuochen Yang, Pengyang Shao, Kelvin Ng +4
An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across si
agentbenchmark - arxiv:2609.32493 · cs.LGSoFT: Soft Targets for Generalizable LLM Fine-TuningHuihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi +6
Distillation enables student language models to acquire new capabilities from expert teachers. However, integrating knowledge from multi-teacher, multi-domain demonstrations into a single student remains challenging. We study supervised fine-tuning (SFT) in this setting, where students must acquire
agentic - arxiv:2609.32492 · cs.AIBeyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM ProgramsHaoran Shou, Haoyue Liu, Yu Huo, Kun Zeng +1
Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design
benchmark - arxiv:2609.32490 · cs.AIRepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent SystemsYuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang +1
LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool us
multi-agentagent frameworkagent systemtool usebenchmark - arxiv:2609.32488 · cs.LGWhen Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual ProjectionsMaojun Sun, Yancheng Yuan, Jian Huang, Ruijian Han
Dense retrieval powers retrieval-augmented generation, semantic search, and question answering, yet the theoretical basis for choosing between shared and dual query-document projections remains unclear. We introduce a bias-variance theory for low-rank bilinear scoring. Shared projections induce posi
retrieval-augmented - arxiv:2609.32485 · cs.LGWhat Should Federated LoRA Share? FedSAIL via Input-aware Subspace AlignmentJunye Du, Shuaida He, Long Feng
Federated low-rank adaptation (LoRA) requires identifying an update structure that is shared across heterogeneous clients. Prior work reports strong similarity among trained LoRA projection matrices across clients; however, such agreement may be largely induced by common initialization and collapses
benchmark - arxiv:2609.32481 · cs.LGJEPA Learns What the Mask Leaves UnrecoverablePeng Xie, Amr Alanwar
Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose suppor
v-jepa - arxiv:2609.32474 · cs.CLPC-SubMax: Efficient Prompt Compression via Regularized Submodular MaximizationZiyi Zhang, Shuang Cui, Haotian Zhang, Xiaoyu Wang
While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle'' phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, m
long-contextbenchmark - arxiv:2609.32473 · cs.AIVPEvolve: A Self-Evolving Virtual Process Engineer for Computational LithographyTianyi Li, Wenxuan Dong, Donger Luo, Nan Wang +4
Optical proximity correction (OPC) recipes grow as engineers add local rules to repair newly discovered lithography hotspots. Each correction can interact with existing rules, while lessons from commercial-tool trials remain scattered across code and logs. \system combines a Virtual Process Engineer
self-evolvingbenchmark - arxiv:2609.32472 · cs.CLAdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep ResearchKailin Jiang, Lei Liu, Jian Xi, Yangqi Chen +5
Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work
ragbenchmark - arxiv:2609.32471 · cs.ROGlowTact: Simple and Compact Vision-Based Tactile Sensing with High Sensitivity and Spatial ResolutionYuxiang Ma, Megha Tippur, Pengfei Ye, Sandra Q. Liu +3
Vision-based tactile sensors (VBTS) provide rich contact information for robotic manipulation, but existing designs can be hard to simplify and adapt to the size and constraints of humanoid fingertips. We introduce \textbf{GlowTact}, a pressure-responsive vision-based tactile sensing mechanism that
manipulationhumanoidtactile - arxiv:2609.32470 · cs.LGOn the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning ModelsShuoyuan Wang, Beier Luo, Hao Zeng, Chengyao Yu +4
Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distributio
benchmark - arxiv:2609.32469 · cs.LGPULSE: Identifying Demonstration-Utility Features with Sparse AutoencodersChenduo Hao, Chuanbao Gao, Pinjun Zeng, Jingze Zhu +3
In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce
benchmark - arxiv:2609.32467 · cs.LGBison: Cross-Dataset Learning for Unseen-Compound Perturbation PredictionYunfan Liu, Kasra Ghorbani, Yufei Huang, Zicheng Liu +5
Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark eight datasets with 16,771 compounds, wit
benchmark - arxiv:2609.32465 · cs.LGPolyStepOR: Learning to Decide Without Optimal DecisionsViet The Nguyen, Gunther Gust, An Thai Le
Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed optima and extends to in-constraint predicti
benchmark - arxiv:2609.32464 · cs.LGAdapting Nonstationary Multi-output Gaussian Processes to Bayesian OptimizationZikai Xie
Multi-objective Bayesian optimization (MOBO) commonly relies on independent Gaussian processes (GPs) with stationary kernels, limiting its ability to represent nonstationary structure and share information between objectives. However, expressive nonstationary GPs do not necessarily make reliable BO
benchmark - arxiv:2609.32462 · cs.CVCan Motion-Language Models Ground Structure? STRIDE for Evaluating the EvaluatorsLixing Tan, Qing Xia, Yuting Guo, Shuai Li +1
Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evalu
benchmarkevaluatorevaluation protocol - arxiv:2609.32460 · cs.CVREMEDY: How Far Is Video Generation from Medical Education World Models?Lixing Tan, Yanghao Zhou, Qing Xia, Yuting Guo +2
Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current gene
world modelbenchmark - arxiv:2609.32458 · cs.CLStreamlined Reflective Evolution for Task-Adaptive Self-Refinement PipelinesXiaofan Zhou, Lu Cheng
Reflective prompt optimization improves large language model (LLM) systems without updating model weights, but fixed architectures constrain how self-refinement is organized. We introduce Workflow-Designing Agents (WDA), a framework for streamlined reflective evolution of task-adaptive self-refineme
agenticself-refinementbenchmark - arxiv:2609.32457 · cs.LGWrite Back the $Δ$: Revisiting the Same Tokens with Fresh RepresentationsWencheng Ye, Anning Hu, Xiangdong Zhang, Tianyi Wang +4
Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computation, or modify the residual stream usin
benchmark - arxiv:2609.32456 · cs.CVSeeing Parts, Reasoning about Worlds: Visual Inference under Partial ObservationWei Wang, Wenqiao Zhang, Yutong Lin, Jun Xiao +1
World modeling under partial observation requires reasoning about the complete worlds that remain compatible with limited visual evidence. Occluded objects and unseen regions can leave several world states possible; additional views can exclude alternatives and strengthen the conclusions supported b
world model - arxiv:2609.32453 · cs.RODRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation PoliciesXinyu Zhao, Yixiang Shan, Tao Yang, Runyu Lei +4
Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation win
manipulationmemorymemory modulepost-training - arxiv:2609.32449 · cs.CLSelf-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual ReportsPhongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demon
benchmark - arxiv:2609.32448 · cs.AIForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion DistillationJunming Liu, Jicheng Wang, Yifeng He, Hao Chen +1
Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through distillation without giving up their native
benchmark - arxiv:2609.32447 · cs.LGLength-Independent State Tracking Under a Parallel ScanJulien Brandoit, Arthur Fyon, Thomas Braipson, Tom Clara +4
Learning robust and scalable finite-state tracking is fundamental to sequence processing. While linear recurrent neural networks (RNNs), linear attention, and state space models enable scalable parallel training through affine recurrences, their theoretical expressivity guarantees assume idealized a
benchmark - arxiv:2609.32444 · cs.LGRethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct ItTianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang +3
We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To acco
benchmark - arxiv:2609.32434 · cs.AIFrom Latents to Wires: Surgical Post-Editing on Large Language ModelsJiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu +3
Given a large language model (LLM), can whoever holds the weights name a semantic target (e.g., the model's identity), locate the model components that produce it, and edit them so that the target no longer appears while other capability is preserved? We call such an edit on a trained model a post-e
post-training - arxiv:2609.32430 · cs.AIMulti-Agent System Search via Active Substructure-aware Policy OptimizationBeicheng Xu, Bowen Fan, Weitong Qian, Lingching Tung +1
LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically
agentmulti-agentagent systembenchmark - arxiv:2609.32429 · cs.LGPrismQuant: Optimal Null-Space Rotations for Grouped QuantizersYanlong Chen, Yining Chen, Song Zhang, Amirhossein Habibian +1
Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The a
memory - arxiv:2609.32428 · cs.AIAuthorization Closure Graph: Minimal Repair for LLM Agents with Evolving User InstructionsQingzhuo Wang, CaiYi Wang, Jinglu Meng, Ruiyang Qin +3
Tool-using large language model (LLM) agents increasingly perform state-changing actions that require user authorization. Yet existing approaches do not provide a principled mechanism for selectively updating prior authorization when only part of an instruction changes. To this end, we propose an Au
llm agent - arxiv:2609.32426 · cs.LGHoTS: Homophily-Aware Temperature Scaling for Graph Neural Network CalibrationInwoo Tae, Yoontae Hwang, Yongjae Lee
For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrators either apply one global temperature or
benchmark - arxiv:2609.32424 · cs.AICyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain ProvenanceQi Chen, Fushuo Huo, Hangli Shen, Jingcai Guo +2
Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulne
long-contextagentllm agentmulti-agentagent systembenchmark - arxiv:2609.32423 · cs.AIPluginRSI: Recursive Improvement of Agent Harnesses with Reusable PluginsYaorui Shi, Yuchun Miao, Yuxin Chen, Jiayuan Zhang +4
The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomiz
agent - arxiv:2609.32416 · cs.RORE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy DistillationJiawei Zhang, Xiangrong Zhang, Rui Song, Huanbin Zhou +2
Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the superv
embodiedcode-as-policy - arxiv:2609.32407 · cs.AIOpening LLM Judges: Recovering Preference Signals Beyond the Final VerdictSourabrata Mukherjee, Sunayana Sitaram
LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal re
benchmarkevaluator - arxiv:2609.32405 · cs.CVToward On-Chip Training of Spiking Neural Networks for Dense Event-Based VisionMaxime Vaillant, Axel Carlier, Lai Xing Ng, Christophe Hurter +1
Event cameras provide low-latency, asynchronous visual sensing for resource-constrained robotics. Spiking neural networks (SNNs) process event streams naturally, but training deep SNNs with backpropagation through time (BPTT) requires substantial memory and remains difficult on neuromorphic hardware
memorybenchmarkevent camera - arxiv:2609.32401 · cs.AIShared Worlds, Private Minds: Structured Memory for Long-Form Writing as World CreationQiuyu Tian, Xiaowen Gu, Hang Su, Jianghan Chao +8
LLM agents that write long-form fiction need an explicit memory of the evolving storyworld to keep new events consistent with established facts. Such memory must keep heterogeneous narrative information distinct, integrate story developments across granularities, and recover dependencies that a writ
memoryllm agentbenchmark - arxiv:2609.32400 · cs.AISkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime FeedbackPengyu Zhu, Jingyi Yang, Yi Liu, Li Sun +1
Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill m
agent - arxiv:2609.32396 · cs.CLFA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy ConditionsWei Chu, Yuanzhe Dong, Ke Tan, Dong Han +10
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices o
benchmark - arxiv:2609.32395 · cs.LGMemory as a cache: Exact context reuse and deletion by constructionShengyao Wang, Jiang Liu
The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose
memory - arxiv:2609.32394 · cs.AIBeyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box OptimizationMinghao Li, Rui Tan, Ruihang Wang
Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate
manipulationagentllm agentagentic - arxiv:2609.32391 · cs.LGSCLATE: a Substrate for Continual-Learning Agent Training and EvaluationYoungmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan +2
Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training fra
memoryagentbenchmark - arxiv:2609.32390 · cs.AIReward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face IncidentMurat Ozer, Bulent Erenay, Ibrahim Berber
The July 2026 intrusion into Hugging Face production infrastructure showed how reward hacking can become an external cybersecurity incident when a capable agent encounters weak containment boundaries. This study develops a probabilistic risk model linking five stages: reward hacking, containment esc
agentai agent - arxiv:2609.32389 · cs.CVRefCompose: Multi-Reference Image Generation via LoRA-Conditioned DiffusionSai Sri Teja Kuppa, Parth Shinde, Priyadharsan Balaji S, Jinka Harshavardhan +1
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tok
memory - arxiv:2609.32385 · cs.AIDashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard AnalysisChuhan Zhang, Qi Xie, Ziyue Wang, Jianing Yin +3
Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failur
agentbenchmark - arxiv:2609.32379 · cs.LGMeasurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect AuditingWeicheng Xue
What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 2
agentllm agent - arxiv:2609.32378 · cs.AIAuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of AuthorityShaojin Chen, Huihao Jing, Wun Yu Chan, Wenbin Hu +6
LLM-based agents are increasingly deployed with authority over consequential resources and decisions in real systems. These agents often operate alongside human and LLM-based participants who hold different forms of authority. Yet workflow roles, permission settings, and review mechanisms do not nec
agentagent system - arxiv:2609.32367 · cs.LGSTRIDE: State-Transition Representation via Increment Dynamics and EvolutionYuchen Xiong, Siming Huang, Jianfeng Sun
We introduce STRIDE (State-Transition Representation via Increment Dynamics and Evolution), which defines states through derivative fingerprints and learns local functions for state transitions (qpairs), recasting continuous forecasting as transition prediction. Trailing convolution windows estimate
benchmark - arxiv:2609.32365 · cs.LGGraph Memory: Spectral Associative Memory via Dirichlet EnergyZhaoyang Shi
Dense associative memories have traditionally focused on storing and retrieving vector-valued patterns. Many modern machine learning problems, however, are naturally graph-structured, requiring memory mechanisms for relational patterns, graph diffusion geometries, community structures, and graph-bas
memory - arxiv:2609.32363 · cs.LGDiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series ForecastingWeiwei Ye, Dongyuan Li, Hangchen Liu, Haotong Jiang +2
Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model (DDPM)-based approaches have shown promise by equipping the dif- fusion process with pretrained mean and variance estimators to accommodate
benchmark - arxiv:2609.32362 · cs.CVStegGNN: Learning Graphical Representation for Image SteganographyAbhinav Kumar, Shorya Singhal, Agam Pandey, Tushar Kumar +1
Image steganography refers to embedding secret messages within cover images while maintaining imperceptibility. Recent advances in deep learning - primarily driven by Convolutional Neural Networks (CNNs) and architectures such as inverse neural networks, autoencoders, and generative adversarial netw
benchmark - arxiv:2609.32361 · cs.LGBlack-Box Auditing of Epistemic Reliability in Multi-Agent Debate DistillationDerui Wang, Zewei Shi, Rayne Holland, Ruoxi Sun +3
Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradation, in which adaptation preserves monitor
multi-agentbenchmark - arxiv:2609.32354 · cs.ROProactive Motion Planning for Human-Robot CooperationElena Basei, Edoardo Lamon, Matteo Saveriano, Daniele Fontanelli +1
This abstract addresses the incorporation of human motion prediction into proactive and dynamic human-aware motion planning, with the goal of enabling safe collaboration between humans and robots. A deep learning, graph-based model is used to forecast human motion and is integrated into a planning f
manipulator - arxiv:2609.32353 · cs.LGFewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token ReductionJunxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu +2
Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step fur
benchmark - arxiv:2609.32352 · cs.CVEyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial GroundingGujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang +6
Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to sy
benchmark - arxiv:2609.32344 · cs.AIALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMsShanfeng Huang, Zhou Fang, Song Xiao, Hai Du
For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confiden
external memorybenchmark - arxiv:2609.32343 · cs.CVOpenMASC: An Open-Source Pipeline for Cross-Trajectory Metal-Aware Sampling and Correction in Accelerated MRIZhengyi Lu, Ming Lu, Chongyu Qu, Junchao Zhu +10
Metal implants corrupt MRI measurements throughout $k$-space, yet existing accelerated MRI methods assume clean data and most metal artifact reduction approaches assume fully sampled acquisitions. No public dataset provides paired $k$-space and images with and without metal for the same anatomy, and
agent - arxiv:2609.32341 · cs.LGA Comparative Analysis of Attention versus State-Space Models for In-Context LearningEnes Arda, Semih Cayci, Atilla Eryilmaz
Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical
memory - arxiv:2609.32339 · cs.AIEnabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent ContextFeng Liang, Yupeng Li, Runhao Zeng, Francis C. M. Lau +1
Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent decides to retrieve i
memoryagentbenchmark - arxiv:2609.32336 · cs.ROWSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM GuidanceHanlin Zhang, Yuquan Wang, Tianwei Zhang, Zhenglong Sun
Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsiste
world model - arxiv:2609.32333 · cs.CVProgressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMsShanfeng Huang, Zhou Fang, Song Xiao, Hai Du
Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's
benchmark - arxiv:2609.32327 · cs.AIHyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop ReasoningZicheng Zhao, Linhao Luo, Junnan Dong, Haoran Luo +3
Large language models (LLMs) have shown strong capabilities, with retrieval-augmented generation (RAG) supporting complex multi-hop reasoning by retrieving evidence distributed across documents. Graph-based approaches exploit connections among evidence, and hypergraph-based retrieval further preserv
retrieval-augmentedbenchmark - arxiv:2609.32325 · cs.LGActive Feature Acquisition With Incomplete Training DataReza Rezvan, Valter Schütz, Han Wu, Linus Aronsson +1
In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by
benchmark - arxiv:2609.32322 · cs.LGNot All Errors Matter: Decision-Relevant Prediction Error Predicts Planning QualityLinhao Wang, Yiyan Fan, Dongjin Huang
World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different
world modelevaluation protocol - arxiv:2609.32318 · cs.LGWhat Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review AgentsZhaowei Han, Xiang Zhang, Lingxiao Guan, Danqi Hu +3
Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We
ai agentbenchmarkleaderboard - arxiv:2609.32316 · cs.CVOne Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent PredictionAng Zou, Runzhe Zheng, Zhigang li, Zhen Yang +4
Traffic lights are a key regulatory signal for autonomous driving at urban intersections, yet existing traffic signal perception is still predominantly formulated as instance-level detection or color recognition. Such formulations identify where traffic lights are and what colors they display, but l
benchmark - arxiv:2609.32313 · cs.ROMemTransfer: Benchmarking Memory Beyond Matched Experience in Embodied Decision-MakingHaiming Tang, Xianjie Dai, Gujie Shao, Zuyi Guo +4
Memory lets an embodied agent reuse past experience, yet retaining useful information does not ensure that the agent can apply it when conditions change. We present MemTransfer, a benchmark comparing six memory representations, a working-memory baseline and five representations of past experience, u
embodiedmemoryepisodic memoryagentembodied agentbenchmark - arxiv:2609.32312 · cs.AIDelayed Supervision for Test-Time Language ModelsJinha Kim, Taksh Kothari
Test-time language models adapt a compact memory while processing the input sequence. This perspective encompasses nonlinear fast-weight learning in LaCT, associative delta-rule updates in DeltaNet, and generalized delta-rule state updates in RWKV-7. Training these models to predict the next token d
memorypost-training - arxiv:2609.32303 · cs.LGTrain4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model MergingJingyuan Huang, Zuming Huang, Yucheng Shi, Zhongzhi Li +3
Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinf
agentic - arxiv:2609.32297 · cs.AIAgentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story WorldsYu Pan
A agentic story world is a dynamic system simulating who learned what, when, and from whom -- yet the standard design gives each character a private memory stream. A shared event is therefore stored once per witness, large duplication will be incurred in terms of storage. We present Agentsensus, a s
memorymulti-agentagentic - arxiv:2609.32296 · cs.LGFUND: Density Flow for Sampling Unnormalised DistributionsVikas Kanaujia, Vipul Arora
Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances, like Boltzmann Generators, offer a pro
benchmark - arxiv:2609.32295 · cs.AIGLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM AgentsWei Zhu, Yiming Wang, Rui Wang, Lixing Yu +2
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalib
agentllm agent - arxiv:2609.32292 · cs.ROAffordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language NavigationXuekang Yang, Lu Chen, Shuang Luo, Jialing Zhu +3
Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the
embodiedagent - arxiv:2609.32289 · cs.LGNeural ODEs Meet Concurrent Learning: Stable Online Learning with Lyapunov GuaranteesOmkar Sudhir Patil
Neural ODEs learn dynamics from trajectory losses, but their adjoint gradients lack the regressor-times-parameter-error structure on which Lyapunov analyses of online adaptation rest, so training on streaming data comes without stability guarantees. We show that this structure is in fact present: th
online learning - arxiv:2609.32288 · cs.LGA Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling ExponentsIndranil Halder, Rastri Dey, Cengiz Pehlevan
Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our contro
post-training - arxiv:2609.32280 · cs.CVHeroFrame-Bench: Reference-Anchored Evaluation via Rubric--Ranking Co-Evolution for Movie Hero Frame SelectionWeitai Kang, Hanieh Deilamsalehy, Yumo Xu, Dewang Sultania +2
Hero frames are in-film stills used as source imagery for theatrical posters, streaming cover art, film database listings, and other promotional placements. As the first visual entry point, they shape audiences' initial impressions of the movie and their subsequent willingness to watch it. Selecting
benchmark - arxiv:2609.32276 · cs.LGHyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation ModelingPeiyu Zhang, Heng Ping, Nikos Kanakaris, Yucheng Zhao +4
Multi-label classification (MLC) requires predicting multiple relevant labels for each instance, where a central challenge is modeling complex label dependencies arising from co-occurrence patterns. Existing approaches are limited in capturing high-order label correlations, relying on implicit learn
benchmark - arxiv:2609.32275 · cs.LGTopology-Adaptive Hyperbolic Graph Attention Networks Guided by the Hyperbolic Sombor IndexHaifang Cao, Boan Tao, Xiyuan Gao, Timing Li +2
Hyperbolic geometry has emerged as a principled space for representing hierarchical graphs. However, existing hyperbolic graph neural networks typically rely on shared curvature configurations and feature-driven attention, failing to explicitly exploit local hierarchical topological patterns. To bri
benchmark - arxiv:2609.32274 · cs.AIWhen Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill UseAnjie Xu, Zhiyu Zhang, Ruiqing Ding, Fengli Xu +1
Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts? We introduce SkillDelta, a framework for predicting task-condit
agentbenchmark - arxiv:2609.32269 · cs.LGBefore Answering: Evidence Sufficiency under Size-Matched Memory ConstructionJoyanta Jyoti Mondal, Md. Shifatul Ahsan Apurba, Mridul Banik, Md Masud Al Mahmud +1
Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memor
memorybenchmark
02 US SEMI · SEC 8-K FILINGS
2 itemsscanned: NVDA / AVGO / MRVL / COHR / LITE / AMD / TSM / SMCI / ANET / CRDO / POWL / VECO
03 HUMANOID · COMPANY NEWS
54 itemsscanned: figure-ai / 1x / boston-dynamics / unitree / apptronik / sanctuary-ai / neura-robotics / agility-robotics / physical-intelligence / agibot
Figure AI (10)
- Figure AISeptember 17, 2026Helix 2.5: Zero-Shot 30-Home Generalization
- Figure AIAugust 25, 2026Introducing Index: Building The World’s Largest and Most Diverse Physical Dataset
- Figure AIOctober 09, 2025Introducing Figure 03
- Figure AISeptember 03, 2026Figure and Nscale Sign Strategic Partnership For Up to 100,000 GPUs on the NVIDIA Vera Rubin Platform
- Figure AIJuly 08, 2026Notice Regarding Unauthorized Attempts to Sell Figure Stock
Boston Dynamics (10)
Unitree 宇树 (10)
- Unitree 宇树Components
- Unitree 宇树Kung Fu Meets Spring, Unitree SFG Robots Present "Cyber Real Kung Fu" in the Year of the Horse2026-05-31Media Coverage
- Unitree 宇树Welcoming Myanmar President Min Aung Hlaing to Unitree2026-08-05Media Coverage
- Unitree 宇树Unitree founder Wang Xingxing graces the cover of Time magazine2026-08-05Media Coverage
- Unitree 宇树Unitree Announces H2 Plus, an NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research2026-06-01Media Coverage
Apptronik (1)
Sanctuary AI (6)
- Sanctuary AIPress ReleaseProduct UpdatesSanctuary AI Expands Physical AI Strategy to Industrial Robotics, Demonstrating Production-Ready AI PerformanceRead More
- Sanctuary AICorporate NewsDaniel Friedmann Appointed CEO of Sanctuary AIRead More
- Sanctuary AIPress ReleaseZeon Invests in Sanctuary AI and Partners to Advance Specialized Materials for Dexterous RoboticsRead More
- Sanctuary AIThought LeadershipWeb Summit Reflections: Canada’s Physical AI Moment Can’t WaitRead More
- Sanctuary AIProduct EvolutionSanctuary AI Demonstrates Zero-Shot In-Hand Manipulation on Hydraulic HandRead More
Agility Robotics (10)
- Agility RoboticsAgility’s Humanoid Deployment ProcessPras VelagapudiAugust 04, 2026
- Agility RoboticsThe Realistic Pathway to HomeInsightMay 26, 2026
- Agility RoboticsAgility and AIInsightMarch 16, 2026
- Agility RoboticsAgility Gets a New BrandInsightMarch 5, 2026
- Agility Robotics2026: The Automation EvolutionInsightJanuary 16, 2026
智元 AgiBot (7)
- 智元 AgiBotAGIBOT and ASD Deploy 100 AGIBOT X2 Humanoid Robots Across 100 Retail Stores2026-09-29
- 智元 AgiBotAGIBOT and Chimelong Launch Large-Scale Embodied AI Theme Park with More Than 300 Robots2026-09-23
- 智元 AgiBotAGIBOT Demonstrates WITA-Omni in Live Human-Robot Interaction at The Greater Bay Area Film Concert2026-09-21
- 智元 AgiBotAGIBOT Launches AGILE 2.0, Advancing Perception-Driven Locomotion for Humanoid Robots2026-09-14
- 智元 AgiBotAGIBOT Releases GE-Act 2.0, Providing the First Systematic Validation of a Scaling Path for Native World-Action Models2026-09-11