PHYSICAL AI · 2026-09-30

Physical AI Brief

Daily cross-source signals for the Physical AI supply chain — silicon photonics, CPO, VLA models, humanoid hardware, embodied AI. Three streams, one page, zero filler.

785 items today · 729 arxiv · 2 SEC 8-K · 54 humanoid · 0 CN photonics

01 ARXIV · PHYSICAL AI PAPERS

729 items
  1. arxiv:2609.38178 · cs.RO
    Skill-Space Shooting for Autonomous Robot Policy Improvement
    Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao

    Robots deployed in the physical world must be able to improve beyond their initial training as they encounter new situations and failures. For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction. Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks. Yet completing tasks this way does not itself teach a task policy to overcome its own failures; that requires turning these behaviors into learnable correc

    robot policyagentic
  2. arxiv:2609.38177 · cs.CV
    Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
    Jaewoo Jung, Hyeonseo Yu, Honggyu An, Jisang Han +9

    Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which

    benchmark
  3. arxiv:2609.38173 · cs.RO
    In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks
    Minxing Li, Minghao Han, Weizhi Zhao, Hanwen Wang +9

    We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Bui

    manipulation
  4. arxiv:2609.38172 · cs.RO
    Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation
    Zihan Wang, Zhen Wu, Pieter Abbeel, Rocky Duan +5

    Teaching humanoids loco-manipulation skills, such as carrying diverse objects, via visual imitation is a promising path toward generalist robots. However, collecting diverse, high-quality interaction videos, such as clips that clearly show a person's full body and unoccluded interactions with objects, poses a practical barrier to scaling this approach. We propose PRISM, a real-to-sim-to-real framework that overcomes this limitation by amplifying a handful of real videos into a large, diverse training set. PRISM first generates hundreds of diverse "counterfactual" human-object interaction video

    manipulationhumanoidsim-to-real
  5. arxiv:2609.38170 · cs.CV
    Adversarial Training for Pixel Diffusion
    Xin Lin, Zhifei Zhang, Yuqian Zhou, Haitian Zheng +3

    Pixel diffusion models generate RGB images directly, avoiding the bottleneck of an autoencoder, yet their outputs still systematically underrepresent fine-scale natural-image statistics. We show that adversarial learning provides an effective post-training correction for this deficiency. Starting from a pretrained model, we retain its original diffusion or flow-matching objective and add an adversarial loss to the predicted output at non-high-noise timesteps, leaving the model architecture and sampling procedure unchanged. To our knowledge, this is the first systematic study of adversarial pos

    post-training
  6. arxiv:2609.38169 · cs.LG
    STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
    Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu +5

    Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitude

    memorypersistent statepost-trainingbenchmark
  7. arxiv:2609.38166 · cs.LG
    LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization
    Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li +9

    Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challen

    memorylong-context
  8. arxiv:2609.38164 · cs.RO
    Rho: A Foundation for Efficiently Adaptable VLA Models
    Rho Team, Simran Bagaria, Daphne Chen, Dean Fortier +10

    General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtrainin

    vlavla modelmanipulation
  9. arxiv:2609.38163 · cs.RO
    Rethinking Representations for World-Action Modeling
    Haoyi Jiang, Liu Liu, Xinjiang Wang, Zhihao Sun +11

    World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure effective policy learning. These findings motivate ReWAM, a representation-centric world-action model built on pre-trained DINO features. Feature Calibration and a Temporal Representation Bottleneck organize these features into compact world states suited to dynamics modeling. Act

    embodiedrobotwinworld model
  10. arxiv:2609.38155 · cs.LG
    Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
    Hui Ren, Lei Fan, Henry Pao, Han Guo +4

    Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of th

    memorybenchmark
  11. arxiv:2609.38154 · cs.CV
    LongLive-Plug: Once-for-All Distillation for Video Generation
    Shuai Yang, Luozhou Wang, Wei Huang, ZhiFei Chen +8

    Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context e

    world modellong-context
  12. arxiv:2609.38152 · cs.CV
    FracGen: Learning How Objects Stretch and Tear with Physics-Informed Video Generation
    Trong-Tung Nguyen, Jiahan Zhang, Anand Bhattad

    We introduce FracGen, a fracture-aware video generation model that produces plausible, controllable fracture dynamics from a single image of an intact object, conditioned on physics signals. To train FracGen, we build FracSim, a fracture-aware simulation framework that augments material point method (MPM) simulation with a continuum damage model, producing paired fracture videos and dense, pixel-aligned physical fields at no additional cost beyond standard rendering. FracGen leverages these maps in two ways: it is trained to jointly predict them alongside RGB video, encouraging the model to ca

    benchmark
  13. arxiv:2609.38147 · cs.AI
    Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning
    Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen +8

    As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work wit

    persistent memoryagentagenticbenchmark
  14. arxiv:2609.38143 · cs.LG
    Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
    Cheng Qian, Kunlun Zhu, Beibin Li, Zhenhailong Wang +1

    Agent performance depends on both reasoning ability and the environment in which it acts. We study test-time AI-for-AI, asking how a Builder can learn to construct better execution environments for a Target while both models' weights remain fixed. To make the Builder's experience reusable, we introduce Meta-Skill: principles specifying when support is needed and what resources to provide. The Builder learns these principles from Target's execution feedback on the development set, then uses the frozen skill bank to construct harnesses for unseen tasks. Across Harness-Bench and NewtonBench, full

    agentai agentself-improvement
  15. arxiv:2609.38140 · cs.CV
    Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
    Yu Xu, Yuxin Zhang, Xiao Yang, Haotian Yang +6

    Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing frag

    world modelbenchmark
  16. arxiv:2609.38137 · cs.CL
    LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning
    Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen, Xi Ye

    Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this paper, we introduce a benchmark for evaluating both the effectiveness and efficiency of long-context harnesses. Our tasks require diverse retrieval strategies, including lexical search and semantic matching, together with strategic and adaptive reasoning over global and local context. Much of the c

    long-contextlong contextbenchmark
  17. arxiv:2609.38136 · cs.CV
    CLeaR: A Unified Framework for Resolving the Leakage-Degradation Dilemma in Style Transfer
    Teng Zhou, Yunhao Chen

    Style transfer aims to render target content in the style of a reference image, but existing methods often suffer from content leakage, where objects, layouts, or semantics from the style reference appear in the generated output. Although prior data-driven and training-free methods can reduce leakage, they often face a leakage-degradation dilemma: stronger content suppression may weaken style fidelity, while richer style preservation may reintroduce unwanted reference content. We identify this dilemma across the full style-transfer pipeline, including feature separation, feature-space groundin

    llm-as-judge
  18. arxiv:2609.38133 · cs.RO
    Multi-Agent Flow Matching with Decoupled Generative Guidance
    Ruoyu Lin, Magnus Egerstedt, Fabio Pasqualetti

    Generative modeling is widely used for producing diverse objects from complex, multimodal distributions. However, its expressivity does not, in general, come with formal guarantees that the generated objects satisfy hard constraints or requirements. In multi-agent generation, this problem becomes more challenging because a hard requirement can depend on multiple agents, while each agent may need to determine its own guidance input without relying on the simultaneously computed guidance inputs of other agents. To this end, we introduce DeGG-Flow, a general framework for multi-agent flow matchin

    agentmulti-agent
  19. arxiv:2609.38123 · cs.CV
    HelixWorld: A Real-time Interactive Audio-Visual World Model
    Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang +12

    World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher

    helixworld model
  20. arxiv:2609.38121 · cs.LG
    WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
    Jiale Chen, Vage Egiazarian, Eldar Kurtić, Torsten Hoefler +1

    KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck, we introduce WUSH-KV for low-bit KV-cache quantization. It adapts WUSH, which constructs a data-aware transform from the second-order statistics of both factors in a matrix product to reduce quantization error. WUSH-KV uses calibration data to construct separate key and value transforms, with the value transform folded into the model weights and the key transform applied after RoPE. The transforms can be paired with clipped quantizers. For one

    memorylong-context
  21. arxiv:2609.38120 · cs.AI
    Stochastic World Models for Verifying Vision-Based Neural Feedback Systems
    I. Samuel Akinwande, Mykel J. Kochenderfer, Clark Barrett

    Verifying a vision-based neural feedback system requires a model of the observations its controller acts upon. Such a model must capture the variation the sensor produces, while remaining tractable for closed-loop analysis. Generative adversarial networks (GANs) have served as perception surrogates, but they are large, reproduce complex scenes poorly, and are hard to verify. We explore stochastic world models as a richer class of perception surrogates. We train a world model with physically grounded latents, built from operations that standard verifiers bound. It reproduces held-out frames mor

    world modelbenchmark
  22. arxiv:2609.38119 · cs.CV
    VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents
    Jinfa Huang, Jianming Xu, Jingyang Lin, Zhengyuan Yang +1

    Long-form video understanding requires multimodal agents to iteratively gather evidence over many reasoning steps. However, most existing agentic methods suffer from semantic thrashing: as append-only working memory grows, attention to key evidence collapses, and the agent loses access to what it has already found. First, we provide a structural argument showing that append-only memory can incorporate newly observed target evidence, but cannot remove accumulated noise or prevent ordered context growth without a rewrite operator. Second, motivated by this analysis, we propose VideoLoop, a multi

    memoryagentagentic
  23. arxiv:2609.38113 · cs.MA
    IMPACT: Modeling Socially Interdependent Movement in a Generative Multi-Agent Simulation of a Pompeian Household
    Tianqi Liu, Nayoung Kim, Julia Sebastien, Kathryn Gleason +2

    Simulations of archaeological sites can make interpretations of past cultural practices observable and examinable. Generative multi-agent simulations offer a bottom-up approach to modeling how people collectively moved through and used historical spaces. However, current agents designed to simulate everyday life often plan and act independently, limiting their ability to capture how movement depends on others' actions. We introduce IMPACT (Interdependent Movement Planning through Inter-Agent Constraints and Triggers), an architecture that uses culturally specific roles and obligations to defin

    multi-agent
  24. arxiv:2609.38108 · cs.LG
    Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
    Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos López de Prado +1

    Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Sea

    agentllm agentbenchmark
  25. arxiv:2609.38107 · cs.AI
    Correct Answers, Invalid Traces: What Verifiable Grade-School Math Reveals About Chain-of-Thought Traces
    Ratish Puduppully, Pranabendu Misra, Paarth Iyer, Durgesh Kalwar +2

    Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthetic grade-school mathematics benchmark designed to study thinking traces and used to support claims of learned reasoning and planning. Crucially, iGSM exposes the exact quantities and dependencies that a correct solution should use, allowing generated traces to be checked programmatically step by st

    agentbenchmark
  26. arxiv:2609.38105 · cs.RO
    FORM: Robot Manipulation through Direct Material Law Identification
    Stepan Tretiakov, Ruihan Zhao, Cheng-Hsi Hsiao, Xingjian Li +5

    When interacting with an unfamiliar deformable material, a robot lacks prior knowledge of its physical properties and how it will respond to applied forces and motion. Rapid online identification is therefore essential for reliable manipulation. We present FORM (From Observed Response to Material laws), which identifies material properties from a single robot interaction and reuses the recovered model to plan manipulation under new actions and geometries. We use weak-form momentum balance to convert observed material motion and contact forces into linear equations in the unknown material param

    manipulation
  27. arxiv:2609.38104 · cs.LG
    Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
    Panagiotis Theodoropoulos, Nan Jiang, Xintong Duan, Ali Hasan +3

    Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To res

    memorypost-training
  28. arxiv:2609.38096 · cs.LG
    Tail-Influence Sampling for CVaR Policy Evaluation
    Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov +1

    Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-In

    policy evaluation
  29. arxiv:2609.38095 · cs.LG
    Probe-Space Preconditioning for Fast and Stable Zero-Order Training
    Francois Chaubard, Mykel J. Kochenderfer, Chris Ré

    Backpropagation (BP) dominates deep learning but imposes a massive memory tax. For example, training OPT-30B with Adam requires $\approx$ 600GB of GPU memory (assuming batch size 8 and sequence length 2048). Alternatively, zero-order optimization (ZOO) trains in inference-mode (requiring only $\approx$ 60GB for the same model): no stored activations, no gradients, and no optimizer states. However, ZOO convergence has lagged behind BP. In this work, we evaluate two methods to close this gap. First, we show that reallocating training compute budget from many steps to large effective batch sizes

    memorypost-trainingbenchmark
  30. arxiv:2609.38093 · cs.AI
    Character Training for Risk-Averse Agents
    Arav Dhoot, Punya Syon Pandey, Jamie Johnson, Daniel Tan +2

    Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, finding that persona traits provide a robust mechanism for instilling risk preferences. To do this, we construct a model constitution describing constant absolute risk aversion (CARA) over an agent's resources and instill it through on-policy distillation. Despite never seeing the benchmark's decision

    ai agentbenchmark
  31. arxiv:2609.38087 · cs.RO
    CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments
    Tan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger, Cuc T. Trinh +5

    Behavior Foundation Models (BFMs) give humanoids a promptable policy over a latent behavior space, enabling one single vector to represent a motion to imitate, a pose to reach, or a reward to maximize. Forward-Backward representations successfully produce such spaces, but at the cost of hundreds of GPU-hours for a single robot. Moreover, when the training process is repeated for a second robot, it produces a second space unrelated to the first, resulting in embodiment-specific latents that do not unify or transfer. We address these problems with CrossBFM, treating the latent space as the trans

    humanoidwhole-body control
  32. arxiv:2609.38086 · cs.CV
    VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents
    Zheng Jiang, Houde Qian, Yiming Chen, Ling Li +4

    Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introduce VISTA, which internalizes collective visual experience through on-policy distillation by turning observations from same-input rollouts into shared supervision. Collective visual experience distillation (CVED) organizes these observations with their interaction context and alig

    agent
  33. arxiv:2609.38081 · cs.LG
    Traversing the solution space of neural networks with Hessian Null Space Continuation
    Ann Huang, Mitchell Ostrow, Zhouyang Lu, William T. Redman +2

    On a single task, deep networks can learn many solutions, depending on their optimizer, training data, architecture, and hyperparameters. Many of these solutions are mode-connected: rather than isolated points in weight space, they are connected by low-loss regions. Yet how their internal computation varies within these regions is unknown. A parallel line of work has identified the degeneracy of neural representations: many networks reach similar training loss with distinct internal structures. However, it is unclear how these solutions are related in weight space. We unify these subfields and

    memory
  34. arxiv:2609.38078 · cs.RO
    MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
    Bingxuan Li, Siqi Song, Yizhuo Wu, Jiarui Yao +2

    Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot mo

    vision-language-actionembodiedmanipulationliberomemoryagentic
  35. arxiv:2609.38070 · cs.AI
    Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
    Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng +6

    As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introdu

    benchmark
  36. arxiv:2609.38066 · cs.LG
    Alpha Diffusion Language Models: Factorization Alone Is Not the Problem
    Nikita Gushchin, Dmitry Baranchuk, Alexander Korotin

    Discrete diffusion language models can generate multiple tokens in parallel, but reducing the number of denoising steps can lead to inconsistent predictions. Standard cross-entropy training fits conditional token marginals, whereas parallel generation requires consistent joint predictions. We introduce Alpha Diffusion Language Models (AlphaDLM), trained with a sequence-level alpha loss that recovers cross-entropy in the limit of vanishing alpha and has a joint-mode optimum at alpha one. Our analysis characterizes how the objective and factorization jointly determine the fitted distribution. We

    benchmark
  37. arxiv:2609.38065 · cs.LG
    Jaxolotl: A Unified High-Performance Benchmark Suite for LTL-Based Multi-Task RL
    Mathias Jackermeier, Jacques Cloete, Alessandro Abate

    Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to agents, and has been successfully adopted for training generalist multi-task policies. However, differences in implementations, task distributions, and evaluation protocols make existing methods difficult to compare, while high computational costs limit the scale and statistical reliability of experiments. We introduce Jaxolotl, a unified high-performance benchmark suite for multi-

    benchmarkevaluation protocol
  38. arxiv:2609.38059 · cs.RO
    WorldLine: Action-Driven Visual Simulation for Robotic Manipulation
    Shenghe Zheng, Wenbo Li, Jiyao Zhang, Bin Xia +3

    Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning

    embodiedmanipulationrobotwinaction-conditionedpolicy evaluation
  39. arxiv:2609.38057 · cs.RO
    EVO-WAM: Evolving World Action Models through Video-Action Verification
    Shiyang Zhou, Xionghao Wu, Wenbo Li, Shenghe Zheng +10

    Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories

    robotwin
  40. arxiv:2609.38051 · eess.SY
    Non-Holonomic Gradient Play: Leafwise Nash Equilibria, Stability, and Deception
    Mahmoud Abdelgalil, Miroslav Krstic, Jorge I. Poveda

    We study generalized learning dynamics in multi-agent systems whose joint state evolves on a manifold and whose agents act through state-dependent, potentially nonholonomic vector fields. Under a bundle-splitting condition, we show that these dynamics admit an intrinsic representation as projected Riemannian gradients, giving rise to a class of \emph{nonholonomic gradient play} dynamics. We characterize the local stability of its equilibria through an intrinsic linearization that explicitly captures the effects of the Riemannian connection and the nonholonomy of the actuation frame. The framew

    manipulationagentmulti-agentagent system
  41. arxiv:2609.38049 · cs.LG
    Improving Function Space Flow Matching with Kernel Optimal Transport
    Fred Xu, Thomas Markovich, Barbora Barancikova, Yizhou Sun

    Generative models for function-valued data, such as time series and solutions of partial differential equations, must learn distributions over infinite-dimensional spaces. Functional Flow Matching (FFM) extends Flow Matching to this setting, learning a velocity field whose flow transports a Gaussian prior to the data distribution, but it inherits the independent endpoint pairing of standard Flow Matching: in each batch, prior and data samples are matched arbitrarily, so the conditional bridge must traverse both the shared global structure of the dataset and instance-specific residuals. In func

    benchmark
  42. arxiv:2609.38046 · cs.RO
    EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation
    Yiming Jiang, Chen Jin, Chongyang Xu, Yilun Chen +2

    Egocentric human demonstrations offer an accessible source of task experience, but differences in body scale and controller response, together with missing robot states, limit their value as humanoid training supervision. We present EgoAlign, a data-construction framework that converts these demonstrations into action and state supervision compatible with a general-purpose, continuous whole-body controller, without collecting physical-robot demonstrations. Using the target-robot model and simulator, EgoAlign guides demonstration collection through execution feedback. It preserves locomotion re

    manipulationhumanoidteleoperationwhole-body control
  43. arxiv:2609.38043 · cs.AI
    UserProxyBench: Evaluating LLM User Simulators for Agent Benchmarks and Training
    Ashish Jain, Armaan Sandhu

    Interactive agent benchmarks and multi-turn reinforcement learning increasingly place a second language model in the role of the user. This simulated user controls what information the agent receives and when, yet current benchmarks score only the agent and do not directly measure whether the user correctly executed its assigned role. We introduce UserProxyBench, an evaluation layer over the tau-bench family, and the User Fidelity Score (UFS), which measures adherence to the benchmark's private user instructions using task-grounded rubric criteria scored independently of agent success. Holding

    agentagent benchmarkbenchmark
  44. arxiv:2609.38024 · cs.AI
    Retrieval-Augmented Skill Optimization via Cross-Harness Adaptation
    Jaewon Chu, Ji Soo Lee, Jihwan Park, Dohwan Ko +7

    An agent skill is a reusable, actionable natural-language artifact that guides an agent to perform a task effectively under a given harness. Recent studies have explored the optimization of agent skills, contributing to a growing collection of publicly available skills spanning diverse tasks, domains, and harnesses. Despite millions of publicly shared skills, existing skill optimization methods largely overlook this accumulated knowledge, instead relying solely on expensive agent rollouts to iteratively refine skills for a target task. To address this, we propose \textbf{Retrieval-Augmented Sk

    retrieval-augmentedagentagent benchmarkbenchmark
  45. arxiv:2609.38023 · cs.AI
    PE-EK-PINN: Physics Embedding with Evolving Kernel for Scalable Physics-Informed Neural Networks
    Huiwen Zhang, Feng Ye, Chu Ma

    Physics-Informed Neural Networks (PINNs) embed governing equations into deep learning, but enforce them only through loss residuals, leaving highly oscillatory wave behavior to be discovered by optimization. As a result, methods that achieve relative $L_2$ errors below $10^{-3}$ on standard manufactured Helmholtz benchmarks can fail on practical radiation problems involving singular excitations, absorbing boundaries, and wave fields spanning tens of wavelengths. Architectural physics embedding addresses this limitation by factorizing the field into analytically derived oscillatory kernels and

    benchmark
  46. arxiv:2609.38021 · cs.AI
    Auditable Long-Term Memory: A Deterministic Retrieval Chain Measured at 479/475 of 500 on LongMemEval-S
    Christopher J. Chanhnourack

    We evaluate an auditable long-term memory system on LongMemEval-S. Its retrieval chain uses hybrid candidate retrieval, cross-encoder reranking, coverage-first packet compilation, and deterministic reasoning scaffolds; an LLM is used only as a replaceable final reader. The chain places all gold sessions in the candidate pool for 468/470 answerable questions and produces gold-complete packets for 462/470. With a Claude Opus reader called through an unpinned CLI alias, two 500-question passes score 479/500 and 475/500 under GPT-4o. The 72 answerable knowledge-update rows used a substantively mod

    memoryagentic
  47. arxiv:2609.38020 · eess.SY
    Grid Demand Flexibility Assessment of AI Data Centers via Batch Workload Temporal Shifting
    Suntao Su, Liang Du, Shengyi Wang

    The rapid growth of artificial intelligence (AI) data centers has introduced new challenges to power system operation. As their power demand becomes larger and more variable, quantitatively characterizing their demand flexibility is increasingly important for effective power system coordination. However, heterogeneous workload characteristics and resource requirements make this flexibility difficult to characterize directly. This paper proposes a framework for assessing the grid-compatible demand flexibility of AI data centers via batch workload temporal shifting. An averaging-based resource u

    memory
  48. arxiv:2609.38018 · cs.LG
    Prompts Live on an Arc: Gaussian Curricula in Fisher--Rao Coordinates for Rollout-Efficient GRPO
    Mei Okonkwo, Pixel Nomand, Julian Berg, Elena Voss +4

    Group relative policy optimization (GRPO) learns only from prompts whose sampled responses disagree: a group that is entirely correct or entirely incorrect has zero reward variance, contributes no gradient, and still consumes its rollouts. Prompt-selection methods reduce this waste by steering sampling toward intermediate pass rates, but they choose the target, its width, and the uncertainty model heuristically, in raw pass-rate or logit coordinates. We show that GRPO comes with a natural coordinate for pass rates: the arc length $ψ=\arcsin\sqrt{p}$ on the Bernoulli Fisher--Rao manifold. In ar

    benchmark
  49. arxiv:2609.38019 · cs.CV
    Beyond Lip Sync: Reference-Grounded Oral Refinement for Audio-Driven Portrait Animation
    Bangxun Tang

    We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rather than a generic one. Existing lip-sync systems follow the audio closely and keep the face recognizable, yet the mouth they render is an average mouth: the shape and texture of the lips, the arrangement of the teeth, and how much of them shows as the mouth opens are not that person's. The problem persists because nothing in current training or evaluation asks for the person's own mouth: perceptual losses accept any plausible mouth, face identi

    evaluation protocol
  50. arxiv:2609.38005 · cs.AI
    Diagnosing and Improving Probabilistic Reasoning in Large Language Models
    Huaman Sun, Dingcheng Wang, Jason Hartline, Jessica Hullman

    Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL intervention

    benchmark
  51. arxiv:2609.38004 · cs.LG
    No Scale Left Behind: Multi-Scale Autoencoder with Bi-directional Attention for Time Series Anomaly Detection
    Jiaheng Guo, Haochen Zhang, Yu-Chao Huang, Jinhao Duan +2

    Time series anomaly detection (TSAD) plays a crucial role in healthcare, finance, industrial monitoring, and other sectors. Within and between these settings, anomalies span vastly different temporal scales, from sub-second point spikes to multi-hour drift patterns. However, most existing TSAD methods commit to a single temporal granularity, and multi-scale designs either analyze different scales in isolation or are constrained to a predefined coarse-to-fine hierarchy, both failing to sufficiently capture multi-scale interactions. To resolve this limitation, we propose Multi-Scale Autoencoder

    benchmark
  52. arxiv:2609.37993 · cs.AI
    BITEM at the NTCIR-19 R2C2 Task: Predicting Confidence from Agentic RAG Pipeline Signals
    Julien Knafou, Luc Mottin, Alexandre Flament, Paul van Rijen +2

    The BITEM team entered both subtasks of the NTCIR-19 R2C2 task with a single agentic pipeline, in which a model searches, reads and records evidence over a movie corpus while an orchestrator holds the record and rules on what may be submitted. A claim is admitted only once an entailment cascade has checked it against the passage it cites, and an answer is released only once enough checked evidence stands behind it. Each question is run three or four times, every pass retrieving from a corpus stripped of what the earlier passes have already seen. The confidence filed with each answer is compute

    ragrag pipelineagentic
  53. arxiv:2609.37989 · cs.LG
    TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models
    Deqing Fu, Huangyuan Su, Rajat Sen, Taman Narayan +3

    Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, task descriptions, and auxiliary files that carry dataset semantics. Meanwhile, self-evolving machine learning engineering (MLE) agents train models from scratch on each dataset, yet jointly searching over features, architectures, and hyperparameters is noisy and prone to overfitting. We introduce TabFM-Auto, which pairs a tabular foundation model, TabFM, with a language model agent that evolves the data pipeline around it. Guided by dataset metad

    agentself-evolvingbenchmark
  54. arxiv:2609.37988 · cs.AI
    KV-Kaizen: Learning Context-Adaptive Cache Compression Choices
    Joao Monteiro, Louis Béthune, Anastasiia Filippova, Sonia Laguna +2

    As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produ

    memorylong-context
  55. arxiv:2609.37986 · cs.CV
    ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals
    Xuyi Hu, Francesco Palandra, Shangzhe Wu, Daniel Cremers +2

    Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithf

    quadrupedbenchmark
  56. arxiv:2609.37976 · cs.LG
    $S^3$: Spectral Null-Space Swap Makes Reasoning Models Efficient
    Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang +1

    LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null spac

    post-training
  57. arxiv:2609.37972 · cs.LG
    Dagger: Decoupling-based Model Stealing Attack against Graph Neural Networks
    Ying Song, Xiaowei Jia, Balaji Palanisamy

    As Graph Neural Networks (GNNs) are widely deployed as Machine Learning-as-a-Service (MLaaS) APIs, model stealing attacks have emerged as a critical security threat. By querying a victim model's black-box API, an adversary can construct a functionally equivalent surrogate model, compromising proprietary intellectual property and downstream security. Existing GNN stealing attacks, however, rely on overly permissive assumptions, such as soft-label outputs, large query budgets, full-graph query access, and prior knowledge of victim backbones that rarely hold in real-world deployments. In this wor

    benchmark
  58. arxiv:2609.37970 · cs.RO
    PhysWAM: Physically Consistent World Action Model for Autonomous Driving
    Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang +10

    World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$

    agent
  59. arxiv:2609.37969 · cs.CV
    SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video
    Haozhe Liu, Tian Ye, Shuchen Xue, Yitong Li +8

    High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinforcement learning (RL) post-training, and a final one-step distillation. We introduce Refiner-Bench,

    post-trainingbenchmark
  60. arxiv:2609.37968 · cs.LG
    SelfSearch: Reward-Free Search for Self-Improving Agents
    Jungwoo Yang, In Jin Kong, Yohan Jo

    Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experienc

    agentllm agentself-improvingself-improvementbenchmark
  61. arxiv:2609.37959 · cs.LG
    TabFM: A Zero-Shot Foundation Model for Tabular Data
    Weihao Kong, Erez Louidor Ilan, Shuxin Nie, Taman Narayan +5

    Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representations that transfer zero-shot to real-world tasks. Across all 51 benchmark datasets in TabArena (38 cla

    benchmark
  62. arxiv:2609.37953 · cs.AI
    Topological Coherence for Self-evolving Multi-agent Systems
    Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen +3

    Complex tasks inherently couple workflow structure, agent responsibility, collaboration, and memory access: task regions delimit responsibility and tool scope, cross-region dependencies give rise to handoffs, and ownership boundaries delimit private and selectively shared memory. Existing methods can jointly optimize agent and communication structures, yet such optimization does not by itself require responsibility, handoff, and memory boundaries to remain consistent with task dependencies. We term this requirement topological coherence. We introduce TOCOMAS, a Topology-Coherent Multi-Agent Sy

    memoryagentmulti-agentagent systemself-evolving
  63. arxiv:2609.37950 · cs.CV
    Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
    Bingjun Luo, Jialin Guo, Siqi Li

    Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improvement. We introduce Video-RSI, a framework for recursive self-improvement in which a video understanding agent uses its own language model to revise its harness. Through active video investigation, the model revisits the original training videos to test competing failure explanat

    agentself-improvementbenchmark
  64. arxiv:2609.37944 · cs.LG
    Identifiability Guarantees for Drivers and Dynamics of Delayed Physical Systems
    Julien Boussard, Antoine Débouchage, Théo Saulus

    A wide range of methods have been proposed, including physics-informed neural networks, which are powerful but do not guarantee identifiability of the dynamics, symbolic regression, which requires a set of precomputed operations, and causal discovery, which is more principled but usually relies on strong assumptions that physical systems may violate. In this work, we develop a theory-grounded method and prove that under a set of permissive assumptions, the structural drivers and drift of stochastic delayed differential equations are identifiable. Our method outperforms others on a benchmark fo

    benchmark
  65. arxiv:2609.37941 · cs.LG
    An Efficient Machine Learning Approach for Degradation Forecasting in AEM Water Electrolysis
    Marco Veneriano, Ani Gjergji, Sebastiano Bellani, Andrea Riva +2

    This study provides a data-driven analysis of a novel dataset of single-cell Anion Exchange Membrane water electrolyzers (AEMWE), operated under constant current load across multiple heterogeneous experimental campaigns. We train and evaluate a range of machine learning models with different complexity, including linear baselines, LSTMs and CNNs, to perform medium-term forecasting of the cell voltage degradation curve. The models are assessed within a rigorous training and evaluation framework specifically designed for heterogeneous industrial data.

    evaluation framework
  66. arxiv:2609.37938 · cs.RO
    Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark
    Yuedong Tan, Lei Qi, Yu Liu, Di Wen +12

    Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring whether local video understanding actually transfers. We introduce EgoGears, a complementary single- and multi-video benchmark designed to diagnose this transition. It contains 567 single-video and 1,487 multi-video questions derived from 126 human-collected egocentric recording

    embodiedbenchmarkleaderboard
  67. arxiv:2609.37937 · cs.CV
    Look Closer: Patch-wise Supervision for AI-Generated Image Detection
    Zhida Zhang, Tao Wu, Siyu Liu, Jie Cao

    How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the r

    leaderboard
  68. arxiv:2609.37935 · cs.LG
    Post-Anomaly Detection Inference for Deep SVDD
    Cao Le Cong Thanh, Dang Quang Vinh, Vo Nguyen Le Duy

    Deep Support Vector Data Description (Deep SVDD) has become a prominent framework for unsupervised anomaly detection by learning latent representations that compactly characterize normal data around a center. Despite its empirical success, anomaly decisions produced by Deep SVDD are typically made solely based on anomaly scores without rigorous statistical guarantees, thereby limiting their reliability in safety-critical and high-stakes applications where false positives must be strictly controlled. In this paper, we propose PADI (Post-Anomaly Detection Inference), a novel framework that equip

    benchmark
  69. arxiv:2609.37930 · cs.LG
    Learning What to Remember: Long-horizon Counterfactual Memory Optimization
    Jiaming Tang, Mingyan Liu, Armin Sarabi

    Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before the rewrite. We introduce Memory Gain Policy Optimization (MGPO), which isolates the incremental value of each memory rewrite by crediting it for its marginal contribution to current and future downstream utility. This turns delayed memory utility into a direct learning signal f

    memory
  70. arxiv:2609.37924 · cs.LG
    Time-Anchored Diffusion Language Models: Latent-Space Caching for Fast Generation
    Joel Anto Paul, Litu Rout, Aditya Akella, Sanjay Shakkottai

    Recent work on anchored diffusion language models improves denoising by shaping an intermediate latent space with supervised important-token targets. In this work, we introduce time-based (self-supervised) anchoring, which learns and reuses latent anchors without requiring such targets. Our key observation is that anchors encode persistent properties of the clean sequence, such as its semantic intent, global structure, or intermediate plan. Although their hidden representations become stale as the token canvas evolves, their semantic content remains useful across nearby diffusion times. This i

    benchmark
  71. arxiv:2609.37923 · cs.CV
    EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory
    Ziyun Zeng, Hang Hua, Shaden Alshammari, Rogerio Feris +2

    Agents can learn from past executions, but enabling different agents to reuse and build on one another's experience remains challenging. We introduce EpiCon, a shared multimodal memory framework for agent collective learning without updating host model parameters. EpiCon links question-level memory evolution to a persistent experience bank through two independently trained 2B models: a memory controller and a tree self-organizer. The controller jointly refines textual guidance and visual evidence across attempts and selectively includes visual memory. The self-organizer consolidates lessons hi

    memoryagentbenchmark
  72. arxiv:2609.37922 · cs.RO
    WayFinder: Hierarchical Visual-Language-Action for Zero-Shot Waypoint Generation and Low-Level Kinematic Control
    Timothy K Johnsen, Marco Levorato

    Visual Language Action (VLA) models offer unprecedented generalization for autonomous robots; however, their real-world deployment is frequently bottlenecked by unreliable execution and the prohibitive computational cost of fine-tuning for specific robot embodiments and tasks. To bridge this gap, we propose WayFinder, an end-to-end, closed-loop hierarchical VLA framework that circumvents the need for fine-tuning by decoupling high-level task reasoning from low-level kinematic control. WayFinder utilizes a zero-shot, offboard Multimodal Large Language Model (MLLM) policy to process linguistic c

    vla
  73. arxiv:2609.37918 · cs.LG
    SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation
    Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami

    Reasoning across videos requires aligning events, matching identities, comparing motion, and integrating partial observations. Evaluating these capabilities and testing how to improve them requires both reliable labels and targeted supervision. We introduce SYNCR, a simulator-grounded framework that connects these two needs through shared task generators. Built on Habitat, Kubric, and CLEVRER, SYNCR derives answers from environment state and provides 4,000 evaluation questions and 15,960 training questions over disjoint videos, spanning eight cross-video reasoning tasks. Visual ablations and h

    benchmark
  74. arxiv:2609.37914 · cs.AI
    The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
    Gonçalo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke +1

    Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM). EM has been linked to persona-like representations, where fine-tuning might reduce loss by amplifying a harmful or 'evil' persona. It remains unclear which properties of the training data drive this effect: whether all harmful examples contribute approximately equally to misalignment and whether different models are equally affected by the same fine-tuning examples. In this work, we use training da

    post-trainingbenchmark
  75. arxiv:2609.37907 · cs.CV
    Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics
    Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio

    Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training. Inverse Dynamics Models (IDMs) have thus been proposed to infer inputs from frames. Large (up to 1B parameters) IDMs trained on $\sim$1K-2K gameplay hours demonstrate feasibility and cross-environment generalization at this scale, but researchers do not clarify what the key components are to recover individual actions and often report only aggregate accuracy that can mask rare-action failu

    embodiedembodied agent
  76. arxiv:2609.37905 · cs.LG
    Beyond Interaction Capacity: Estimator Scaling with Recursive Models for CTR Prediction
    Shivang Chopra, Fotis Iliopoulos, Zsolt Kira, Gaurav Menghani

    Click-Through Rate prediction, a core task in recommendation and advertising systems, relies on modeling interactions among sparse categorical features. Explicit cross networks are a central paradigm for CTR prediction, and recent progress has largely come from increasing the interaction capacity of a single predictor through deeper cross networks and more expressive cross operators. We revisit whether continually increasing interaction capacity remains the most effective way to improve predictive performance, and find that its benefits quickly exhibit diminishing returns even as capacity cont

    benchmark
  77. arxiv:2609.37902 · cs.AI
    You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
    Liang He, Jingbo Wen, Yixiong Chen, Yue Yang +3

    Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does no

    evaluator
  78. arxiv:2609.37898 · cs.AI
    Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
    Youling Huang, Tiankuo Xu, Jiaji Liu, Tong Zheng +9

    Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early train

    agentic
  79. arxiv:2609.37891 · cs.LG
    It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
    Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov +6

    Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, eg, with reasoning traces to address cold-start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called synthetic data on knowledge and skill acquisition of language models, including sma

    post-training
  80. arxiv:2609.37889 · cs.LG
    ReCAP: Retrieval-Guided Capability Reuse for Multimodal Continual Instruction Tuning
    Tao Hu, Zhinuo Zhou, Xialiang Tong, De-Chuan Zhan +1

    Multimodal continual instruction tuning (MCIT) aims to enable multimodal large language models to acquire new capabilities from sequential tasks while preserving previously learned knowledge. Existing methods primarily mitigate catastrophic forgetting by constraining parameter updates or separating task-specific adaptations. However, continual adaptation can also benefit from external knowledge that provides domain-specific information and reusable reasoning patterns for solving diverse instructions. For example, to answer "How many red cubes are to the left of the sphere?", domain knowledge c

    benchmark
  81. arxiv:2609.37882 · cs.LG
    How Many Labels Does a Language Need? Annotation Budgets and Cross-Lingual Pooling for African-Language Text Classification
    Bhanu Prakash Vangala, Sowmya Guda, Navya Vangala

    Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no pretrained weights and no accelerator. Monolingual learning curves at budgets from 25 to several thousand labels show that topic classification reaches 90\% of it

    benchmark
  82. arxiv:2609.37874 · cs.CV
    EndoPrior-GS: Dynamic Endoscopic Reconstruction with a Joint Texture Prior
    Jiaqi Huang, Shidong Wang, Tong Xin, Kabita Adhikari

    Dynamic endoscopic reconstruction is fundamental to robotic surgery and computer-assisted interventions. While 3D Gaussian Splatting (3DGS) realises real-time rendering, its application to deformable intraoperative environments remains constrained by spurious geometry and varying illuminations. To address these limitations, we introduce EndoPrior-GS, a novel pipeline that explicitly couples frame-extracted vision heuristics and estimated depth maps. EndoPrior-GS derives a joint texture prior from a tool-filtered valid tissue mask, a non-specular photometric filter, and anatomical structural sa

    benchmark
  83. arxiv:2609.37871 · cs.RO
    ExceptionDrive: A Planning-Oriented Counterfactual Corner-Case Benchmark for Autonomous Driving
    Ziyi Luo, Zhe Sun, Yehao Lu, Lei Zhou +4

    Average performance on routine driving benchmarks does not establish planner reliability under rare, safety-critical hazards. We proposed ExceptionDrive, a counterfactual planning benchmark that uses VLM-assisted screening, localized multi-view editing, and quality auditing to insert hazards into real nuScenes scenes while preserving their context. Its 21 tasks span six safety families and define hazard or conflict regions, local safety constraints, and acceptable responses. Because hazard insertion can invalidate the recorded human trajectory, our reference-free protocol evaluates edited pred

    agentbenchmark
  84. arxiv:2609.37868 · cs.LG
    Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR
    Doohyuk Jang, Yoonsik Park, Gyouk Chu, Sihwan Park +1

    Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, successful trajectories missing from one model's rollouts may already have been discovered by another. Indeed, we observe that heterogeneous models often succeed on complementary prompts, creating opportunities for mutual learning without a designated stronger teacher. To exploit this complementarity, we

    benchmark
  85. arxiv:2609.37864 · cs.AI
    AgentBug-Smith: Automatically Reproducing Real-World Harness Bugs in Agentic Systems
    Yiming Cheng, Alfin Wijaya Rahardja, Mengshi Zhang, Zihao Chen +2

    Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and fixed number of executable harness bugs while requiring hundreds of human hours to construct. This work presents AgentBug-Smith, an automated harness bug reproduction approach that continuously discovers and reproduces real-world harness bugs from open-source agentic systems. Across different backbone LLMs, AgentBug-Smith consistently outperforms existing bug reproduction techniq

    agentagenticself-improvingbenchmark
  86. arxiv:2609.37858 · cs.LG
    Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning
    Tianhao Qian, Ziming Hong, Chongyang Gao, Kezhen Chen +1

    Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce

    benchmark
  87. arxiv:2609.37853 · cs.CL
    AnthroDial: Benchmarking LLM Anthropomorphism in Autonomous Social Interaction
    Wentao Liu, Xi Chen, Siyu Song, Biao Yuan +12

    Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or persona consistency. Agents must autonomously decide whether, when, and how to communicate while adapting to evolving contexts, goals, and relationships. Existing research, however, lacks a unified approach to enabling, evaluating, and improving such capabilities in continuous, open-ended interaction. We introduce AnthroDial, a unified framework for developing anthropomorphic social agents from three complementary aspects: MindFlow, a lightweight i

    benchmark
  88. arxiv:2609.37849 · cs.AI
    Is manual software optimization a thing of the past?
    Pavlin G. Poličar, Martin Špendl, Tomaž Hočevar

    Scientific software is increasingly required to process larger datasets while maintaining acceptable execution times. Software optimization traditionally requires substantial expertise in programming, algorithms, and numerical methods. Recent advances in large language models (LLMs) offer the possibility of automating much of this process. We investigate whether LLM-based agents can autonomously achieve substantial performance improvements in scientific software, including mature implementations that have already been extensively optimized by human developers. We tasked an LLM-based agent with

    agentautonomous agent
  89. arxiv:2609.37848 · cs.LG
    Evaluation Choices Shape Biomedical ML Claims: A Pediatric Pneumonia Benchmark Case Study
    Bhanu Prakash Vangala, Sowmya Guda, Latha Peddi, Navya Vangala

    Biomedical machine learning papers often compress model performance into one headline number. That number can look like a property of the model even when it depends strongly on how the benchmark was evaluated. We study this problem on the widely used Kermany pediatric chest radiograph dataset using nine image classifiers and a controlled evaluation protocol. Under the same protocol, the eight pretrained backbones differ by only 0.026 AUROC. In contrast, changing whether the backbone is frozen or fine-tuned changes AUROC by 0.044 on average, and changing the decision threshold changes balanced

    benchmarkevaluation protocol
  90. arxiv:2609.37834 · cs.AI
    Mixture of Self-Improving Branches For Agent Harness Optimization
    Haoyu Dong, Yuhang Zhou, Zihao Lin, Yifan Wu +5

    Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal poli

    agentagenticself-improvingself-improvementbenchmark
  91. arxiv:2609.37831 · cs.CV
    ReCaVSR: One-Step Streaming Diffusion Video Super-Resolution with Recycled Latents and Learned Cache Routing
    Xijun Wang, Xin Li, Suhang Yao, Zirui Lang +2

    Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise generative fidelity. We propose ReCaVSR, a Wan2.2-based, one-step framework for streaming VSR that builds on two observations: recycled SR latents retain local temporal context, reducing the need for full historical Key-Value (KV) caches; and individual transformer layers benefit from distinct temporal scopes. ReCaVSR combines three complementary designs: (i) layer-wise cache routing with recycled SR latents: each DiT layer learns its KV-cache tempo

    memorybenchmark
  92. arxiv:2609.37825 · cs.LG
    Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR
    Kun Liang, Chenming Tang, Clive Bai, Weijie Liu +4

    Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges on reliable value estimation, a difficult task requiring the critic to both assess progress toward a correct solution and anticipate an evolving policy's future behavior; errors in either can compromise credit assignment and destabilize online training. In this paper, we revisit

    benchmark
  93. arxiv:2609.37819 · cs.AI
    Making Duplicate Reimbursement Unrepresentable: A Verified Ethereum E-Invoice System for Humans and AI Agents
    Jia Cai

    Electronic invoices are replacing paper invoices worldwide, but today's centralized architectures leave three problems unsolved on the consumption side: an invoice can be submitted for reimbursement repeatedly, authenticity is difficult for recipients to verify, and data is siloed at a central authority that forms both a performance bottleneck and a single point of failure. This paper presents the design, formal analysis, and implementation of a complete blockchain-based electronic invoice system on Ethereum. We formalize the invoice lifecycle as a guarded labeled transition system and prove,

    ai agentbenchmark
  94. arxiv:2609.37818 · cs.AI
    Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue
    Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang +4

    Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the perception-reasoning gap. In addition, CoT may not fully capture acoustic cues in words, and generating it adds inference latency. To address these limitations, we introduce LoopSLM, which builds on looped Transformers for latent reasoning, reusing a decoder block to refine hidden states with acous

    benchmark
  95. arxiv:2609.37810 · cs.RO
    Explore, Execute, Evolve: A Skill Acquisition and Reuse Loop for Embodied Agents
    Sicheng Xie, Yitong Chen, Haidong Cao, Shunlin Lu +2

    Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, execut

    vision-language-actionembodiedtactileliberoagentembodied agent
  96. arxiv:2609.37801 · cs.CV
    ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding
    Thomas A. O'Shea-Wheller

    The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with tracklet similarity matching to robustly deal with object occlusions. However, this strategy is nevertheless vulnerable to erroneous track reclassification and identity switching, as detection confidence scores dictate association priority. To address this, I present a simple enhancement of the ByteTrack architecture, named ByteTraX, that optimises track continuity via a single unified matching threshold

    benchmark
  97. arxiv:2609.37793 · cs.RO
    MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation
    Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao +8

    World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canva

    manipulationliberorobotwinbenchmark
  98. arxiv:2609.37791 · cs.AI
    A neural network that maintains and retrieves memories based on context
    Hayoung Song, JeongJun Park, Qihong Lu, Giacomo Vedovati +3

    Every day, people continuously infer situational context and adjust the way they understand and remember the world. Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this modulation remains limited. Here, we train a recurrent neural network (RNN), augmented with an episodic memory buffer, to infer context using Bayesian inference as it continuously makes predictions of upcoming scenes while watching naturalistic movies. When the inferred context modulates the RNN's recurrent connectivity (the basis of worki

    memoryepisodic memory
  99. arxiv:2609.37788 · cs.AI
    A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
    Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo

    Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category.

    benchmark
  100. arxiv:2609.37786 · cs.CV
    CHOQOLATE: Organizing Concept Bottleneck Latent Spaces with Choquet Integrals
    Rémi Kazmierczak, Johanne Cohen, Marianne Clausel

    Concept Bottleneck Models (CBMs) built on vision-language models such as CLIP represent a latent space as human-understandable concepts. These representations are unfaithful: related concepts are entangled, so individual scores do not reflect their intended meaning. We propose CHOQOLATE, an interpretable-by-design layer based on 2-additive Choquet integrals, which merges correlated concepts into compact nodes. Across four datasets, CHOQOLATE achieves a favorable accuracy-interpretability trade-off, with weight-sparse and semantically coherent nodes. A closed-form gradient derivation, backed by

    benchmark
  101. arxiv:2609.37783 · cs.CV
    A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System
    Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam +9

    Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and phone photographs may be used to establish presence, sequence, causation, damage, or identity, yet contemporary generative systems allow non-experts to alter or fabricate such images through ordinary prompt-based interfaces. Existing image-forensics benchmarks provide important resources for face manipulation, classical tampering, and general synthetic-image detection, but they ar

    manipulationbenchmark
  102. arxiv:2609.37782 · cs.CL
    Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings
    Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang +3

    Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first pa

    long-contextbenchmark
  103. arxiv:2609.37776 · cs.RO
    Geometry-Preserving Human-to-Robot Upper-Body Motion Retargeting from Monocular Video
    Xiaoyu Yang, Sen Han, Da Li, Nan Wu

    Monocular RGB video provides an accessible source of human demonstrations for upper-body robot motion, yet video-driven human-to-robot transfer remains challenging because body and hand motion are recovered at different spatial scales, human and robot kinematics differ substantially, and fine distal motion is difficult to preserve across embodiments. We present a geometry-preserving motion-retargeting framework that integrates unified body--hand reconstruction with morphology-independent geometric transfer. Frame-wise body estimates, video-level observations, and detailed hand evidence jointly

    dexterousgrasp
  104. arxiv:2609.37775 · cs.CV
    HiRAE: Hierarchical Representation Autoencoding with Residual Budgets
    Xuanyu Zhu, Yan Bai, Yang Shi, Yihang Lou +4

    Pretrained visual representations support image generation, but may not fully preserve the fine-grained details needed for faithful reconstruction. Meanwhile, intermediate encoder layers contain complementary visual details, but learning to fuse them for reconstruction can produce a latent distribution that is difficult to model. Existing fusion methods require empirical tuning of layer selection or staged optimization of fusion and decoding, increasing configuration effort or training complexity. We introduce HiRAE (Hierarchical Representation Autoencoder), which learns a hierarchical fusion

    evaluation protocol
  105. arxiv:2609.37773 · cs.AI
    OmniVCBench: Benchmarking Evidence-Grounded Multimodal Reasoning Towards AI Virtual Cells
    Manyu Li, Xunkai Li, Yongfu Xiong, Yi Liu +2

    Artificial Intelligence Virtual Cells (AIVCs) are envisioned as scientific agents that simulate cellular responses, explain underlying mechanisms, and support hypothesis-driven discovery. Existing AIVC benchmarks, however, operate primarily at the simulation layer, motivating complementary evaluation of how models interpret experimental evidence and formulate biological hypotheses. We introduce OmniVCBench, a figure-centric, source-traceable benchmark for the interpretation component of an AIVC. It contains 6,077 curated single- and multi-subfigure question--answer pairs derived from figures a

    benchmark
  106. arxiv:2609.37772 · cs.RO
    Urgent Actions Go First: Urgency-Aware Denoising for Real-Time VLA Control
    Zibo Wang, Haochen Han, Pengzhen Ren, Mingtong Dai +1

    Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies. We exploit this asymmetry to introduce Urgency-Aware Denoising (UAD), a novel inference-time framework that allocates denoisi

    vision-language-actionvlamanipulationbenchmark
  107. arxiv:2609.37771 · cs.RO
    Faster and Better? Benchmark Bugs and Design Limitations Distort the Evaluation of Vision-Language-Action Acceleration
    Qiwei Chen, Kaijun Zhou, Nuohui Shi, Zhiyang Li +2

    Simulated manipulation benchmarks are the standard tool for evaluating vision-language-action (VLA) policies and the acceleration methods that reduce their inference latency for on-robot deployment. On these benchmarks, we observe that some training-free acceleration methods, which approximate the baseline policy's computation, achieve higher measured success rates than the baseline itself. Success rates alone cannot establish whether such gains come from better task execution or from evaluation flaws. We therefore investigate two kinds of benchmark flaws behind these gains: bugs, where the im

    vision-language-actionvlamanipulationliberorobotwinbenchmark
  108. arxiv:2609.37759 · cs.CV
    Selective Channel Restoration for Backdoored Vision-Language Models
    Shuming Liu, Zhifang Zhang, Suqin Yuan, Khin Mi Mi Aung +2

    Vision-language models (VLMs) exhibit strong multimodal capabilities but remain vulnerable to backdoors implanted through poisoned fine-tuning data. Existing defenses often require extensive parameter updates during fine-tuning or incur per-query overhead during inference. To address these limitations, we propose Perturb-Select-Restore (PSR), a post-training defense that performs sparse updates to the projection interface and introduces no additional computation during inference. We reveal that backdoored VLM projectors are substantially more sensitive to bounded perturbations than clean VLM p

    post-training
  109. arxiv:2609.37755 · cs.CL
    Which papyrus HTR is good enough? Character-error-rate tolerance of four papyrological tasks on Greek texts
    Anton Repushko, Elena Chepel

    Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would let scholars discover documents and literary works that have so far gone unread. Recognition systems for Ancient Greek papyri are in statu nascendi, and how accurate they must be for a given papyrological task has not been examined. To answer this and set a benchmark for Greek papyrus HTR, we test a range of character error rates (CER) against four papyrological tasks, using published editions as ground truth. Methods: From 63,846 current editio

    benchmark
  110. arxiv:2609.37751 · cs.AI
    Cross-Entropy Guided Routing in Mixture-of-Experts Large Language Models
    Yury Nahshan, Nati Daniel, Jacob Goldberger, Yoli Shavit

    Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized with load balancing terms and learn affinity scores through the language-model objective. However, these objectives do not provide direct alignment between routing affinities and token-level error. We introduce token-error supervision for sparse routing in two forms. The first form predicts an error score per expert. The affinity-weighted aggregate of these scores is aligned to the next-token cross-entropy loss, while the individual scores

    benchmark
  111. arxiv:2609.37750 · cs.CV
    Multi-Site Real-World Performance of Commercial AI for Pulmonary and Incidental Pulmonary Embolism Detection
    Aawez Mansuri, Mohammadreza Chavoshi, Theodorus Dapamede, Wasif Bala +12

    Pulmonary embolism (PE) is a leading cause of cardiovascular mortality, yet the real-world performance of FDA-cleared AI detection models remains incompletely characterized. We retrospectively evaluated two FDA-cleared AI algorithms from a single commercial platform (Aidoc Medical BriefCase), one for PE triage on dedicated CT pulmonary angiography (CTPA; n = 30,678) and one for incidental PE (iPE) detection on routine contrast-enhanced CTs (n = 37,191), across a 17-facility academic health system. Reference-standard labels were extracted from radiology reports using a validated LLM pipeline (9

    benchmark
  112. arxiv:2609.37743 · cs.AI
    ContextRender: From Execution Dependencies to Agent Context
    Savini Kashmira, Jayanaka L. Dantanarayana, Lingjia Tang, Jason Mars

    LLM agents performing long-horizon tasks accumulate tool results that later steps may need. Passing the full history to every invocation is costly even when it fits within the context window, while reducing it risks omitting needed information. Existing context management methods can overlook how earlier tool results are used in subsequent execution, leaving needed information out of context. We introduce ContextRender, which manages context through a persistent graph of execution dependencies. We develop Tool-Flow Analysis to track how later operations reuse information from earlier tool resu

    agentllm agent
  113. arxiv:2609.37725 · cs.LG
    Context Language Models
    Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li +9

    We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59%

    agentmulti-agentagent system
  114. arxiv:2609.37721 · cs.RO
    CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
    Sen Wang, Liu Liu, Xinjiang Wang, Zequn Chen +6

    Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple act

    manipulation
  115. arxiv:2609.37713 · cs.CL
    Billiger.de Products: A Bilingual Entity Matching Benchmark
    Aaron Steiner, Ksenia Elagin, Ralph Peeters, Johannes Knopp +1

    Existing product matching benchmarks primarily contain English-language product data and are often dominated by a single product category, such as electronics. This paper introduces Billiger.de Products, a bilingual German and English entity matching benchmark covering thirteen consumer product categories, including difficult-to-handle categories such as clothing and furniture. The benchmark data originates from the German price comparison platform billiger.de. Following the design of WDC Products, the benchmark offers multiple variants that differ in the fraction of corner cases, the size of

    benchmark
  116. arxiv:2609.37712 · cs.CV
    PolyOCR-Venus: Unified OCR Foundation Models for Text-Centric Visual Intelligence
    GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu +21

    Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, localize, and reason over textual information in complex visual environments. However, existing OCR systems often excel at only some tasks and struggle to balance recognition, parsing, and reasoning across scenarios. In this report, we present PolyOCR, a family of unified OCR foundation models of varying scales. PolyOCR combines a shared instruction-following framework with a large-scale data engine that converts heterogeneous visual resources into qu

    benchmark
  117. arxiv:2609.37709 · cs.CV
    VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation
    Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki, Yutaka Matsuo +1

    Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple references must be composed under multiple and heterogeneous visual-instruction images. To address this gap, we introduce VIF-Bench, a benchmark of 1,241 tasks designed to assess the edge of model capabilities in this joint setting by covering: (i) multi-reference generation (up to 7)

    benchmark
  118. arxiv:2609.37708 · cs.LG
    Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics
    Ojas Shirekar, Yash Surange, Agustinas Jučas, Chirag Raman

    Human social behaviour is not a collection of independent motions, but a jointly organised process in which group dynamics and individual variation continuously shape one another. Yet existing social motion models often prioritise plausible trajectories while leaving interaction state implicit, limiting their ability to transfer across groups, tasks, and partial-observation regimes. To address this gap, we introduce Bilevel Representations for Agent Interaction Dynamics (BRAID), a hierarchical sequential latent-variable model for generative multi-person interaction. BRAID explicitly formulates

    embodiedlatent dynamicsagentagent system
  119. arxiv:2609.37702 · cs.LG
    Width Expansion as a Method for Class Incremental Learning
    A. L. S. Conde, Y. Elkhatib, C. M. Ranieri

    Class Incremental Learning (Class-IL) requires models to learn new classes over time while preserving previously acquired knowledge without access to past data or task identity. This setting intensifies the stability-plasticity dilemma and makes catastrophic forgetting a central challenge. Existing approaches include regularization, knowledge distillation, replay, and architectural expansion. However, many expansion methods rely on explicit task identifiers or predefined growth strategies, limiting their applicability when task boundaries are unavailable at inference time. This work proposes a

    memory
  120. arxiv:2609.37700 · cs.AI
    Locating Answer-Correctness Signals in Frozen Large Language Models
    Yuansen Liu, Yixuan Tang, Anthony Kum Hoe Tung

    Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search

    retrieval-augmented
  121. arxiv:2609.37694 · cs.LG
    GARDiff: Graph-Aligned Residual Diffusion for Probabilistic Multivariate Time-Series Forecasting
    Rui Han, Min Yang, Xu Zhang, Xinghao Yang +2

    Diffusion models have recently shown strong potential for probabilistic multivariate time-series forecasting by modeling complex conditional distributions. Recent decoupled diffusion frameworks further separate forecasting into deterministic prediction and stochastic residual generation, making it natural to derive dependency graphs from deterministic representations and use them to guide residual diffusion. However, we show that this direct structural transfer is unreliable. Although deterministic-derived graphs encode useful global dependency priors, they exhibit substantial edge-level misal

    benchmark
  122. arxiv:2609.37690 · cs.CV
    Honeycomb: Constant-Size Scene Memory Representation for Video World Models
    Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen, Yufeng Weng +5

    Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memory systems accumulate RGB observations or latent features, causing storage requirements to grow as generation proceeds. We introduce **Honeycomb**, a video world model built on **HexMemory**, a compact low-rank representation that stores scene features in a fixed-size memory comprising six spatial and spatiotemporal planes. A feed-forward writer maps each newly generated video chunk to plane features. As the spatial coverage or temporal range expands, HexMemory

    world modelmemory
  123. arxiv:2609.37686 · cs.AI
    EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
    Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun +18

    Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from

    autonomous agentbenchmark
  124. arxiv:2609.37682 · cs.CV
    Med-RADIO: Reducing All Medical Domains Into One via Multi-Teacher Distillation
    Chu Zhang, Haoyu Jiang, Hongyuan Zhang, Hongbin Liu +1

    The rapid expansion of large-scale medical datasets and computational resources has driven significant progress in medical foundation models. Given the inherent heterogeneity of medical imaging modalities, current research mainly follows two paths: specialized models optimized for specific modalities, and generalist models designed to handle multiple modalities. However, medical generalist models suffer from both insufficient training data scale relative to natural image generalists and inadequate domain-specific depth relative to medical specialists. Empirically, generalist models establish a

    benchmark
  125. arxiv:2609.37681 · cs.RO
    ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback
    Ahmed Nader Ahmed, Omar Moured, Mughni Irfan Mohammed Abdul, Muhayy Ud Din +1

    Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scena

    manipulation
  126. arxiv:2609.37677 · cs.RO
    Learning Expressive and Compositional Motion Representation via Spectral Skills
    Feiyang Wu, Chenxiao Gao, Chen Yang, Ye Zhao +2

    Robotic foundation models offer a promising path toward general-purpose humanoid robot control, often through hierarchical architectures. However, their effectiveness depends on the command interface between the planner and the controller, which must support accurate execution while remaining easy to predict, and ideally allow new behaviors to be composed from prior ones. In this work, we introduce spectral skills, a latent representation of this interface that meets these requirements through predictive representation learning. By design, spectral skills compactly encode short motion segments

    humanoid
  127. arxiv:2609.37673 · cs.AI
    KUPAS MASTER: Distilling the Tacit Expertise of Master Practitioners into Agent-Ready Experience Corpora
    Changmian Wang, Yuchao Ma, Xuchao Lu, Chen Zhang +12

    Experienced professionals know more than just facts and conclusions. They know which cues matter, why a judgment is reasonable, and which action to take. Routine work records often leave out this tacit knowledge, making it difficult for Large Language Model (LLM) agents to use professional experience effectively. We introduce KUPAS MASTER, an experience engineering platform built around nine-layer cognitive corpus construction. It turns heterogeneous work records and practitioner interviews into traceable, reusable experience corpora for agents. Six case elements preserve the task process: con

    retrieval-augmentedragagent
  128. arxiv:2609.37669 · cs.AI
    Retrieve, Reproduce, Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection
    Sabrina Kaniewski, Tim Krämer, Julius Bächle, Markus Enzweiler +2

    Retrieval-Augmented Generation (RAG) is increasingly used to enhance Large Language Model (LLM)-based software vulnerability detection by grounding predictions in retrieved vulnerability knowledge, such as vulnerability reports. However, existing RAG-based software vulnerability detection (RAG4SVD) systems are often evaluated using proprietary models, which challenges open science and reproducibility. Further, studies use different datasets, custom knowledge bases, different backbone models, and diverse metrics, which hinders meaningful cross-system comparison. In this work, we study six open-

    retrieval-augmentedbenchmark
  129. arxiv:2609.37666 · cs.RO
    Semantic Map Sharing and Capability-Aware Coverage Planning for AI-Native 6G Robotic Coordination
    Abdulqader Dhafer, Qi Wang, Zhou Daniel Hao

    Search and Rescue (SAR) operations increasingly deploy heterogeneous teams of aerial and ground robots. However, conventional coverage methods typically do not translate perceived terrain into platform-specific reachability, while continuous image exchange imposes a high communication cost. We propose an edge-centric, semantic-aware coverage planning framework that integrates aerial terrain perception, robot-specific traversability reasoning, and payload-efficient semantic state sharing. Aerial observations are converted into compact semantic grid maps, enabling reachability-constrained area d

    benchmark
  130. arxiv:2609.37664 · cs.LG
    Learning Causal Normalizing Flows from Incomplete Data via Observed-Data Likelihood
    Trung-Dung Hoang, Alceu Bissoto, Tim Flühmann, David Herzig +3

    Causal Normalizing Flows (CNFs) enable causal inference from observational data given the causal structure, but they assume fully observed training data. We introduce MissCNF, which trains CNFs directly on incomplete data by maximizing the marginal likelihood of each partially observed sample, without discarding rows or constructing a completed dataset. Thanks to the causal structure encoded in the autoregressive factorization of CNFs, only missing variables in the ancestral closure of the observed set are integrated out, while the others are dropped without computation. We further establish t

    benchmark
  131. arxiv:2609.37661 · cs.CL
    Corpus-Guided Dual-Path Propagation for Graph Retrieval-Augmented Generation
    Baoxian Liu, Tong Wei

    Graph-based retrieval-augmented generation supports multi-hop retrieval by organizing corpus information into graphs. However, existing relation-free graph retrieval methods rely primarily on query-sentence similarity to search for evidence. This can exclude useful bridging evidence with low query similarity and activate incidental entities unrelated to the reasoning chain. In this paper, we propose a simple and effective approach called NexusRAG, which augments the relation-free Tri-Graph with a corpus-level entity neighborhood structure derived from joint entity co-occurrence and semantic si

    retrieval-augmentedbenchmark
  132. arxiv:2609.37658 · cs.AI
    EnterpriseBench: Benchmarking LLM Agents on Enterprise-Level Strategic Reasoning and Decision-Making
    Min Yang, Yichen Pan, Jinghua Piao, Dandan Song +2

    LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowledge, and financial QA, leaving interactive and long-horizon decision-making underexplored. To bridge this gap, we introduce EnterpriseBench, a benchmark that evaluates LLM agents across this spectrum, from static question answering to dynamic decision-making. Specifically, Enterp

    agentllm agentbenchmark
  133. arxiv:2609.37656 · cs.CV
    Tracing the Evidence: Faithful Token Attribution Through Vision-Language Reasoning
    Bowen Yuan, Danny Wang, Ruihong Qiu, Zijian Wang +1

    Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response by assigning scores that rank image and prompt tokens by how much the model relies on them, such that removing higher-ranked tokens causes the likelihood of the generated response to drop more rapidly. However, existing token-attribution methods have been developed mainly for text-based language models, and our empirical study reveals two challenges when complex mu

    benchmark
  134. arxiv:2609.37655 · cs.CV
    Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding
    Jiayu Ying, Qijian Tian, Ruijie Xu, Xinnan Zhu +3

    Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce Exemplar2VQA, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously d

    embodiedsim-to-realmulti-agentbenchmark
  135. arxiv:2609.37647 · cs.AI
    Evaluating and Benchmarking the System One Model Jev
    Tobias Deußer, Lorenz Sparrenberg, Rafet Sifa

    Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed options, a position on a rubric, or the probability that a statement is true, with probabilities the vendor describes as calibrated. Such models target small decisions in information access pipelines, such as routing queries, checking grounding, moderating content, or rating against a rubric. We evaluate Jev (jev-1.13.0) zero-shot on 37 datasets spanning classification, routing, natural language inference, reading comprehension, commonsense reason

    benchmark
  136. arxiv:2609.37644 · cs.AI
    Beyond a single latent space: a dual-latent world model for long-horizon planning
    Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He +5

    Latent world models often struggle with long-horizon planning despite accurate short-term predictions. Recursive rollouts accumulate errors, while distance concentration in high-dimensional latent spaces can weaken goal discrimination. We introduce the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning through distinct state representations and dynamics models. The low-level model predicts action-conditioned transitions, while the high-level model uses learned macro-actions to plan over longer temporal spans. We also propose Long-Horizon Representation L

    world modelaction-conditioned
  137. arxiv:2609.37633 · cs.LG
    RLTL;DR: Self-improvement by Internalizing Self-generated Feedback
    Michael Kirchhof, Eleonora Gualdoni, Andrew Szot, Khashayar Gatmiry +5

    The common paradigm of reinforcement learning with verifiable rewards (RLVR) is to let agents make multiple attempts at a task, and optimize towards the successful ones. This becomes problematic in the realms of self-improvement, where tasks are so difficult that the agent has a low or even no chance of success, and where there are no teacher models or example solutions to distill from. In this paper, we introduce RLTL;DR. After each failed attempt, we show the policy the verifier outputs and let it write its own feedback, in the form of a single TL;DR insight. The next rollout is conditioned

    agentself-improvementeval
  138. arxiv:2609.37632 · cs.LG
    ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation
    Fariza Rashid, Duc Van Le, Rahat Masood, Gustavo Batista +2

    Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or unrepresentative. To address this issue, we propose ProCTI, a diffusion-imputation framework that augments local conditioning with retrieved global dataset-level

    benchmark
  139. arxiv:2609.37628 · physics.optics
    Operational theory for photonic circuits: generalizing linear optics beyond quantum theory
    Ismaël Septembre, Matthias Kleinmann, Martin Plávala

    We use the analogy between the quadrature operators in quantum optics and the phase-space coordinates in a mechanical harmonic oscillator to generalise the theory of linear optics beyond quantum theory. For this, we introduce a framework for analysing photonic circuits within generalised probabilistic theories based on quadrature operators. Using this framework, we ask whether phenomena like the Hong-Ou-Mandel effect, Mach-Zehnder interferometry are specific to quantum theory or also occur across a broader class of theories. We find the latter: the vanishing photon coincidences at the output o

    mach-zehnder
  140. arxiv:2609.37626 · cs.AI
    SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving
    Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun +5

    No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long prompts favor context parallelism. Reasoning, agentic, and RL-rollout workloads make a fixed choice untenable: a batch that begins as many short requests ends as a few very long ones, so the best layout changes while the same requests run. Serving engines nevertheless fix one layout at launch, because changing it has meant draining requests and restarting workers. We present SPLASH, a serving sys

    memoryagentic
  141. arxiv:2609.37616 · cs.LG
    Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
    Abhinav Rajeev Kumar, Paras Chopra

    Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and

    post-training
  142. arxiv:2609.37605 · cs.LG
    TomoTransformer: Towards a Foundation Model for CT Reconstruction
    AmirEhsan Khorashadizadeh, Benjamín Béjar

    Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts miss

    benchmark
  143. arxiv:2609.37602 · cs.RO
    When to Adapt: Multi-Signal Domain Shift Detection for Efficient Training-Free Adaptation in Open-Vocabulary Segmentation
    Michele Antonazzi, Alejandra C. Hernandez, José Araujo, Olov Andersson +1

    Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapte

    benchmark
  144. arxiv:2609.37599 · cs.RO
    BlenDAgger: Blended Shared Control for Interactive Imitation Learning
    Cailyn Smith, Geoffrey Sun, Henny Admoni, Zackory Erickson

    Robot policies are frequently trained from human corrections, yet teleoperating a robot to provide corrections is burdensome, and human demonstrators are not always optimal. We propose Blended DAgger (BlenDAgger), an approach for collecting data to train imitation learning policies by using shared control to blend the policy's and demonstrator's actions during interventions. By blending human and policy actions, we aim to improve the autonomous performance of manipulation policies. We validate our approach across five manipulation tasks, two in the real world and three in simulation. Our appro

    manipulation
  145. arxiv:2609.37591 · cs.RO
    Credit-Guided Policy Improvement for Test-time Adaptive Vision-Language Navigation
    Yang Li, Sijia Zhang, Yihan Li, Aming WU +3

    Test-time adaptation for vision-language navigation (TTA-VLN) enables pretrained policies to adapt online to unseen environments using only test-time observations and interaction history. However, distribution shifts can distort local action preferences and lead to off-course decisions. Existing methods rely on predictive uncertainty, trajectory-level feedback, or accumulated adaptation experience to correct such deviations. These signals, however, do not directly reveal whether an executed action supports instruction-guided progress toward the goal. Moreover, a plausible corrective signal doe

    sim-to-realbenchmark
  146. arxiv:2609.37590 · cs.AI
    FOCUS: Training-Free Decision-Preserving Context Compression for LLM Agents
    Shantanu Dixit, Anson Bastos, Xuchao Zhang, Chetan Bansal +1

    LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's fu

    context compressionllm agentagenticbenchmark
  147. arxiv:2609.37587 · cs.LG
    ReLMem: Learning Recurrent Memory for Longitudinal EHR Modeling
    Zijie Meng, Xiwei Dai, Yingying Zhang, Jian Wu +2

    Longitudinal electronic health record (EHR) modeling requires integrating new visits with an expanding patient history. Yet the continual accumulation of clinical information imposes increasing computational and memory costs on large language models (LLMs) when they process and retain complete patient histories. A practical alternative is visit-wise recurrent compression, which incorporates each incoming visit into a compact, continually updated patient memory. However, under a fixed memory budget, successive updates must integrate new information without progressively losing critical historic

    memory
  148. arxiv:2609.37583 · cs.RO
    RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation
    Shifeng Bao, Fanding Huang, Yihan Lin, Youhe Feng +6

    Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies

    manipulationagentself-improving
  149. arxiv:2609.37577 · cs.CL
    Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency
    Bruno Brocai, Maria Becker

    Large Language Model judges are widely used to rank texts and text-generating systems through pairwise comparison, and their reliability is typically assessed via three proxies: position bias, transitivity, and pairwise agreement (self- or human-labeled). Because these proxies drive judge selection and benchmarking, a substantial literature reporting that judges perform poorly on them risks steering practitioners away from otherwise capable evaluators. We argue this assessment is misleading. Under the Bradley--Terry geometry underlying pairwise aggregation, each proxy is dominated by close-ran

    benchmarkevaluator
  150. arxiv:2609.37576 · cs.CV
    Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment
    Yu Zhao, Jiarui Wang, Huiyu Duan, Ye Zhao +4

    With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study semantic understanding, quality perception, and authenticity identification in isolation, while largely neglecting responsibility detection. This leaves a gap in unified and comprehensive validation. To bridge this gap, we introduce SQUARE-Bench, a comprehensive benchmark that systema

    benchmarkevaluator
  151. arxiv:2609.37574 · cs.CL
    MERGE: Multi-LLM Ensemble for Retrieval via Generative Enrichment
    Tzu-I Ho, Yung-Yu Shih, Shang-Yu Su, Dongzhe Wang +1

    Large Language Models (LLMs) are increasingly used to enrich user queries in information retrieval (IR) so that a standard retriever such as BM25 can bridge vocabulary gaps with the target corpus. Any single LLM, however, is limited by its training data and architectural biases, and its enrichment behavior depends on hand-crafted prompts that must be re-engineered for each new model -- an expensive and poorly scalable process. We present MERGE (Multi-LLM Ensemble for Retrieval via Generative Enrichment), a two-stage framework: three heterogeneous 7-8B open-source LLMs independently produce can

    benchmarkevaluator
  152. arxiv:2609.37569 · cs.CV
    Decompose Radicals, Then Reward: Fine-Grained Inspection for Accurate Chinese Text Rendering
    Yazhen Xie, Xingsong Ye, Zhineng Chen

    Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcripts with target strings. Such rewards overlook the compositional nature of Chinese writing: an ideograph consists of reusable components arranged through explicit spatial relations, yet OCR evaluates it as an atomic character. Consequently, visually different radical-level errors may receive equally coarse feedback, encouraging glyphs that merely resemble the target instead of faithfully reproducing its internal structure. We employ Ideographic

    post-training
  153. arxiv:2609.37568 · cs.CL
    Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
    Yu Zhang, Pingrui Zhang, Xuefeng Bai, Pengfei Zhang +2

    Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficient

    benchmark
  154. arxiv:2609.37567 · cs.AI
    Concealing LLM-Based Multi-Agent Topology via Phantom Structure Injection
    Longzhu He, Zelang Wen, Xinfeng Li, Sen Su +1

    Driven by the rapid advancement of large language models (LLMs), LLM-based multi-agent systems (MAS) have emerged as a powerful paradigm for collaborative reasoning over complex tasks. A key design element of MAS is the communication topology, which governs information flow among agents and often encodes proprietary knowledge about the system architecture. However, recent work has shown that such topologies can be inferred even in black-box settings by exploiting semantic dependencies in observable reasoning traces, posing significant risks of intellectual property leakage and exposure of syst

    multi-agentagent systembenchmark
  155. arxiv:2609.37566 · cs.MA
    RAVEN: Receiver-Conditioned Action-Value Encoding for Finite-Alphabet Multi-Agent Communication
    Shuwei Sun, Chenxi Wang, Jian Huang, Weiyun Ru +1

    A message drawn from a small alphabet helps a teammate only if it keeps the distinctions that change that teammate's next decision. We show that scoring messages by action values averaged over the receiver's situation can erase exactly these distinctions, and we propose RAVEN (Receiver-conditioned Action-Value ENcoding), which trains a four-symbol, one-step-delayed channel to preserve each receiver's centered action-value profile within the receiver's own context. The sender never needs to know that context: the receiver decodes every symbol with its private information. We give two estimators

    multi-agent
  156. arxiv:2609.37565 · cs.LG
    A Model-Agnostic Physics-Guided Adapter for Few-Shot Transfer of Coastal Flood Prediction Models to Unseen Regions
    Bilal Hassan, Areg Karapetyan, Samer Madanat

    Deep learning surrogates can produce high-resolution coastal flood maps orders of magnitude faster than physics-based hydrodynamic simulators, yet transferring them to new coastal regions remains costly, since generating target-region data for fine-tuning typically requires numerous time-consuming simulations. To tackle this bottleneck, we introduce the Physics Adapter (PA), a compact, architecture-agnostic adaptation interface that enables efficient few-shot transfer of flood prediction models across diverse coastal regions. PA predicts peak water level through a differentiable wet/dry respon

    benchmark
  157. arxiv:2609.37560 · cs.RO
    RoboFin3D: A Sim-to-Real Platform for Robotic Surface Finishing
    Haowei Wen, Shangtao Li, Vaibhav Sanjay, Philip Huang +2

    Grinding and sanding are fundamental processes in industrial robotic surface finishing. However, physical trials are expensive and consume workpieces, making reproducible experiments difficult. We present RoboFin3D, a sim-to-real platform built on Isaac Sim and the Newton physics engine, that provides physics-based grinding and sanding simulation for cheap and repeatable robotic surface finishing experiments. RoboFin3D utilizes a signed distance field (SDF) to model the changing geometry of the workpiece, enabling contact computation, live updates and rendering without an intermediate mesh. It

    sim-to-real
  158. arxiv:2609.37559 · cs.CV
    APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
    Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu +8

    To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 549 sessions, 104 trajectories, and 2,719 candidates, spanning both objective and open-ended question

    memorypersistent memorybenchmark
  159. arxiv:2609.37557 · physics.app-ph
    Möbius Reparametrization of Multiport-Network Models of PIN-Diode-Programmable Metasurfaces for Accurate Low-Order Neumann Approximations
    Philipp del Hougne

    Accurate models of programmable metasurfaces based on multiport-network theory (MNT) account for mutual coupling (MC) through a configuration-dependent matrix inversion. The latter's high computational cost during gradient-based optimization (GBO) can be alleviated via a finite-order Neumann approximation. We show that the accuracy of this approximation depends strongly on the (tacitly) chosen MNT parametrization. For 1-bit-programmable meta-elements (e.g., meta-elements based on PIN diodes), we identify a closed-form Möbius reparametrization that depends only on the meta-element's two load st

    memorybenchmark
  160. arxiv:2609.37555 · cs.LG
    Benchmarking graph-based models for in-silico toxicity prediction in drug discovery
    Noel Suarez-Barro, Manuel Lama, Juan C. Vidal

    Drug discovery is a costly and high-risk process, where toxicity-related failures remain a major cause of attrition in both preclinical and clinical stages. As a result, accurate early prediction of chemical toxicity is essential to reduce downstream costs and improve compound prioritization. In this context, graph deep learning (GDL) has emerged as a powerful paradigm for toxicity prediction, leveraging molecular graph representations to learn directly from chemical structure with improved expressivity over traditional approaches. Despite the growing number of proposed models, current literat

    benchmarkevaluation protocol
  161. arxiv:2609.37554 · cs.RO
    Risk-Aware Semantic Grounding for Trustworthy LLM-Based Robot Planning
    Łukasz Sobczak, Nur Keleşoğlu, Sławomir Piotr Nowak

    Large language models (LLMs) are increasingly used as high-level planners in robot navigation, but their outputs may become unreliable when instructions are ambiguous, unsupported by the environment, or semantically inconsistent. This paper presents a Risk-Aware Semantic Grounding framework for trustworthy LLM-based robot planning. Unlike existing LLM-based planners that primarily optimize plan generation, we formulate semantic grounding reliability as a multi-dimensional risk estimation problem. The proposed architecture explicitly models grounding uncertainty through ambiguity, hallucination

    benchmark
  162. arxiv:2609.37552 · cs.RO
    Wrench-ACT: Enhancing Robot Policies for Contact Rich Behavior Using Direct Wrench Control
    Johannes Hechtl, Yannik Blei, Simon Ball, Reihaneh Mirjalili +4

    While contact-rich manipulation requires deliberate regulation of interaction forces, recent approaches to robot manipulation learning predominantly represent actions as target positions or poses. Even methods that incorporate force sensing either use it solely as an observation or, when predicting forces as part of the output, rely on a hybrid force controller. In this paper, we propose an imitation learning policy that predicts wrenches as its sole action output for direct use by a pure force controller. Our studies suggest that force-domain imitation learning depends critically on data coll

    manipulationteleoperationaction chunking
  163. arxiv:2609.37544 · cs.AI
    How Can Recommendation Feedback Evolve Agent Memory?
    Shanwen Mao, Mingming Li, Hao Zhang, Zhiheng Li +3

    Content-generation agents continuously receive impressions, clicks, conversions, and negative feedback from recommendation systems, providing real-world outcome signals for memory evolution. However, these signals are delayed and noisy, confounded by audience composition, placement, and recommendation policies, and may result from the combined influence of multiple memories, making accurate attribution difficult. Existing methods rely primarily on immediate feedback or semantic retrieval and therefore struggle to reliably translate recommendation outcomes into memory fitness. To address this c

    memoryexternal memoryagent memoryagentbenchmark
  164. arxiv:2609.37543 · cs.CL
    RunyaNER: Auxiliary Language Selection for Runyankore NER
    Prosper Arineitwe Asiimwe, Francois Meyer, Jan Buys

    Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words

    benchmark
  165. arxiv:2609.37539 · cs.AI
    SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
    Renxi Wang, Mingshan Hee, Fajri Koto, Timothy Baldwin +1

    Skills equip LLM agents with professional knowledge and guidance to complete long-horizon and complex tasks. Although skills have been widely adopted in recent agent paradigms and harnesses, how to synthesize reliable training data and how to train agents for skill use remain underexplored. In this work, we propose SkillGym, an automatic pipeline to build verifiable environments, collect trajectories, and train skill-use agents. SkillGym first crawls a large volume of skills from the internet, then keeps those whose workflows can run reproducibly offline. A builder-reviewer pipeline is used to

    agentllm agentbenchmark
  166. arxiv:2609.37533 · cs.CL
    E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
    Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets +1

    Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the

    benchmark
  167. arxiv:2609.37530 · cs.RO
    RawVLA: Embodied Neural Image Signal Processor For Robotic Manipulation
    Shuhong Liu, Heng Zhou, Lingfeng Qian, Yuhao Fang +6

    Vision-language-action (VLA) models typically operate on RGB images produced by a fixed camera image signal processor (ISP), leaving the imaging pipeline outside the learning and evaluation loop. We systematically examine the consequences of this overlooked design choice across five fundamental ISP dimensions: gain, sensor noise, chromatic response, tonal response, and bit depth. Our analysis reveals that RAW-to-RGB processing materially shapes both action prediction and manipulation success, with different ISP dimensions exerting substantially different effects. Guided by these findings, we i

    vision-language-actionvlaembodiedmanipulationbenchmark
  168. arxiv:2609.37522 · cs.LG
    Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers
    Xiaohan Yi, Wen Luo, Yani Huang, Junfeng Zhan +3

    On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combine

    agent
  169. arxiv:2609.37519 · cs.RO
    Video2STL: Grounding VLM-Generated Temporal Specifications for Robot Learning
    Merve Atasever, Keyan Azbijari, Cagan Bakirci, Bo-Ruei Huang +5

    Video-based policy learning is particularly promising, as it illustrates target behaviors without requiring action annotations or embodiment-matched demonstrations. A central challenge is deciding what information should be transferred from the video to the robot. Existing approaches commonly convert visual observations into scalar similarity or value signals, or ask foundation models to directly generate reward code. These approaches can make the temporal structure of a task difficult to inspect, ground, and reuse. We present Video2STL, a framework that converts observation-only videos into p

    manipulationquadruped
  170. arxiv:2609.37515 · cs.LG
    Hierarchical Compression of Vision-Language Model Benchmarks
    Hyunjong Ok, Seunggu Kang, Jaeho Lee

    Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical frame

    benchmark
  171. arxiv:2609.37510 · cs.LG
    From Dissonance to Orchestration: Teacher Intervention in On-Policy Distillation
    Yuhao Wang, Ruiyang Ren, Yinan Zhang, Ruiqing Zhang +2

    On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The pre

    benchmark
  172. arxiv:2609.37509 · cs.LG
    ScaGNN: a Graph Neural Network for Multiple Scattering Simulations
    Rémi Marsal, Stéphanie Chaillat, Alexandre Chapoutot

    The boundary element method (BEM) provides an efficient numerical framework for solving multiple scattering problems in unbounded homogeneous domains. By restricting the discretization to the domain boundaries, it substantially reduces computational complexity. The procedure first consists in determining the solution trace on the boundaries of the domain by solving a boundary integral equation. Then, the volumetric solution can be recovered at low computational cost using a boundary integral representation. As the first step of the BEM represents the main computational bottleneck, we present S

    benchmark
  173. arxiv:2609.37446 · cs.LG
    Demistifying Data and Simulator Assumptions in Supervised Causal Discovery
    Pingchuan Ma, Rui Ding, Bojun Huang, Shuai Wang

    Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pairs are typically simulated, making the simulator both a source of supervision and a carrier of assumptions about causal graphs, mechanisms, and noise. Understanding the resulting predictions therefore requires examining how these assumptions supplement the information available in observational data, which may be compatible with multiple causal graphs. This paper examines that relationship across representative methods available through June 202

    benchmark
  174. arxiv:2609.37443 · cs.AI
    Learning to Retrieve Missing Evidence for Long-Term Memory QA
    Yi-Xuan Deng, Yi Zhang, Wei Liu, Chao Xue +1

    Long-term memory enables language models to use past interactions in future conversations. However, evidence needed to answer a question may be scattered across distant turns, while the question itself omits clues needed to locate it. Retrieved facts can reveal these clues, motivating retrieval decisions conditioned on evidence already found. We introduce MERA (Missing-Evidence Retrieval Augmentation), which separates globally searchable memory from a question-specific evidence state. Verified evidence guides subsequent retrieval without restricting access to the global memory. We train a ligh

    memory
  175. arxiv:2609.37441 · cs.RO
    Anisotropic Representations Improve Planning in JEPA World Models
    Mingu Kang, Yoori Oh, Sookyung Kim, Joonseok Lee

    Latent world models learn action-conditioned dynamics in representation space and often score candidate actions by Euclidean distance to a goal representation. Joint training typically regularizes the representation to prevent collapse, but the resulting representation geometry also determines how terminal errors are weighted during planning. We show that accurate prediction and noncollapsed representations do not guarantee a task-aligned latent planning cost: isotropic Gaussian regularization can induce a geometry that ranks feasible outcomes differently from the task cost. To address this mi

    world modelaction-conditioned
  176. arxiv:2609.37435 · cs.LG
    Variational Augmented Invertible Koopman Autoencoder for probabilistic time series forecasting
    Anthony Frion, Lucas Drumetz, Guillaume Tochon, Mauro Dalla Mura +2

    Neural Koopman autoencoder models have been shown to successfully build a latent embedding with linear dynamics for arbitrary dynamical systems, enabling strong performance in long-term time series forecasting. However, these models usually work in a deterministic setting, which does not allow the quantification of the uncertainty of their predictions. Thus, we propose the new Variational Augmented Invertible Koopman AutoEncoder (VAIKAE), in which the latent embedding follows a Gaussian distribution instead of being deterministic. A key property of the VAIKAE architecture is that it leverages

    benchmark
  177. arxiv:2609.37433 · cs.RO
    FP2: Equipping Robotic Foundation Models with Force Control
    Hongjie Fang, Shirun Tang, Junjian Hu, Shidong Zhang +7

    Robotic foundation models (RFMs) are increasingly capable of general-purpose manipulation, yet reliable physical interaction remains challenging in contact-rich settings. We present FP2, a lightweight downstream interface that equips task-adapted RFMs with explicit force control while preserving their action-generation capability. FP2 adopts an action-regulation decomposition: the task-adapted RFM serves as a foundation policy responsible for task-level action generation, while a high-frequency force control policy focuses solely on interaction regulation. To condition force regulation on the

    manipulation
  178. arxiv:2609.37432 · cs.LG
    Looped Actor: Depth-Recurrent Reasoning Models for Reinforcement Learning
    T. Konstantin Rusch, Tim Seyde, Jared Boyer, Zach J. Patterson +1

    Looped reasoning models repeatedly apply a shared set of parameters, enabling more computation without increasing the model size. These models also support input-dependent computation by dynamically deciding when to stop looping. Motivated by the recent success of looped transformers in language modeling and reasoning, we investigate whether dynamic looping can similarly benefit sequential decision-making. We provide a complexity-theoretic motivation for this approach by showing that there exist Markov decision processes in which a state-adaptive policy achieves the optimal return with asympto

    manipulation
  179. arxiv:2609.37426 · cs.CV
    LazySloth: Bounded LLM-based Lazy Tree Search for Fast Long Video Comprehension
    Arka Mukherjee, Kaleen Shrestha, Larissa Zhu, Maja Matarić

    Modern vision-language models (VLMs) have shown promising results in long-video understanding due to the rich semantic information they can capture. However, most methods focus on coarse captioning of extracted image frames that are computationally inefficient and require models with large context windows. While past work has explored efficient methods through multimodal retrieval-augmented generation (RAG), they rely on lossy embeddings that lose temporal context and fine-grained detail. Few works to date have investigated how VLM-based query-relevant information retrieval can be optimized. W

    retrieval-augmentedagenticbenchmark
  180. arxiv:2609.37419 · cs.RO
    Towards Spatial Perception for Heterogeneous Robot Collaboration in Subterranean Mining Environments
    Mario Alberto Valdes Saucedo, Akash Patel, Christoforos Kanellakis, George Nikolakopoulos

    The autonomous extraction of deep mineral deposits in abandoned underground mines is fundamentally a multi-agent integration problem. No single platform simultaneously offers the mobility to traverse kilometers of degraded drifts and the sensing payload required to characterize an ore body. This article presents the onboard perception pipeline that bridges two heterogeneous agents within the PERSEPHONE autonomous mining mission. Which consist of a lightweight Explorer robot that maps an unknown mine and generates a 3D scene graph of inspection targets, by running a zero-shot, vision-language s

    scene graphmulti-agent
  181. arxiv:2609.37416 · cs.LG
    Scale Sensitivity in Low-Bit Post-Training Quantization: Curvature of the Quantization Error Landscape
    Jonas von Berg, Massimiliano Datres, Carlo Kneißl, Gitta Kutyniok

    Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; the common max-based choice degrades sharply at low bit-widths. We study how sensitive this objective is to the scale. For a layer with i.i.d. Gaussian weights and calibration activations of sufficiently large effective rank, we prove that, as the width grows, the normalized round-to-nearest loss converges with high probability, uniformly over all scales, to the mean-squared error of a uniform quantizer applied to a standard Gaussian; we verify th

    post-training
  182. arxiv:2609.37407 · cs.CV
    Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation
    Xianghan Wei, Xiaoda Yang, Zhi Wang, An Pan +6

    While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target

    retrieval-augmentedagentagentic
  183. arxiv:2609.37405 · cs.AI
    Complexity-Aware Evaluation of LLM Comprehension
    Ali Mohammadi Esfahani, Nafiseh Kahani, Samuel A. Ajila

    Large language models (LLMs) are increasingly used for software engineering tasks that require understanding existing source code, including behavior prediction, function explanation, debugging, and code review. However, aggregate benchmark accuracy can conceal how model reliability changes as source code becomes structurally more complex. This paper presents a complexity-aware framework for evaluating LLM code comprehension using cyclomatic complexity, nesting depth, branching factor, and Halstead volume. We evaluate DeepSeek-Coder-V2 and Llama through two complementary tasks: automatic input

    benchmark
  184. arxiv:2609.37402 · cs.AI
    Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
    Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang +4

    Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We further observe that routing quality often saturates well before all query--model feedback is collec

    benchmark
  185. arxiv:2609.37400 · cs.CV
    BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling
    Xiaojian Shen, Dahu Shi, Jianrong Zhang, Hai Li +6

    Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structur

    benchmark
  186. arxiv:2609.37398 · cs.RO
    Direct Experience World-Model Optimization: Learning the World Beyond Action Imitation
    Xiangcheng Zhan, Zirui Chen, Yicheng Zhao, Ziteng Gao +1

    World-Action Models (WAMs) couple action generation with predictions of how physical interactions unfold. However, current post-deployment learning paradigms typically improve behavior without requiring better world predictions. Especially in dexterous manipulation, small execution errors can compound in high-dimensional action spaces, hindering policy improvement and pushing interactions beyond the world model's training distribution. Motivated by this, we propose Direct Experience World-Model Optimization (DEWO), a post-deployment learning paradigm for WAMs that, alongside action imitation,

    manipulationdexterousworld model
  187. arxiv:2609.37392 · cs.LG
    Learning Macroscopic Dynamics without Reconstructing Microscopic States
    Zhichao Han, Yue Zhao, Qianxiao Li

    Modeling the temporal evolution of macroscopic properties of complex systems is an important scientific task. To predict this evolution without full microscopic simulation, a common approach encodes microstates into compact latent states, learns their evolution, and reads out macroscopic predictions from the latent trajectory. These latent states are often learned through microstate reconstruction. However, with limited latent capacity, reconstruction can favor high-variance microscopic details over information needed for macroscopic prediction. Yet jointly learning latent states and their tra

    latent dynamics
  188. arxiv:2609.37391 · cs.LG
    Rethinking Soft Tokens for Parallel Decoding in Diffusion Language Models
    Kodai Kawamura, Kenji Kawaguchi, Anji Liu

    Diffusion language models (DLMs) enable parallel generation by predicting and committing multiple tokens at each denoising step, yet they can generate individually plausible but mutually inconsistent tokens. Recent work shows that \emph{soft tokens} can mitigate this issue by representing uncertain positions with continuous embeddings built from the model's predictive distribution at the previous decoding step. However, although soft tokens are commonly understood as preserving predictive uncertainty, how soft-token feedback improves parallel decoding has not been systematically examined. In t

    benchmark
  189. arxiv:2609.37386 · cs.CV
    Anatomy-Aware Prediction of Bronchoscopic Accessibility from 3D CT
    Linkai Peng, Cuiling Sun, Bin Wang, Jamie Rowell +9

    Pre-operative planning for bronchoscopy is critical for the diagnosis of lung lesions. Current accessibility assessment relies on subjective manual inspection of CT scans, which is time-consuming and prone to inter-observer variability. In this paper, we formalize bronchoscopy accessibility prediction as a novel supervised learning task and present the first end-to-end framework to address it. We propose an Anatomy-Aware Mixture-of-Experts (MoE) model that integrates specialized modules: a CT Expert for local morphological features, a Lobe Expert for anatomical priors, and a Path Geometry Expe

    benchmark
  190. arxiv:2609.37384 · cs.LG
    MoTIF-X: A Multimodal Tokenized Framework for Interpretable and Extensible Molecular Representation Learning
    Linqing Mo, Jiayu Zhou, Bin Chen

    Molecular representation learning is central to computer-aided drug discovery. Molecular graphs, SMILES strings, and 3D conformations provide complementary structural information, yet many multimodal approaches encode these views independently and align them only at a later stage, limiting fine-grained cross-modal interaction and substructure-level interpretability. To address these limitations, we introduce MoTIF-X, a motif-centered framework that uses graph-grounded chemical motifs as shared anchors for multimodal integration and interpretation. Its first pretraining stage learns motif repre

    benchmark
  191. arxiv:2609.37378 · cs.LG
    Do-JEPA: From Masking to Intervention in Latent World Models
    Hossein Resani, Javen Qinfeng Shi

    Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action $a$ and under a reference action $a_{\varnothing}$, and train the model to predict the difference $Δz=z^{a}-z^{a_{\varnothing}}$ between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action ente

    world modelbenchmark
  192. arxiv:2609.37374 · cs.CV
    MG-Thinker: Bi-Axial Self-Reflection for Multi-Image Reasoning Grounding
    Heyu Huang, Chi Chen, Zonghao Guo, Yuhua Li +2

    Reinforcement learning (RL) has recently delivered substantial gains in multimodal reasoning, opening a promising route for fine-grained visual perception. Yet for multi-image reasoning grounding (MRG), reasoning over real-world multi-image contexts toward pixel-precise localization, existing RL-based approaches overlook two characteristics intrinsic to this paradigm: a coarse-to-fine hierarchical reasoning pattern, and heterogeneously distributed task--sample difficulties. In this work, we present MG-Thinker, a post-training RL framework that advances a new MRG paradigm featuring such hierarc

    post-trainingbenchmark
  193. arxiv:2609.37372 · cs.CV
    Think Before You Score: Thinking Reward Model for Visual Generation
    Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang +8

    Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that explicitly determines what matters for each case before judging how well the candidate performs. Following this principle, we propose the Thinking Reward Model (TRM), which formulates case-adaptive rubrics, performs rubric-guided assessment, and produces fine-grained pointwise rewards

    benchmark
  194. arxiv:2609.37361 · cs.CL
    SemOPT: Fixing Semantic Errors in LLM-based Optimization Modeling via Reward-Guided Search
    Zetong Zhou, Wentao Zhang, Jingyuan Wang, Yifan Yang +2

    Operations research supports decision-making in domains such as energy, economics, and healthcare. Solving operations research problems typically begins with optimization modeling, which translates a natural-language problem description into executable solver code. LLMs offer a promising way to automate this process, but they remain prone to errors. In practice, these errors can be divided into two categories: syntactic errors refer to solver code that fails to run successfully or is judged infeasible by the solver; semantic errors refer to solver code that successfully returns an objective va

    benchmark
  195. arxiv:2609.37359 · cs.RO
    Encore: Few-Shot Agentic Discovery of Manipulation Strategies
    Yifan Kang, Zihan Wang, Zhiwen Fan, Bangya Liu

    Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the

    manipulationliberogrippergraspagentagentic
  196. arxiv:2609.37356 · cs.AI
    Teaching LLMs to Generate Challenging MILP Instances via Solver Feedback
    Jitin Singla, Parikshit Pareek, Pratik Jawanpuria, Parag Singla

    Generating optimization instances that are both feasible and computationally challenging is crucial for benchmarking solvers and training learning-based optimization algorithms. Existing non-LLM generators rely on seed instances or parameter tuning, resulting in high test-time computational cost, while existing LLM generators lack explicit hardness measures. Recent reinforcement learning methods with verifier feedback evaluate only binary correctness, which is misaligned with generating challenging problems. We note that an optimization solver reports the cost of solving at several stages of i

    self-playbenchmark
  197. arxiv:2609.37353 · cs.AI
    Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation
    Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang +7

    Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insuff

    agentbenchmark
  198. arxiv:2609.37351 · cs.LG
    Port-Hamiltonian Latent Deliberation: Mitigating the Deliberation Drift Cliff in Test-Time Compute Scaling
    Zeyu Jia

    Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains at short horizons (K <= 4), their reasoning collapses when extrapolated to deeper thinking steps (K >= 16), dropping by 22% to 62% across standard logical benchmarks. We resolve the trilemma among expressivity, Lyapunov stability, and computational efficiency in test-time latent

    benchmark
  199. arxiv:2609.37349 · cs.CV
    TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG
    Yalun Wu, Bingzhou Wang, Boyang Wang, Peiying Wang +4

    Multi-step visual retrieval-augmented generation (RAG) answers complex questions by repeatedly retrieving visual evidence, updating an intermediate state, and deciding whether to continue searching or answer. Yet retrieving relevant evidence does not ensure its effective use throughout the reasoning trajectory. As multi-step reasoning progresses, redundant sources occupy context capacity needed for missing evidence, observations tied to resolved requirements or unproductive searches linger in context, and visual sources are revisited with insufficient detail for fine-grained reading. We term t

    memoryretrieval-augmentedragevaluation protocol
  200. arxiv:2609.37348 · cs.RO
    DROM: A Language-Guided Diffusion Framework for Multi-Skill Robotic Manipulation
    Vincenzo Pomponi, Rocco Felici, Paolo Franceschi, Stefano Baraldo +4

    Learning robust manipulation policies for diverse, long-horizon tasks from limited demonstrations remains a fundamental challenge in robotics. We present DROM, a language-guided diffusion framework that enables robots to learn, represent, and compose multiple manipulation skills within a single generative policy. DROM leverages Dynamic Movement Primitives (DMPs) to augment a small set of expert demonstrations into expressive multi-skill datasets, substantially reducing data collection while improving spatial generalization beyond the demonstrated workspace. Building upon Motion Planning Diffus

    manipulationfranka
  201. arxiv:2609.37339 · cs.CV
    UGO: Unified Architecture for General Multi-Object Tracking by Segmentation
    Jer Pelhan, Alan Lukezic, Matej Kristan

    General multi-object tracking (GMOT) tracks all instances of a user-specified category from a single first-frame exemplar. Prior work relies on bounding boxes and surrogate training, and struggles with non-rigid objects, crowded scenes, and distractors. We introduce UGO, a unified GMOT tracker that pairs a pretrained exemplar-conditioned detection head with an instance-propagation head in a common architecture. A novel training-free, energy-minimization consolidation method converts overlapping proposals into exclusive pixel-wise masks and detections, resolving over-segmentation, duplicates, a

    memorybenchmark
  202. arxiv:2609.37334 · cs.RO
    Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
    Sohyun Lee, Yoonjae Baek, Jaesang Won, Jinnyeong Kim +4

    Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the action commanded by a VLA and the motion executed by the robot. To stress-test VLA robustness across e

    vision-language-actionvlavla policybenchmark
  203. arxiv:2609.37330 · cs.LG
    The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators
    Thomas Deixelberger, Markus Steinberger

    Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a fea

    sim-to-real
  204. arxiv:2609.37324 · cs.AI
    VISTA: Value-Informed Event Appraisal for Multimodal Emotion Conflict
    Jiale Dai, Liuxian Ma, Xiaoke Niu, Wenjing Zhang +4

    Conflicting emotional cues can be individually valid: a subdued voice may reflect a blocked goal while a smile satisfies a social obligation. Their interpretation depends on what the event means to the person. We introduce VISTA (Value-Informed Semantic Trust Arbitration), a learned seven-field appraisal interface that conditions modality arbitration on concerns, event relations, and expression conditions while retaining a joint-evidence residual. A log-odds decomposition separates emotion expectation from cue diagnosticity, motivating an interface that lets appraisal change how evidence is in

    benchmark
  205. arxiv:2609.37321 · cs.AI
    PowerMarketJax: A JAX Benchmark Suite for Multi-Agent Reinforcement Learning in Power Markets
    Zhanhua Pan, Xin Qin, Xiao Liu, Zhilong Cao +2

    Power markets are a natural testbed for multi-agent reinforcement learning (MARL), where multiple self-interested participants repeatedly submit bids. A market-clearing mechanism then determines dispatch and prices subject to power grid constraints and market settlement rules. However, existing MARL environments typically focus on a single market setting, implement simplified clearing mechanisms, or rely on CPU-based optimization solvers that slow large-scale training and limit the systematic study of bidding strategies and market behavior. We introduce PowerMarketJax, a benchmark suite for MA

    multi-agentbenchmark
  206. arxiv:2609.37317 · cs.CV
    What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation
    Sieun Hyeon, Yejoon Lee, Mintaek Lim, Woojin Kim +2

    Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's

    benchmark
  207. arxiv:2609.37315 · cs.LG
    Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments
    Rohith Reddy Bellibatlu, Zichong Wang, Wenbin Zhang

    Tool-using agents are entering settings where a wrong action carries real cost, and the benchmarks certifying them grade what each simulated tool call reports having done, assuming the tool did what its interface advertises. The audit taxonomies we survey publish no category for that assumption, and a defect beneath a score is present on every rerun. We treat a tool's advertised surfaces as an executable contract, check the implementation against it, and trace each score's provenance through the task files and evaluator code to the verdicts that derive from state a defective tool should have w

    agentagent benchmarkbenchmarkevaluator
  208. arxiv:2609.37311 · cs.AI
    ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents
    Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin +2

    Recent Recommendation Agents (RecAgents) offer a promising alternative by shifting recommendation to an active, user-side paradigm, where generative agents autonomously perceive external platforms, reason over user preferences, and execute decisions. However, existing RecAgents still suffer from two critical limitations: brittle item perception based on noisy and heterogeneous item pages, and inefficient long-context reasoning over extended user histories and multi-step interaction traces. To address these challenges, we propose a novel recommendation agent framework, termed as ReMem, that com

    humanoidmemorymemory modulelong-contextlong contextexternal memory
  209. arxiv:2609.37307 · cs.RO
    Remember What You Did: Action-History Memory with Dual-Expert Denoising for Long-Horizon Vision-Language-Action Policies
    Yaxin Zhao, Dianye Huang, Chenwei Wang, Chenguang Yang +1

    Vision-language-action (VLA) models have driven rapid progress in robotic manipulation, demonstrating strong fine-grained control and promising performance on long-horizon tasks. However, many existing VLAs lack explicit access to interaction history, making them vulnerable to perceptual aliasing: similar current observations and robot states at different task stages may induce action ambiguity and lower success rate. Existing methods incorporate temporal or progress cues through feature conditioning, action-prior modification, or sampling guidance. However, methods that jointly fine-tune memo

    vision-language-actionvlamanipulationliberomemorymemory module
  210. arxiv:2609.37304 · cs.LG
    MetaCtrl: Your Large Language Models Can Reason Better and More Concisely with a Metacognitive Controller
    Zhibin Wen, Tao Han, Lei Bai, Can Li +1

    Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not always lead to better results and can introduce substantial redundant reasoning on simple problems. Conversely, aggressively shortening reasoning can degrade performance on difficult ones. Effective reasoning therefore requires dynamically deciding when additional computation is useful based on the reasoner's capabilities and evolving solution state. Existing approaches often rely on predefined budgets or intervention rules, retrain the target

    benchmark
  211. arxiv:2609.37298 · cs.LG
    Scaling Full Conformal Image Classifiers
    Julio Silva-Rodríguez, Ender Konukoglu

    Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a l

    benchmark
  212. arxiv:2609.37294 · cs.LG
    SafeLLM4SE: Statistical Evaluation and Reporting for LLM-based Software Engineering Systems
    Francisco Ortin

    Large language models (LLMs) are increasingly used for software engineering tasks, yet their stochastic behavior challenges the validity, reproducibility, and comparability of their evaluations. Conventional practices such as reporting a single output, an average score, best-of-N, or pass@k performance can obscure variability and estimation uncertainty, potentially leading to misleading conclusions about system reliability. This article presents SafeLLM4SE, a practical methodology and reporting standard for statistically principled evaluation of LLM-based software engineering systems. Rather t

    benchmark
  213. arxiv:2609.37292 · cs.RO
    Recovering the View: Benchmarking Physical Active Vision for Occlusion Recovery in Robotic Manipulation
    Kaijun Luo, Yudi Huang, Qijun Zhong, Xinshuai Song +2

    Physical active vision allows robots to change their viewpoint when task-relevant observations become unreliable, yet existing manipulation benchmarks provide limited support for studying how policies recover from occlusion during execution. We introduce BAVO-Bench (Bimanual Active Vision under Occlusion), a bimanual active-vision benchmark that systematically controls external visibility through Clean, Stage Occlusion, and Random-time Occlusion conditions, enabling evaluation of both manipulation performance and active visual recovery. Building on this setting, we present A-FAR (Active Future

    manipulationbenchmark
  214. arxiv:2609.37287 · cs.CV
    VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics
    Bo Lv, Mao Zheng, Zheng Li, Fangxu Liu +2

    Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to comprehensively assess this capability. To systematically evaluate this capability, we introduce VISTA-Bench, covering 22 languages and 10 domains, and develop an image-specific rubric evaluation protocol. The benchmark combines sampling for language and scenario coverage with model-a

    benchmarkevaluation protocol
  215. arxiv:2609.37283 · cs.CV
    SAM Meets VLM: Parameter-Decoupled Full-Parameter Training for Unified Medical Reasoning and Segmentation
    Xuyang Cao, Enyou Liu, Jun Zhao, Zhuoyun Liu +2

    Medical multimodal large language models (MLLMs) are increasingly expected not only to answer clinical questions, but also to localize the visual evidence behind their predictions. A common strategy connects a vision--language model (VLM) with SAM-style segmentation through a special <SEG> token, yet full-parameter training of this unified architecture is difficult because image-level reasoning and pixel-level segmentation impose different requirements on the shared representation space. To address this issue, we propose a parameter-decoupled training framework for unified medical reasoning an

    benchmark
  216. arxiv:2609.37279 · cs.AI
    Transolver-$σ$: Joint Spectral-Physical Subspace Modeling for Neural PDE Solving
    Haonan Shangguan, Hang Zhou, Haixu Wu, Yuezhou Ma +2

    Neural solvers offer efficient surrogates for numerical simulation of partial differential equations (PDEs). For time-dependent problems, strong one-step accuracy does not necessarily translate into reliable autoregressive rollout. We observe that a solver based only on physical-state modeling can achieve lower one-step error, whereas its spectral-only counterpart can become more accurate at later rollout steps. Motivated by this observation, we present Transolver-$σ$, a neural PDE solver based on joint spectral--physical subspace modeling. Within each block, adaptive physical-state interactio

    benchmark
  217. arxiv:2609.37267 · cs.AI
    Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
    Jio Oh, Seunghyun Do, Young-Jun Lee, Steven Euijong Whang +1

    Proactive LLM agents can turn idle compute into useful support before users ask. Yet even correct work can misread user context, impose review costs, or undermine trust. This work proposes foundations for designing, realizing, and evaluating proactive LLM agents around three joint principles (3T): Task Capability, anticipating relevant needs and correctly performing useful work; Temporal Allocation, allocating compute according to resource availability and when results are needed; and Trust, sustaining users' confidence and appropriate reliance on the agent. We connect these objectives to a de

    agentllm agent
  218. arxiv:2609.37264 · cs.CV
    UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception
    Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan +8

    Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separate problems, with different task definitions, supervision formats, datasets, and evaluation protocols. This fragmentation limits the learning of transferable object-affordance semantics across visual and geometric spaces. We propose Token Router for Tasks, a multitask training paradigm for MLLM-based systems that routes contextual hidden states to task-specific branches without requiring the language head to generate predefined markers. Routed stat

    embodiedbenchmarkevaluation protocol
  219. arxiv:2609.37250 · cs.RO
    V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
    Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu +3

    World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the la

    vision-language-actionliberov-jepavjepa
  220. arxiv:2609.37239 · cs.LG
    Differentiating Bisimulation Metrics: A Framework for Parametric Markov Chain Fitting via Bicausal Optimal Transport
    Sergio Calo, Amy Zhang, Javier Segovia-Aguas, Anders Jonsson

    Many problems in sequential decision-making, such as imitation learning from observations, state-space compression, world-model learning, and sim-to-real transfer, can be reduced to learning a model such that a notion of distance with respect to the target process is minimized. We consider this general framework and consider the bisimulation metric, equivalently Bicausal Optimal Transport (BOT), as the notion of distance to minimize. We show that BOT, since it can be formulated as a linear program (LP), is differentiable with respect to the model dynamics. We then derive an exact closed-form g

    sim-to-real
  221. arxiv:2609.37236 · cs.AI
    Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
    Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi +1

    An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on wh

    agentbenchmark
  222. arxiv:2609.37233 · cs.AI
    DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis
    Yuan Li, Hanyun Jiang, Guowei Tian, Chengpeng Wang +1

    Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated

    benchmark
  223. arxiv:2609.37226 · cs.LG
    Follow the Entities: A Corpus Map for Agentic Search
    Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam

    Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often m

    agentllm agentagenticbenchmark
  224. arxiv:2609.37223 · cs.CL
    CredWise: A Controlled Agentic Decision-Intelligence Framework for Explainable and Auditable Credit-Risk Assessment
    Aakash Kumar Tiwari

    Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with other evidence. This paper presents CredWise, a decision-support framework that integrates credit-risk prediction, probability calibration, explainable artificial intelligence, policy retrieval, SQL analytics, and controlled agent-based workflows. An XGBoost model is trained on Lending Club data (1,345,310 loans, 18 features) using a temporal split: 2007--2016 for training, 2017 for validation, and 2018 for testing. On the 2018 test set, the calibr

    agentagenticbenchmark
  225. arxiv:2609.37221 · cs.AI
    OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization
    Chenxing Wei, Sichen Liu, Lizhao Liu, Ningyuan Sun +5

    Large language models (LLMs) are increasingly deployed to solve complex scientific and practical problems via iterative optimization. However, dynamically coordinating diverse search mechanisms as candidate quality, failure modes, and resource budgets evolve remains a critical open challenge. Targeted empirical diagnostics reveal that mechanism effectiveness is highly state-dependent. Motivated by this, we analyze how individual decisions drive final outcomes, decomposing the expected terminal improvement under a shared budget into cumulative decision opportunities minus cumulative selection l

    benchmark
  226. arxiv:2609.37216 · cs.AI
    CRJudgeBench: Can AI Detect Plausible but Invalid Code Reviews?
    Yue Pan, Jiawei Li, Ziyuan Zhang, Xiangxin Zhao +1

    Large language models can generate plausible code-review comments, but such comments may contain technically incorrect claims that mislead developers. We study technical trustworthiness judgment: determining whether a review comment's core technical claims are correct and applicable to the reviewed code in its repository context. Existing code-review benchmarks primarily evaluate review generation, issue discovery, or general comment quality, but do not directly assess whether an agent can determine the technical trustworthy of an individual review comment. To fill this gap, we introduce CRJud

    agentagenticbenchmark
  227. arxiv:2609.37203 · cs.AI
    Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning
    Qili Zhang, Qianren Mao, Hanze Cai, Kaiming Zhao +10

    Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to co

    benchmark
  228. arxiv:2609.37198 · cs.AI
    V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization
    Haoran He, Runyuan Cai, Yiming Wang, Lin Yu +1

    Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose

    memoryexternal memory
  229. arxiv:2609.37196 · cs.AI
    ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
    Yanjie Li, Xiangyu He, Xuelong Dai, Bin Xiao

    Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack success rate because they examine content or aggregated outputs rather than authorizing effects, especially for the within-tool attack, which preserves the intended tool but manipulates its arguments. Data-Flow Control such as CaMeL provides stronger guarantees, but incurs substantial time latency that limits practical dep

    llm agent
  230. arxiv:2609.37190 · cs.CV
    HaPRL: Human-Anchored Process Reinforcement Learning for Visual Search Agent
    Zhangquan Chen, Yaoxin Niu, Xiang An, Mingze Sun +4

    Multi-turn visual search agents answer questions about high-resolution images by iteratively deciding where to look. Reinforcement learning for these agents rewards only the final answer, leaving the search process unsupervised. Consequently, faulty routes in which the reasoning process is erroneous yet the final result is correct arise frequently, which in turn leads to ineffective training, i.e., scaling along the wrong paths. In this paper, we introduce HaPRL, the first framework to reinforce the search process with human search behavior. We first build an annotation platform and collect 1K

    agent
  231. arxiv:2609.37189 · cs.LG
    Corruption-Robust Sparse Linear Contextual Bandits with Knapsack Constraints
    Yige Wang, Hanyang Li, Yiming Zong, Wanteng Ma +1

    We study sparse linear contextual bandits with knapsack constraints under joint reward and consumption corruption. Consumption corruption creates a challenge beyond corrupted rewards: it affects not only statistical estimates, but also the recorded budget, resource prices, and stopping decisions that govern future allocation. We develop Robust Optimistic Primal--Dual (ROPD), an estimator-modular framework that combines corruption-aware confidence widths with online resource prices and a budget-safety rule. With concrete sparse implementation, ROPD achieves regret against a clean population-LP

    benchmark
  232. arxiv:2609.37187 · cs.RO
    InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning
    Hongpei Zheng, Hujun Yin

    Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned R

    embodiedaction-conditionedmemory
  233. arxiv:2609.37181 · cs.RO
    EgoHumanoid-V2: Human-to-Humanoid Transfer of Coordinated Whole-Body Skills for Loco-Manipulation
    Jin Chen, Yiming Jiang, Chongyang Xu, Modi Shi +7

    Human demonstrations capture diverse scenes and rich whole-body skills without requiring robot teleoperation. Prior work on egocentric transfer has emphasized scene generalization in loco-manipulation under decoupled control, leaving direct transfer of coordinated whole-body skills less explored. We present EgoHumanoid-V2, the first egocentric human-to-humanoid skill transfer framework for coordinated whole-body loco-manipulation. At its core, coarse-to-fine action alignment combines kinematic reference correction with dynamics-aware refinement. It improves end-effector pose accuracy while pre

    vision-language-actionmanipulationhumanoidteleoperation
  234. arxiv:2609.37171 · cs.CL
    Bridging Semantic Gaps in RAG through Generated Context Knowledge Fusion
    Xinkai Du, Chao Lv, Yalin Sun, Quanjie Han +2

    Retrieval-Augmented Generation has established itself as a fundamental framework in natural language processing, seamlessly integrating information retrieval with the generative capabilities of large language models. However, this process is fundamentally constrained by a critical challenge: semantic space mismatch between queries and retrieved contexts. We propose Knowledge-Aware Semantic Bridging (KASB), a novel framework that improves passage selection quality through semantic space alignment between queries and retrieved documents through intelligent knowledge fusion. Our approach leverage

    retrieval-augmentedrag
  235. arxiv:2609.37170 · cs.LG
    Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
    Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang +4

    Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token dis

    benchmark
  236. arxiv:2609.37169 · cs.LG
    Trajectory Soup: Pushing the Compute-Scaling Frontier of LLM Mid-training via Diverse Trajectories
    Zhehao Huang, Changxin Tian, Qingyuan Yang, Kunlong Chen +4

    Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since additional serial compute yields little further downstream improvement and can even degrade some capabilities, which places a practical ceiling on how much compute mid-training absorbs. We revisit how this compute should be allocated to a single run or multiple similar optimizations. We find that branches forked from a shared checkpoint under various controlled recipe reaches measurably different regions of parameter space, and establish a form of co

    post-training
  237. arxiv:2609.37165 · cs.RO
    Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead
    Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee +6

    Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure fro

    vision-language-actionvlavla policymanipulationlibero
  238. arxiv:2609.37162 · cs.CV
    End-to-End Self-Supervised RGB-T Tracking without Modality Misleading
    Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia +3

    RGB-T object tracking leverages the complementary characteristics of visible and thermal infrared modalities to improve robustness under adverse conditions. Existing supervised methods typically rely on costly modality-aligned bounding box annotations, while most self-supervised approaches follow a two-stage pseudo-labeling paradigm, making tracker training sensitive to pseudo-label quality and preventing joint end-to-end optimization. In this paper, we propose ESMTrack, a fully end-to-end self-supervised RGB-T tracking framework without offline pseudo-label generation or dense frame-level bou

    benchmark
  239. arxiv:2609.37161 · cs.LG
    The Vote Hides the Failure: Aggregation Choice and Noise Robustness in Heart Murmur Detection
    Nicholaus Dismas Ladislaus, Olatunji Damilare Emmanuel, Samuel Chol Buol

    Noise robustness in automated phonocardiogram (PCG) murmur detection, and how it is measured, remains underexamined despite growing interest in low-resource screening. We evaluate two independently reimplemented pipelines, Hierarchical Multi-Scale Convolutional Network (HMS-Net)--CNN, and Bidirectional Long Short-Term Memory (BiLSTM)--LSTM, under controlled, multi-severity noise with noise-augmented fine-tuning and held-out generalization testing. Under matched aggregation, the complete BiLSTM pipeline outperforms the complete HMS-Net pipeline across all conditions in accuracy and Weighted Acc

    memory
  240. arxiv:2609.37156 · cs.LG
    Lucid Dreaming for World Models: Learning to Doubt Imagination and Decide by Trust
    Ziqi Wen, Ting Xu, Lianyu Wang, Xian Lin +4

    World models enable agents to learn and plan in imagination, but predictions beyond their experience can become unreliable and mislead decisions. Existing uncertainty estimates derived from predictions can remain overconfident on unfamiliar state-action pairs. We propose the Lucid World Model (LucidWM), which learns doubt from experience and propagates trust through imagination. By integrating Subjective Logic into categorical latent transitions, LucidWM distinguishes predicted outcomes from their evidential support and assigns each transition a degree of doubt. The complement of this doubt de

    world model
  241. arxiv:2609.37153 · cs.AI
    When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents
    Zifu Tao, Changqing Yin

    Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately adopted. We introduce ToxicBench to measure checking and adoption under numerical, label, schema, and retrieval errors, pairing clean and poisoned observations over fixed source data. In the 118-task GPT evaluation across three adapters, poisoning lowers task success by 26 to 39 pe

    agent
  242. arxiv:2609.37151 · cs.MA
    FlowMAS: Learning Multi-Agent Workflow Topology via Information-guided Generative Flow Network
    Haitao Wang, Chenjing Liang, Haipeng Zhang, Jiawei Hu +5

    Automated multi-agent systems offer clear advantages over manually designed ones in scalability and adaptability, but existing workflow topology methods still face important limitations. Search-based methods are often computationally expensive, textual-gradient-based methods rely on coarse-grained feedback, and existing generation-based methods are not well suited to discrete workflow topologies with complex dependencies. To address these limitations, we propose FlowMAS, a multi-agent workflow topology method based on Generative Flow Networks (GFlowNets). FlowMAS models workflow generation as

    multi-agentagent systembenchmark
  243. arxiv:2609.37150 · cs.RO
    CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera Shifts
    Tianhang Pan, Xuanhao Wang, Yiwen Pang, Bo Zhou +3

    VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerability that limits reliable deployment. To address this vulnerability, existing methods collect paired observations of the same scene from different viewpoints to fine-tune the VLA or train visual adaptation modules. Unfortunately, they require additional data collection and VLA t

    vlaembodiedpi0libero
  244. arxiv:2609.37143 · cs.CL
    LoLBench: Evaluating Coding Agents with Long-Horizon Proposals on Large Software Systems
    Yun Peng, Zihan Wu, Zeyang Zhuang, Xin Zhou +5

    Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents' implementation capability to produce correct code edits from detailed specifications. However, practical modular development tasks also require the perception capability of grounding user intent and high-level design to derive a specification. We introduce LoLBench to evaluate both capabilities through the entire proposal-to-implementation process on large software systems

    agentbenchmark
  245. arxiv:2609.37135 · cs.CV
    Multi-Granularity Language-Guided Imitation Learning via Instruction Decomposition
    Yi-Pei Chiu, Wei-Ta Chu

    Using language instructions as conditions to guide robot policy learning has recently become an important research domain. However, existing language-guided policy learning methods typically use an overall task description to guide the entire demonstration trajectory. For manipulation tasks involving multiple execution stages, these methods assign the same language description to different subtasks, making it difficult to distinguish the behaviors required at different stages. In this work, we propose a multi-granularity language guidance method based on instruction decomposition. The proposed

    manipulationrobot policy
  246. arxiv:2609.37132 · cs.AI
    Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models
    Zheng Zhang, Xinyue Tan, Lufei Li, Xinyi Zhang +2

    On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the quality of supervision constrained by the teacher's ability to exploit privileged information. We ask whether the model's own optimization progress can instead be recycled into a stronger self-teacher. In this paper, we introduce Bootstrapped On-Policy Self-Distillation (B-OPSD)

    self-improving
  247. arxiv:2609.37131 · cs.RO
    ReF-HIL: Shaping the Critic around Human Action Neighborhoods for Efficient Human-in-the-Loop Reinforcement Learning
    Shaoyin Luo, Song Wang, Shibo Xia, Tianle Zhang +4

    Human-in-the-loop reinforcement learning (HIL-RL) offers a promising route to efficient training of robotic manipulation policies by combining autonomous learning with human demonstrations and online corrections. However, insufficient use of successful human experience in value learning prolongs costly real-world training, while persistent imitation penalties can limit value-driven policy improvement. To address these limitations, we propose ReF-HIL, an efficient HIL-RL framework that uses human guidance to accelerate the learning process. Human-Reference-Guided Value Shaping learns an indepen

    manipulationhuman-in-the-loop
  248. arxiv:2609.37128 · cs.AI
    SkillCome: Group Contrast Skill Optimization with Dual Memory
    Haolin Li, Feng Hong, Ang Li, Chilin Fu +5

    Skill evolution improves the capabilities of large language models by analyzing trajectories generated under a given skill and modifying the skill accordingly. Existing approaches typically generate a single trajectory per question. However, this provides insufficient optimization signals since it requires inferring effective skill edits from a solitary path. It is difficult to pinpoint which actions caused the failure in a failed trajectory, or to determine which actions in a successful one should be incorporated into the skill. Furthermore, they rely on a local batch of trajectories for anal

    memoryagenticbenchmark
  249. arxiv:2609.37127 · cs.CL
    LLM unbranding: Erasing Commercial Identity while Preserving Generic Utility
    Kajetan Ożóg, Alicja Wojciechowska, Dawid Malarz, Paweł Batorski +2

    Establishing unbranding as a critical practice to prevent visual logos from acquiring negative connotations is standard in image generation. Large Language Models (LLMs) now face a parallel and emerging challenge. These models frequently generate brand descriptions within diverse contexts. This frequency introduces significant risks, such as trademark dilution, false attribution, and brand defamation. In response, we formally define the novel task of LLM Unbranding. We specifically address the complex challenge of managing trade dress within textual outputs. This involves neutralizing characte

    iterative refinementbenchmark
  250. arxiv:2609.37125 · cs.AI
    When Should Agents Check External State? Budgeting Observations for Stored Intentions
    Zhengkun Di, Bin Shi, Kai Sun, Yiming Xu +1

    Prospective memory allows an agent to retain an intention tied to a future condition, but the stored intention does not reveal whether that condition currently holds. Checking it may require web access, multi-step tool use, and paid calls. Existing systems decide when intentions require attention, but do not allocate the resulting observations under a shared budget. We introduce the first resource-allocation formulation for the external observations required by stored intentions under a shared episode budget. BudgetPM offers two policy variants that share a hard-budget executor. BudgetPM-Stati

    memoryagentagent systemtool usebenchmark
  251. arxiv:2609.37119 · cs.LG
    Unlocking the Critic: Reward-Free Policy Optimization for LLM Post-Training
    Hongyang Li, Xiao Li, Caesar Wu, Said Mammar +2

    Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning. First, we find that instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact: keeping policy updates small an

    memorypost-training
  252. arxiv:2609.37114 · cs.LG
    Interpretable intrinsic dimension estimation through componentwise calibration of distance and angle
    Chih-Hsuan Huang, Chih-Wei Chen, Szu-Chi Chung

    DANCo (Dimensionality from Angle and Norm Concentration) jointly calibrates nearest-neighbor distance and angular statistics and consistently reaches state-of-the-art accuracy on clean intrinsic-dimension (ID) benchmarks. Practical data, however, introduce neighborhood-relative noise and sample-amplitude heterogeneity that can distort these geometric signals. We reformulate DANCo componentwise, retaining separate distance and angular discrepancy curves so that the source of an estimate can be identified and interpreted. For the distance component, we derive a closed-form Kullback-Leibler diver

    benchmark
  253. arxiv:2609.37111 · cs.AI
    Learning from Viable Failure Prefixes: Milestone Viability Potential Policy Optimization for Long-Horizon LLM Agents
    Qi Zhou, Yuanfan Li

    Long-horizon LLM agents require reinforcement learning methods that can assign credit to intermediate decisions under sparse and delayed rewards. Existing group-based methods such as GRPO and GiGPO alleviate this issue by comparing rollout returns or repeated anchor states, but they still fail when the compared returns have no variation. We identify this failure mode as zero-credit failure: during early training, many failed rollouts contain useful prefixes, yet existing methods assign them no task-discriminative advantage. To address this issue, we propose Milestone Viability Potential Policy

    llm agent
  254. arxiv:2609.37107 · cs.CV
    Waypoint-1.5: A Real-Time Video World Model for Consumer Hardware
    Rajit Rajpal, Shahbuland Matiana, Liew Wei Pyn, Anmol Agarwal +10

    We present Waypoint 1.5, a real-time diffusion world model for interactive video generation on consumer-grade hardware. Unlike general video diffusion models, interactive world models (iWMs) must respond to dense user controls under strict latency and throughput constraints. Waypoint 1.5 is pre-trained on 100,000 hours of diverse, control-aligned video game data across hundreds of games, and generates playable video conditioned on full keyboard and mouse input. The model includes two resolution variants that run across a wide spectrum of consumer hardware. To characterize this unique setting,

    world model
  255. arxiv:2609.37105 · cs.LG
    VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses
    Jiexing Qi, Yu He, Jun Liu, Qichen Huang +8

    Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidat

    agentagentic
  256. arxiv:2609.37104 · cs.AI
    What Does Post-Training Change in Multilingual Reasoning?
    Hongyang Li, Xiao Li, Caesar Wu, Grégoire Danoy +1

    Open-source reasoning models provide unequal access to reasoning capability across languages. When a model can solve a problem but cannot deliver a complete solution in the user's language, language becomes an access barrier rather than merely a source of performance variation. We audit Qwen3 checkpoints on competition-mathematics tasks in eleven languages. Across the ten non-English languages, only 15.4-17.9% of problems receive a correct, terminating solution with visible reasoning in the requested language in any of 16 samples, compared with 92.9% in English. To identify the source of this

    post-training
  257. arxiv:2609.37097 · cs.AI
    Breaking the Illusion of Review Reliability under Static Evaluation: SCOPE Fuzzing for LLM-based Scientific Reviewers
    Zhuo Chen, Hao Zeng, Jiawei Liu, Guoxiu He +5

    The rapid growth of submissions and reviewing workload has accelerated the use of large language models (LLMs) in peer review. Prior studies suggest that LLM-based reviewers can penalize content perturbations, such as overclaiming, indicating a certain degree of reliability. Yet these conclusions are largely based on a narrow set of perturbation strategies instantiated with static templates, providing limited evidence of actual reliability. In this paper, we construct a three-level evaluation framework covering perturbations to surface presentation, argumentative logic, and value judgment. Exp

    evaluation framework
  258. arxiv:2609.37096 · cs.CV
    Why MLLMs Struggle to Count: Overcoming Individuation and Aggregation Bottlenecks with ConvStack
    Liwei Che, Yihao Quan, Sen Fang, Hongyi Wang +3

    Multimodal Large Language Models (MLLMs) consistently struggle with fine-grained visual counting, yet the underlying causes remain poorly understood. In this work, we present a mechanistic analysis of this failure mode, identifying two critical bottlenecks inherent to the global attention pipeline of MLLMs. First, we reveal an individuation bottleneck stemming from image patchification: because Vision Transformers process patches independently, they struggle to group fragmented geometric features across boundaries into distinct object representations. Second, we identify a collapse in the subs

    benchmark
  259. arxiv:2609.37094 · cs.MA
    LLM-Based Multi-Agent Systems over Wireless Networks: A Joint Agent--Network Design Perspective
    Chao Hu, Yuan Guo, Guanlin Wu, Yueling Che +2

    As large language models (LLMs) evolve from standalone models into collaborative agents embedded in physical systems, their reasoning and execution are increasingly distributed across wireless edge nodes. In this setting, wireless networks are experiencing a paradigm shift from only providing data connectivity to supporting the multi-agent reasoning workflow itself. The task performance of such network-constrained LLM-based multi-agent systems (MASs) is jointly affected by the multi-agent reasoning dependencies as well as the underlying network connectivity and edge resources. This coupling gi

    multi-agentagent system
  260. arxiv:2609.37090 · cs.CV
    Task-Oriented Visual Feature Compression via Residual Vector Quantization for Device-Edge Multimodal Inference
    Luning Pang, Cheng Yuan, Jiawei Shao, Mingtao Huang +1

    Large multimodal models (LMMs) support diverse visual understanding and reasoning tasks but are often impractical to run entirely on resource-constrained devices. Device-edge co-inference reduces device computation, yet transmitting visual data over bandwidth-limited uplinks can introduce substantial delay. Task-oriented feature compression (TOFC) reduces the payload through feature aggregation and entropy coding. However, continuous-feature coding remains costly, and query-agnostic aggregation may discard task-relevant local evidence. We propose query-guided task-oriented feature compression

    benchmark
  261. arxiv:2609.37089 · cs.CV
    Real2Gym: Building Gyms from Videos, Bringing Skills to Robots
    Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu +3

    Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and gene

    manipulationsim2realfrankaagentagentic
  262. arxiv:2609.37085 · cs.AI
    ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum
    Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi +2

    Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. This balance becomes especially difficult when resource scaling reaches capacity limits, because changes in demand and cluster pressure must then be absorbed without violating client-defined quality ranges. Existing orchestrators mainly adapt resources, placements, or replicas, while analytics requirements such as coverage, sample, and freshness remain fixed. This article presents ARGOS, the Adaptive Reinforcement Learning-Driv

    memory
  263. arxiv:2609.37082 · cs.CL
    Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search
    Jingyuan Ma, Lynx Aster, He Zhang, Siyao Song +4

    Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavi

    memoryagent
  264. arxiv:2609.37070 · cs.RO
    Predictive Safety Curricula for Robust Legged Locomotion
    Ivan Ovinnikov, Pascal Sutter, Christian Gehring, Jordis Herrmann

    Rare but consequential failures can persist in learned locomotion policies for legged robots even when average task performance is high, in part because standard curricula primarily adapt task difficulty rather than the distribution of safety-critical experience. We introduce Predictive Safety Curricula (PSC), a framework for allocating locomotion training experience using learned predictions of future safety cost. PSC trains a distributional safety critic from policy rollouts and uses its predictions to prioritize both terrain contexts and previously encountered randomized events. The resulti

    legged locomotion
  265. arxiv:2609.37067 · cs.RO
    FACT: Fidelity-Aware Construction of Articulated Twins
    Kuixiang Shao, Chuansen Nie, Yinuo Bai, Jiayuan Gu +1

    Visually plausible articulated assets may still fail during contact interactions or exhibit inaccurate motion. We present FACT (Fidelity-Aware Construction of Articulated Twins), an agentic framework that progressively constructs articulated twins to improve geometry, contact, and dynamic fidelity. The agent drives an evidence--diagnosis--revision loop on a shared editable representation, selecting measurements and model edits using quantitative feedback, while numerical tools execute and validate the updates. It reconstructs editable articulated geometry from images through feature planning,

    agentagentic
  266. arxiv:2609.37066 · cs.LG
    Beyond Compression: Diagnosing How Post-Training Changes Mathematical Reasoning
    Hongyang Li, Yiming Zhu, Xiao Li, Caesar Wu +2

    Post-training is central to mathematical reasoning in modern large language models (LLMs), but endpoint pass@1 alone underidentifies what has changed. Gains may reflect newly reachable solutions, cheaper sampling of latent solutions, surface robustness, or memorisation. We compare three post-training paths under a common diagnostic readout: our sufficiently trained off-policy distillation trajectories, released Qwen3 off-policy-plus-on-policy distillation endpoints, and a released DeepSeek-Math endpoint trained with Group Relative Policy Optimisation (GRPO). Our probe uses cross-surface pass@K

    post-training
  267. arxiv:2609.37061 · cs.LG
    A Comprehensive View of Fairness through Distributional Stability
    Gayane Taturyan, Charlotte Laclau, Stephan Clémencon

    We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic c

    benchmark
  268. arxiv:2609.37056 · cs.AI
    Evolving Towards Better Codes: LLM-Guided Search for High-Distance Binary Linear Codes
    Amal Seddas, Vladyslav Shashkov, Maryna Viazovska, Emmanuel Abbe

    Evolutionary program search driven by large language models (LLMs) has produced record-breaking constructions for open problems in combinatorics and beyond. We apply this approach to the longstanding problem of improving the best-known bounds for binary linear codes. Building on the EvoTune evolutionary framework and the ShinkaEvolve codebase, we introduce LinCodeEvolve, which evolves code-construction programs against an exact minimum-distance evaluator. A strategy loop combines diversity-driven search and expert supervision: when progress plateaus, new strategies are used to redirect the sea

    evaluator
  269. arxiv:2609.37055 · cs.CV
    Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation
    Zhenyu Liu, Zhangquan Chen, Keyi Chen, Mingze Sun +3

    Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, suc

    embodiedself-improvingself-improvementbenchmark
  270. arxiv:2609.37054 · cs.AI
    actr: aligning thoughts and responses for multilingual safety in reasoning llms
    Xianhui Zhang, Jian Yu, Chengyu Xie, Chenhang Cui +5

    Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reas

    judge model
  271. arxiv:2609.37053 · cs.AI
    MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows
    Mei Wu, Rui Xie, Runyu Zhang, Yuqiang Li +5

    Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, a

    benchmark
  272. arxiv:2609.37052 · cs.CV
    OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models
    Yuchen Deng, Zidang Cai, Feidiao Yang, Yufei Wang +3

    Omnimodal large language models (Omni-LLMs) encode audio and visual streams into temporally interleaved token sequences for multimodal reasoning. However, processing long audio-visual token sequences incurs substantial prefill costs. Existing compression methods have made progress, but often overlook temporal changes in audio-visual semantic relevance. Motivated by temporal variation and local continuity, we propose OmniRoute, a training-free, two-stage compression framework. First, Temporal Evidence-Guided Budgeting (TEGB) derives chunk-wise modality preferences and initial leading-modality b

    benchmark
  273. arxiv:2609.37047 · cs.LG
    Multi-Depth Temporal Fusion for Feedforward, Locally Trained Spiking Neural Networks
    Aidin Attar, Eleonora Cicciarella, Michele Rossi

    We propose a new spiking neural network (SNN) design to process static images and event streams using time-to-first-spike (TTFS) latencies. Our key research question is which architectural choices best accommodate local and online learning in multi-layer convolutional SNNs. This question is addressed via an original framework combining residual-like connections with multi-depth feature aggregation and consensus. The full SNN pipeline features an early-vision front end, to convert raw visual data into sparse spike latencies, a four-layer convolutional backbone trained layerwise with unsupervise

    online learning
  274. arxiv:2609.37046 · cs.CV
    Speed in the Blind Spot: An Interpretability Analysis of Dynamic Perception in VLMs for Autonomous Driving
    Katharina Winter, Stefan Englmeier, Fabian B. Flohr

    Vision-Language Models are increasingly used in autonomous-driving systems, yet their ability to recover dynamic physical state from visual input remains insufficiently characterized. We study velocity understanding as a controlled diagnostic across three tasks: surrounding-agent speed, current ego speed, and short-horizon future ego-speed proposal. On nuScenes, we evaluate open-weight general-purpose and PhysicalAI VLMs, together with the driving-oriented Alpamayo-1.5 Vision-Language-Action model, using multiple input and output formulations. We combine verbal evaluation with temporal perturb

    vision-language-action
  275. arxiv:2609.37042 · cs.LG
    GleanVID: Complementary Token Selection for Efficient Video Large Language Models
    Shuo Yang, Changbai Li, Rui Tang, Xinyu Zhao +2

    Video Large Language Models (VideoLLMs) have achieved strong video understanding capabilities but incur substantial inference overhead due to the large number of visual tokens. Existing VideoLLM token compression methods largely rely on selection-independent scoring, overlooking cross-frame complementarity and consequently retaining redundant evidence across frames. Instead, we view video token selection as a progressive evidence accumulation process. It aims to retain visual evidence that is individually informative and collectively complementary under a limited token budget. Building on this

    benchmark
  276. arxiv:2609.37038 · cs.LG
    NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters
    Haoran Xu, Xingzhuo Guo, Yuchen Zhang, Jincheng Zhong +2

    Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instan

    benchmark
  277. arxiv:2609.37035 · cs.AI
    Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning
    Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao +7

    Streaming video assistance requires models to answer asynchronous questions from an observed prefix under a fixed context budget. Existing approaches model response timing or compress history, but an online state formed before future questions are known can omit visual details before later questions reveal their relevance; the retained state alone cannot recover them. We introduce Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning. WTI maintains compact natural-language memory entries tagged with source-video time ranges; these entries support dire

    memory
  278. arxiv:2609.37033 · cs.AI
    FedLAFP: Low-Rank Aggregation Meets Full-Rank Personalization in Federated Fine-Tuning
    Mengjun Yi, Huaian Gu, Yinghao Ai, Furao Shen +1

    Federated parameter-efficient fine-tuning enables clients to adapt pre-trained models without sharing raw data or communicating the full model, but statistical heterogeneity makes a single global adapter insufficient for personalized prediction. Existing personalized methods typically use the same low-rank structure for both shared and private adaptation, overlooking their distinct requirements for aggregation and personalization. We propose FedLAFP, a role-aware framework that couples a compact, globally aggregated LoRA branch with a client-private, full-rank-capable RandLoRA branch. The shar

    benchmark
  279. arxiv:2609.37030 · cs.CV
    MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos
    Jiahao Zhan, Yongrui Ma, Qunliang Xing, Xuanyu Zhang +5

    Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, evaluating target objects along object consistency, motion continuity, and physical plausibility. To achieve this, we first introduce VidMotion, a diagnostic dataset of 6,879 videos with designated moving objects and fine-grained annotations including dimension-wise scores and fail

    evaluator
  280. arxiv:2609.37029 · cs.LG
    LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
    Hao-Yuan He, Peng-Fei Liu, Si Shen, Ming Li

    Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding co

    long-contextpersistent state
  281. arxiv:2609.37027 · cs.AI
    Beyond Low-Rank Parameterization: Narrowing the Gap Between LoRA and Full Fine-Tuning via Gradient Decomposition
    Yihao Ouyang, Shiwei Li, Haozhao Wang, Xiandi Luo +4

    Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space directions are constrained by the current parameterization. We characterize the corresponding LoRA-accessible gradient space and show that it coincides with the tangent space induced by the current LoRA parameterization. This characterization yields an orthogonal decomposition of the full

    memory
  282. arxiv:2609.37025 · cs.AI
    AnyAct: Universal Action for Self-Evolving Agents
    Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren +1

    As large language models (LLMs) advance, AI agents are increasingly deployed in open-world environments to tackle complex sequential tasks (e.g., document processing, cross-application collaboration), relying heavily on actions ranging from GUI operations to semantic APIs. However, three core challenges persist: the "scale dilemma" of massive tool ecosystems exceeding LLM context windows, the "non-stationarity" of tool quality due to updates or outages, and the "heterogeneity" of feedback formats (pixels, text, structured data) creating information silos. To address these, we propose AnyAct, a

    ai agentself-evolvingbenchmark
  283. arxiv:2609.37022 · cs.AI
    Physics-Informed Multi-Agent Coordination for Hospital Patient Flow Optimization
    Guoqing Zhang, Rafik Hadfi, Takayuki Ito

    Efficient patient flow coordination across autonomous hospital departments is critical for mitigating overcrowding and balancing resource utilization. While classical queueing theory, specifically open Baskett--Chandy--Muntz--Palacios (BCMP) networks, provides an interpretable mathematical topology for healthcare operations, analytical models rely on stationary assumptions and fixed routing matrices that degrade under state-dependent real-world dynamics. Conversely, centralized reinforcement learning approaches struggle to accommodate the decentralized structure of hospital governance, where i

    multi-agentagent system
  284. arxiv:2609.37017 · cs.CL
    LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration
    Shinan Zhang, Tao Zhang, Qihui Zhu, Mengjie Zhang +6

    LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. However, directly forwarding all sender latents makes the receiver-side context scale with both the number of agents and the reasoning length, increasing computation, memory usage, and collaboration latency. A natural solution is latent compression. But we find that cross-agent redundancy remains unresolved in existing latent compression approaches, which typically compress each sender independently and then concatena

    memorymulti-agentagent systembenchmark
  285. arxiv:2609.37011 · cs.AI
    OPFL: Optimistic Verification of Federated Learning via Empirical Boundary
    Hongxu Su, Jianzhu Yao, Xuechao Wang, Pramod Viswanath

    Federated learning enables multiple clients to collaboratively train models without sharing their private data. However, the lack of visibility into local training makes it difficult to verify whether clients follow the prescribed training procedure or submit malicious updates, such as model poisoning. A natural approach is to replay client training for verification. However, privacy-preserving replay produces numerical results that cannot be directly matched with local client execution because the two run in different environments. We present OPFL, an optimistic verification framework for pri

    manipulation
  286. arxiv:2609.37009 · cs.MA
    An LLM-powered Agent Framework for Heterogeneous Evacuation Behavior Modeling under a Moving Threat in a Public Plaza
    Jian Ma, Runxin Yu, Tianyu Tang, Xiaolian Li

    Modeling heterogeneous evacuation behavior under a moving threat is difficult because human perception, memory, and evidence evaluation are not well captured by fixed rules. We propose a novel LLM-powered agent-based framework to represent these internal decision processes. Each pedestrian agent perceives a private symbolic ASCII view, maintains a Memory-based Knowledge Graph derived solely from individual observations, and makes decisions through persona-conditioned prompts under a common sampling configuration. A compressed decision context with stateless memory preserves trial-and-error exp

    memoryknowledge graphagentllm agentagent framework
  287. arxiv:2609.37004 · cs.CV
    World2Motion: Turning Video World Models into 3D Human Motion Generators
    Tu Fangyuan, Xiangyue Zhang, Yiyi Cai, Yichen Peng +7

    We present World2Motion, a framework that generates scene-aware 3D human motion and corresponding video from a single image and a text prompt. While existing 3D motion generators learn from motion datasets, their generalization is constrained by limited coverage of environments. In contrast, video world models such as Cosmos 3 offer broader environmental priors but are not designed for full-body motion generation; recovering motion from their generated videos requires costly two-stage inference. To address these, we turn Cosmos 3 into a single-stage 3D motion generator. This adaptation has two

    world modelbenchmark
  288. arxiv:2609.37003 · cs.CV
    VesselBench-800K: A Large-scale Perception Benchmark for Multimodal Vessel Detection, Counting, and Density Estimation
    Danfeng Hong, Chenyu Li, Jocelyn Chanussot

    Vessel perception from space is crucial for a wide range of maritime applications, from traffic monitoring to environmental protection. However, most existing datasets predominantly focus on general object detection tasks in optical remote sensing (RS) images. Relying solely on single-modality optical RS images proves inadequate for effectively perceiving vessel objects in complex maritime scenarios, where ever-changing weather conditions (e.g., clouds and rain), the need for day-and-night coverage, and the inherent limitations of a single imaging modality pose significant challenges. To fill

    benchmark
  289. arxiv:2609.37002 · cs.CV
    Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom
    Xijia Tao, Yihua Teng, Xinyu Fu, Cheng Gong +4

    High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in PSisual Parallel Search improves mean accuracy over dedicated zoom-only search in 14 of 15 same-mode

    agenttool usebenchmark
  290. arxiv:2609.37000 · cs.AI
    Cross-Organizational SysML Model Integration: A Survey of Challenges and AI-Supported Tasks
    Zirui Li, Torsten Brix, Stephan Husung

    Cross-organizational collaboration is widely regarded as a key promise of SysML-based Model-Based Systems Engineering (MBSE), yet practitioners still face persistent challenges when exchanging and integrating system models. In parallel, Large Language Models (LLMs) raise expectations for AI-assisted model understanding and integration, while reliability and required human oversight continue to pose challenges. This paper reports the results of an online questionnaire survey with 29 MBSE stakeholders involved in cross-organizational collaboration. Respondents rated eight predefined integration

    human-in-the-loop
  291. arxiv:2609.36996 · cs.LG
    ImbalancE: Inference-Time Latent Search Against Degree Imbalance in Link Prediction
    Alberto Bernardi, Luca Costabello, Christophe Gueret

    Knowledge Graph Embedding models have been extensively used to learn representations of entities and relations in Knowledge Graphs for predicting missing links. However, the quality of the learned representations varies a lot across different areas of the graph. If previous research has loosely linked the problem to relation types or degree bias, we show that it is more widespread and it correlates with the degree imbalance of the entities in test triples. In particular, the prediction of a target entity that has a degree much smaller than the degree of the anchor entity is extremely problemat

    knowledge graphbenchmark
  292. arxiv:2609.36995 · cs.CV
    Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
    Xingtong Ge, Yutong Wang, Lunjie Zhu, Haitao Lin +6

    Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Meanwhile, directly reusing bidirectional score models in causal Distribution Matching Distillation (DMD) creates a mismatch between generation and scoring contexts. We address these challenges with Salt++, a two-stage post-training framework comprising Causal Self-Flow (CSF) and c

    post-training
  293. arxiv:2609.36993 · cs.RO
    All You Need Is Low Fidelity: Zero-Shot Sim-to-Real of Learned Robotic Fish Control
    Liam Maloney, Simon Ramchandani, Mike Y. Michelis, Ronan Hinchet +1

    Complex tasks for underwater robots remain limited by the capabilities of their controllers. Learning a better one for a soft, underactuated robotic fish trades simulator cost against fidelity. We show that an intentionally low-fidelity simulator is enough: a stateless, quasi-steady fluid model with no wake and no added-mass history suffices to learn a \emph{general}, closed-loop controller that transfers to hardware without tuning. Our platform is a soft, single-motor, tendon-driven fish whose policy observes only what the hardware can measure. A staged pipeline grounds the simulator in two i

    sim-to-real
  294. arxiv:2609.36987 · cs.CL
    CypherTurn: A Multi-Turn Benchmark for Conversational Text-to-Cypher Evaluation and the Autonomy Divergence
    Yuzhe Zhang, Weijie Zhu, Haolin Yang, Ziyun Zhang +4

    Graph databases are increasingly queried through natural language, yet every existing benchmark evaluates isolated single-turn queries rather than the multi-turn sessions through which analysts actually work. We introduce CypherTurn, the first benchmark for conversational Text-to-Cypher evaluation, comprising 721 sessions and 5,927 turns across 7 knowledge graphs and 13 conversational phenomena. We evaluate 15 models under a guided oracle protocol and a fully autonomous agentic protocol, yielding four findings. First, the best model reaches only 64.7% execution accuracy, and session-level corr

    knowledge graphautonomous agentagenticbenchmarkleaderboard
  295. arxiv:2609.36985 · cs.LG
    Abductive World Modeling via Causal Representation Learning
    Ziqi Liu, Songhan Yang, Linfan Zhou, Jiatong Liu +3

    The central challenge of world modeling is to learn representations that capture how the world evolves. However, existing world models predominantly represent future states without explicitly capturing the latent causes underlying their evolution, limiting their ability to reason about why and how the world changes. To address this limitation, we propose Abductive World Modeling (AWM), a framework that learns structured causal representations by abductively inferring latent causes from predicted futures. Specifically, we realize AWM through the Hierarchical Abductive State Pyramid (HASP), whic

    world modelv-jepa
  296. arxiv:2609.36984 · cs.AI
    REALHOP: Rethinking Multi-Hop Reasoning Evaluation via Behavioral Auditing
    Jiawen Tao, Xiaokun Yuan, Yaoming Li, Chenxu Liu +3

    Complex questions often require multi-hop reasoning that connects facts distributed across sources or distant regions of a long context through intermediate steps. Benchmarks commonly evaluate this ability with questions built around predefined reasoning chains, treating a correct answer as evidence that the intended composition was used. Yet answer correctness alone leaves open whether success depends on the evidence associated with each intended step: models may instead rely on memorized associations, shorter paths, or partial evidence. We examine this dependence using the Behavioral Necessi

    long-contextlong contextbenchmark
  297. arxiv:2609.36982 · cs.AI
    SRJudge: Empowering Large Language Models with Selective Reasoning for Fine-Grained Knowledge Concept Tagging
    Zhiwei Yang, Jiahua Yang, Huiru Lin, Xing Chen +1

    Knowledge concept tagging aims to assign specific concept or topic labels to educational content, which is essential for both educators and learners in traditional and online teaching practices. Recent work has explored large language models (LLMs) for this task, achieving promising performance. However, LLMs still struggle to select the correct concept from a large-scale candidate set due to the high dimensionality of the decision space. In this paper, we propose a novel three-stage Select-Reason-Judge (SRJudge) framework, which empowers LLMs with selective reasoning capability for fine-grain

    benchmark
  298. arxiv:2609.36980 · cs.CV
    UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching
    Jiajun Le, Yifan Lu, Zizhuo Li, Lei Cao +2

    Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subse

    memory
  299. arxiv:2609.36976 · cs.CL
    AMU:Admission and Memory Update for Personalized Conversations---Structured Memory with SLM Guided Control
    Tao Hwang, Yishi Diao

    Large language models (LLMs) have become the foundation of personalized assistants, but maintaining persistent user memory across long-term interactions remains challenging. Existing memory systems often focus on storage, retrieval, or consolidation, while memory writing remains less controlled: transient requests, duplicate statements, and outdated user states may enter memory and later be retrieved for personalization. In this paper, we present AMU: Admission and Memory Update for Personalized Conversations, an SLM-guided (Small language model guided) structured framework for writing-time me

    memory
  300. arxiv:2609.36975 · cs.CV
    A Dual-Track Curation-and-Classification Framework for Resolving Ground-Truth Label Noise in Operational Sentinel-2 Wheat Area Estimation
    Kasimali Agharia, Ujjwal Kumar Gupta

    Operational estimation of wheat-cultivated area is persistently constrained by discordance between administrative record-keeping and remotely sensed classification products. We address this administrative reference discordance for the 2022 Rabi season in Patiala district, Punjab, India, using a thirteen-timestep Sentinel-2 NDVI time series. A curated 849-sample reference dataset, developed through an iterative rule-based bootstrapping procedure, underpins both a feature sensitivity analysis and an operational classifier. Feature sensitivity independently assessed via Cohen's d and gradient-boo

    benchmark
  301. arxiv:2609.36970 · cs.LG
    Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages
    Katarzyna Woźnica, Katarzyna Rogalska, Zuzanna Sieńko, Mustafa Cavus

    The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different explanations. This creates a critical vulnerability in AutoML: x-hacking, the selective post-hoc choice of a model based on its explanation rather than predictive merit. No existing AutoML framework exposes this risk. We introduce ARSA ML, an open-source Python framework that quantifies Rashomon set structure and predictive multiplicity within AutoML pipelines. Using ARSA ML, we benchmark AutoGluon and H2O across 28 binary classification dataset

    benchmark
  302. arxiv:2609.36967 · cs.RO
    Beyond Token Importance: Preserving Spatial Scaffolds for Efficient Vision-Language-Action Inference
    Jiayu Chen, Shuyong Gao, Jingkai Jia, Xiaosheng Bu +5

    Existing VLA pruning strategies primarily select individual visual tokens according to task-level semantic relevance, while overlooking the spatial information required for robotic manipulation. To examine this limitation, we construct a simple Stride baseline that uniformly samples tokens along the flattened one-dimensional visual sequence, representing a purely geometric pruning strategy. Surprisingly, Stride outperforms semantic pruning and random pruning at certain pruning ratios, but collapses when the token budget is only slightly reduced. We characterize this phenomenon through the spat

    vision-language-actionvlamanipulationlibero
  303. arxiv:2609.36956 · cs.AI
    Controlled Decoding Attacks on Black-Box LLMs
    Jesson Wang, Shawn Li, Wei Yang, Franck Dernoncourt +3

    Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak

    benchmark
  304. arxiv:2609.36950 · cs.LG
    Scalable Diffusion SBI for Compositional Inference under Simulator Misspecification
    Vincent D. Zaballa, Elliot E. Hui

    Simulation-based inference is challenging when many heterogeneous observations must be composed, hierarchical latent structure must be preserved, and the simulator is misspecified relative to observed data. We develop sampling and fine-tuning methods for diffusion-based inference in design-conditional settings, where the same simulator is queried across different experimental conditions $ξ$. We extend compositional score-based inference with a continuous-time diffusion coefficient that accounts for the number of observations, avoiding Jacobian and auxiliary-covariance corrections. We introduce

    benchmark
  305. arxiv:2609.36945 · cs.LG
    Fine-Tuning on Self-Generated and Reward-Weighted Data: Learning Dynamics, Convergence Rates, and Benefits of Off-Policyness
    Zhiwei Wang, Yanxi Chen, Yaliang Li, Bolin Ding

    We study the learning dynamics of fine-tuning a policy model on self-generated and reward-weighted data, with particular focus on a generalized version of REINFORCE -- referred to as RE(S) -- that updates the rollout distribution once every $S \ge 1$ gradient steps. Prior work in bandits and reinforcement learning has developed rich theory for policy gradient methods, and on-policy sampling (i.e., a small $S$, ideally $1$) is often viewed as crucial to their success; yet in prominent application like post-training large language models, reward-guided self-training has proved to be effective ev

    post-training
  306. arxiv:2609.36944 · cs.AI
    Dual-Channel Robust Group-Relative Policy Optimization via Advantage and Sequence-Weight Estimation
    Zhongyi Li, Wan Tian, Xiang Xu, Yutian Xiao +3

    Group-relative policy optimization relies on reward-derived advantages and sequence-level likelihood weights, both of which can be sensitive to localized outliers. Extreme rewards can collapse the contrast among clean responses after group normalization, while token-level log-ratio perturbations can alter sequence weights and clipping decisions. We introduce RoVR-GSPO, a dual-channel robust optimizer that addresses these failure modes separately. Its reward channel combines robust reference estimation with bounded residual credit, while its ratio channel uses differentiable SoftRoVR aggregatio

    long-context
  307. arxiv:2609.36942 · cs.LG
    Safe-by-Design Learning via Energy-based Neural Networks
    Simone Betteti, Morteza Lahijanian, Luca Laurenti

    Learning neural-network models of dynamical systems with safety guarantees is a fundamental requirement for their deployment in safety-critical settings. Safety is commonly established by proving the invariance of a desired subset in state-space, ensuring that every trajectory initialized in this subset remains confined to it for all time under admissible inputs. Existing frameworks, however, either rely on computationally expensive post-hoc verification or employ safety-enforcing mechanisms without formal correctness guarantees. In this paper, we introduce a novel neural architecture grounded

    benchmark
  308. arxiv:2609.36940 · cs.CV
    DispFlow-GS: Displacement Flow Supervision with Motion Disentangling for Monocular Deformable 3D Gaussian Splatting
    Thai Duy Nguyen, Haitian Zhang, Addison Lin Wang

    Accurate dynamic scene reconstruction is important for robotic perception, where temporally consistent representations of dynamic environments are essential. Deformable 3D Gaussian Splatting (3DGS) models dynamic scenes through deformation fields, and recent methods incorporate motion supervision by aligning rendered Gaussian flow with optical flow. However, we find that such Gaussian-flow-based supervision provides only limited improvements in motion modeling. We identify a fundamental limitation of this supervision paradigm, namely a domain gap between rendered Gaussian flow and optical flow

    benchmark
  309. arxiv:2609.36939 · cs.AI
    SCA: Spatial Credit Assignment for Reinforcement Learning of GUI Agents
    Shengtian Yang, Ziyu Xiong, Kaibing Yang, Guangfeng Cai +5

    GUI agents automate tasks on digital devices by grounding language instructions in visual interfaces. Existing group-relative reinforcement learning improves GUI action prediction by comparing the rewards of multiple responses sampled from the same GUI state. However, binary evaluation treats spatially different failed clicks as identical and provides no relative signal when all sampled clicks fail. To address these limitations, we propose Spatial Credit Assignment (SCA), which uses the screen coordinates of sampled clicks to refine group-relative credit. Specifically, SCA predicts each held-o

    benchmark
  310. arxiv:2609.36937 · cs.CV
    WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation
    Sangeyl Lee, Seunghyun Shin, Seungho Park, Wooseok Jeon +1

    Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skeletons or parametric body meshes and struggle to preserve identity-motion binding under inter-person occlusion. To address this limitation, we propose WeLike2Party, a multi-human animation framework built on direct in-context video conditioning without explicit pose or mesh extracti

    benchmark
  311. arxiv:2609.36935 · cs.AI
    CoEM: Empowering Long-Context Reasoning with Commit-on-Evidence Memory
    Jingguang Li, Yebo Wu, Zuyi Guo, Kailang Ma +6

    Long-context reasoning is essential for complex and long-horizon tasks, yet the performance of large language models (LLMs) degrades as context length increases. Recent approaches address this by processing input chunk by chunk while maintaining a bounded textual memory in model context. However, premature information compression can discard critical details essential for subsequent reasoning. In this paper, we introduce Commit-on-Evidence Memory (CoEM), which learns when to convert source evidence into compact memory facts. Specifically, under a fixed context-memory budget, CoEM preserves pot

    memorylong-context
  312. arxiv:2609.36934 · cs.AI
    VLALight: A Vision-Language-Action Model for Traffic Signal Control
    Pan Zhang, Siqi Lai, Kemu Dong, Hao Liu

    Traffic signal control (TSC) is essential for improving urban mobility and reducing congestion. Although roadside cameras are widely deployed at signalized intersections and provide rich visual observations of evolving traffic, existing TSC methods typically rely on manually engineered traffic states or separate perception modules, creating a gap between physical observations and control decisions. We present VLALight, the first vision-language-action (VLA) model for end-to-end traffic signal control from multi-view roadside videos. VLALight directly maps visual observations to coordinated sig

    vision-language-actionvlavla modelagentic
  313. arxiv:2609.36932 · cs.AI
    Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
    Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen +6

    Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning

    benchmark
  314. arxiv:2609.36931 · cs.CL
    Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
    Mario Sanz-Guerrero, Minh Duc Bui, Manuel Mager, Katharina von der Wense

    Reproducibility is essential for scientific research, yet prior work shows that LLM outputs vary with hardware and batching. We identify an overlooked factor: the hidden injection of the current date into system prompts, which users cannot control and which changes every day. Across 9 recent LLMs and 6 datasets spanning multiple-choice QA (MCQA), math reasoning, code generation, and machine translation, performance varies solely with the current date, with deltas of up to 6% on MCQA, 14% on math reasoning, 7% on code generation, and 2.84 BLEU on machine translation. Model rankings also shift,

    leaderboardevaluation protocol
  315. arxiv:2609.36929 · cs.CV
    SFE-VGGT: Source-Free VGGT Distillation for Event-Based Monocular Depth Estimation
    Thai Duy Nguyen, Addison Lin Wang

    Recent event-based depth estimation methods successfully transfer geometric priors from vision foundation models via cross-modal distillation. However, their reliance on synchronized RGB-event pairs or depth annotations during training severely restricts practical deployment. To overcome this bottleneck, we propose SFE-VGGT, a novel source-free framework that distills the geometric priors of VGGT to the event domain without any paired RGB observations. Our core idea is to reconstruct surrogate frames directly from the target event stream to act as a frozen geometric teacher, entirely eliminati

    event camera
  316. arxiv:2609.36928 · cs.RO
    ComManip: Overfitting Manipulation Policies to Comfortable Regions
    Yidan Lai, Xinyi Chen, Shuquan Man, Huiping Zhuang

    Training robot manipulation policies relies on costly robot demonstrations, making large-scale data collection impractical. Meanwhile, to improve policy generalization, existing approaches seek greater diversity in visual observations by varying object placements, viewpoints, and robot configurations during data collection. However, under a limited demonstration budget, this strategy forces the policy to model diverse visual observations, providing insufficient supervision to learn reliable observation-action correspondences under similar local conditions. Our study reveals that policies train

    manipulationopenvla
  317. arxiv:2609.36927 · cs.AI
    Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution
    Hyewon Suh, Thanh Minh Nguyen, Chih-Lun Lee, Darrow Hartman +4

    Many computer tasks recur: the same workflow runs many times, with new inputs and from different starting states. Current computer-use agents re-plan every step of every run, which makes them costly and unreliable on such tasks. We introduce neuro-symbolic computer use, in which a recurring workflow is executed by a learned policy rather than re-derived by an agent on each run. The policy fixes the decisions that are stable across runs (ordering, variables, loops, and branches) in executable code, and delegates observation-dependent decisions, such as grounding and state checks, to neural mode

    agentbenchmarkevaluator
  318. arxiv:2609.36924 · cs.RO
    Track-and-Complete: Learning Humanoid Skills from a Single Failed Human Video
    Sarmad Idrees, Jongeun Choi

    Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until

    humanoid
  319. arxiv:2609.36923 · cs.AI
    PrecogUI: Proactive GUI Agents via Pre-cognitive Simulation and Experience Retrieval
    Bin Kang, Jiarui Ouyang, Li Jiang, Bin Chen +1

    Existing reactive Graphical User Interface (GUI) agents often fail in long-horizon, dynamic scenarios, where unexpected disturbances trigger attention-diverting and cascading failures. To address this, we propose PrecogUI, a pre-cognitive architecture that shifts the paradigm from reactive execution to proactive decision-making. Specifically, we design a Proactive Experience Pool (PEP), which caches recurring anomaly and success patterns as "state-action-result" tuples in a dual-memory repository. Furthermore, we introduce a Proactive Simulation Executor (PSE) that learns to forecast the next

    benchmark
  320. arxiv:2609.36920 · cs.CL
    Benchmarking Automatic Speech Recognition Tools for Iberian Languages
    Fernando López, Pablo Gómez, David Solans, Paulo Villegas +1

    Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Catalan, Galician, Portuguese, Spanish), with German and Turkish as controls. Evaluation uses an 85-hour dataset covering read speech, broadcast media, and audiobooks, assessing accuracy and efficiency via word error rate (WER) and real-time factors (RTF/RTFx). Results show no single model dominates:

    world modelbenchmark
  321. arxiv:2609.36915 · cs.RO
    AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations
    Rui Huang, Yanlin Mu, Lidong Li, Yucong Wang +2

    Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between manipulation and flight, continuously changing observations, and safety-critical physical interactions. These challenges demand diverse training data and systematic policy evaluation, yet collecting demonstrations and evaluating policies directly on physical aerial platforms are c

    vision-language-actionvlamanipulationteleoperationmanipulatorgrasp
  322. arxiv:2609.36914 · cs.CL
    Can Language Models Learn to Forecast Stock Prices
    Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao +10

    Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with verifiable outcomes, not only are realized returns noisy, but even what constitutes a relevant information set for making effective predictions is not obvious a priori: the model must decide which observations to gather and then commit to a numerical judgment before the outcome is k

    tool usetool-usepost-trainingbenchmark
  323. arxiv:2609.36906 · cs.CV
    SafeVantage: Vantage-Aware Memory for Reliable Embodied Decisions
    Sean Hardesty Lewis, Zuyi Guo, Benwang Chen, Zirui Liu +2

    Reliable embodied decisions under partial observability require informative observations and sufficient supporting evidence. However, semantic scores alone do not reveal which viewpoints justify a claim or where additional evidence should be acquired. We introduce SafeVantage, a vantage-aware semantic memory and active acquisition framework that retains each claim's supporting views, camera poses, and estimated target location, keeping positive support distinct from search coverage. A learned candidate-observability model uses claim-grounded geometry to predict target visibility at reachable v

    embodiedmemorysemantic memorybenchmark
  324. arxiv:2609.36903 · cs.AI
    MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
    Ke Wang, Houxing Ren, Zimu Lu, Yunqiao Yang +3

    End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and social-robot reception require a single model to track, contextualize, and respond to multiple speakers over extended durations. Progress is constrained by both data and evaluation: open multi-party speech corpora remain small and are not designed for codec-frame-level full-duplex modeling, while existin

    long-contextbenchmark
  325. arxiv:2609.36901 · cs.AI
    Digital Twin Modeling of Quantum Dynamical Systems: Dissipative Quantum Reservoir Computing
    Abhijit Sen, Bikram Keshari Parida, Shital Chauhan, Mahima Arya +1

    Modeling the response of driven many-body quantum systems from input--output data is difficult: the dynamics are nonlinear, history dependent, and expensive to simulate as system size grows. A paradigmatic case is High-Harmonic Generation~(HHG), where a strong field drives a medium to emit radiation that is highly sensitive to the drive and encodes long-range temporal correlations. We introduce a dissipative quantum reservoir computing~(DQRC) framework that builds a digital twin of such a system, learning its input--output map directly from data while the reservoir---itself a small open quantu

    benchmark
  326. arxiv:2609.36896 · cs.AI
    HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL
    JunHyeok Oh, Zian Jang, Byung-Jun Lee

    Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its actio

    manipulationbenchmark
  327. arxiv:2609.36894 · cs.CV
    DiffReID: Discriminative Diffusion Model for Object Re-Identification
    Yingquan Wang, Pingping Zhang, Dong Wang, Huchuan Lu

    As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbf{DiffReID} for object ReID. It levera

    benchmark
  328. arxiv:2609.36893 · cs.CL
    Momentum-Coupled Rubric Adaptation for Detailed Image Captioning
    Zhenwen Ji, Lei Jin, Shanyong Wang, Jiaming Lu +5

    Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dyna

    benchmark
  329. arxiv:2609.36892 · cs.AI
    Harness Evolution as Learning: Approximation, Generalization, and Optimization Limits of Self-Improving Personal Agents
    Zeyu Gan, Zixuan Gong, Yong Liu

    As the capabilities of large language models (LLMs) continue to advance, increasing attention is turning to how to translate their abilities into useful behavior. Personal agents bring this question into everyday settings, where models are expected to serve individual users and continually adapt to their preferences. With the underlying model held fixed, such adaptation relies on harness engineering: designing and evolving the surrounding layer that manages context, memory, tools, and execution. Despite rapid progress, the factors governing effective harness evolution remain insufficiently und

    self-improvingbenchmark
  330. arxiv:2609.36891 · cs.CV
    ProGuT: Label-Efficient Panoptic Segmentation for Forest Scenes
    Pankaj Deoli, Karsten Berns

    Panoptic segmentation in forest environments is bottlenecked not by semantic quality but by instance separation; existing unsupervised panoptic approaches produce usable stuff maps but near-zero thing quality. Depth or flow-based instance discovery methods needs sensors that are not always available. We present ProGuT (Prototype Guided Training), which produces panoptic pseudo-labels without per-image training masks, needing only unlabeled images and one-time cluster-to-class mapping. ProGuT clusters CLIP patch features, then recovers trunk instances through multiscale geometric prior that fal

    benchmark
  331. arxiv:2609.36890 · cs.LG
    SINO: Scale-Invariant Neural Operator
    Kaichen Ouyang, Chenglei Yu, Chuanrui Wang, Tailin Wu

    In scientific machine learning, physical fields governed by partial differential equations exhibit low-rank structure and scale invariance. When solving equations on coarse grids, missing information leads to the closure problem: modeling unresolved physics to recover lost dynamics. Although closure terms depend on grid resolution, they represent scale-invariant physical laws. A model truly learning physics should capture these mechanisms with low-rank parameterization rather than memorizing grid-specific patterns. Inspired by this, we propose the Scale-Invariant Neural Operator (SINO), which

    benchmark
  332. arxiv:2609.36889 · cs.RO
    All Roads Lead to Rome: Flow-driven Multi-Anchor Exploration for Open-Environment Active 3D Mapping
    Yang Li, Aming Wu, Zihao Zhang, Ziju Han +2

    To advance the development of embodied intelligence, Open-Environment Active 3D Mapping has attracted increasing attention, aiming to perform a long-horizon and shortest trajectory exploration for reconstructing unseen scenarios. Since only limited information about unseen environments is available, methods built on the closed-set assumption, i.e., assuming that the test environments are similar to those seen during training, cannot generalize satisfactorily. In existing active mapping methods, long-horizon exploration is often guided by predicting a coarse long-range goal and then converting

    embodied
  333. arxiv:2609.36888 · cs.AI
    Beyond Sub-Gaussian Detector Scores: Robust Weighted Profile-Loss Change Point Detection for Human-LLM Text Segmentation
    Wan Tian, Zhongyi Li, Yawen Li, Rui Zhang +2

    Mixed human-LLM documents require locating authorship transitions from detector scores whose reliability varies across text units. Existing weighted mean contrasts are vulnerable to extreme scores, while directly replacing means with robust centers obscures how a misplaced boundary changes the population objective. We propose Robust Weighted Profile-Loss Change Point Detection (RWCP), which combines capped reliability weights, Huber profile gains, and narrowest-over-threshold search in reliability coordinates. Our key analysis expresses the population gap between a true and a displaced split a

    benchmark
  334. arxiv:2609.36887 · cs.AI
    WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
    Bo Mao, Hang He, Linting Wang, Lizhi Lin +16

    Recent efforts to scale tool-use post-training have largely centered on the synthesis of executable environments, which constitute only one component of a broader agentic interaction system comprising the environment, task, agent harness, and evaluator. Scaling environments in isolation, however, does not guarantee commensurate gains in model performance, because reliable learning signals depend on coherent interactions among all components of the agentic interaction system. To address this problem, we introduce WEFT (Whole-system Evolution For Tool-use Post-training), which couples scalable a

    agentagentictool-usepost-trainingbenchmarkevaluator
  335. arxiv:2609.36885 · cs.LG
    RNA Design via Conditioned Flow Matching and Finite-Policy Reinforcement Learning
    Zefeng Lin, Xianyong Fang, Tianfan Fu, Xiaohua Xu

    RNA design aims to identify sequences that fold into specified secondary structures. Existing methods formulate the task as target-specific search or conditional generation. However, natural RNA evolution proceeds through sequence variation and selection, with compensatory substitutions, whereas these methods do not explicitly model this process. To address this limitation, we propose a two-stage framework comprising RNA Inverse-Folding Flow (RNA-IFlow) and RNA-IFlow-RL. RNA-IFlow uses structure-conditioned Dirichlet Flow Matching to model coordinated variation across the sequence, while RNA-I

    benchmark
  336. arxiv:2609.36881 · cs.LG
    What You Observe Determines How You Identify Causal Effects: Evaluating Causal Models across Observational Views
    Heejin Jung, Gyeongdeok Seo, Hoyoon Byun, Joseph Lee +1

    Causal foundation models (CFMs) pre-trained on data generated from various structural causal models (SCMs) have been proposed for estimating causal effects from observational data. However, differences in pre-training environments and evaluation protocols make it difficult to assess how their performance depends on the information available for causal identification. To enable controlled comparisons, we introduce CausalIDView, a multi-view benchmark that holds fixed SCM realization and target estimand while varying only the observational view available to the estimator. Each observational view

    benchmarkevaluation protocol
  337. arxiv:2609.36879 · cs.AI
    SKILLLITE: Evidence-Guided Malicious Skill Auditing with Compact LLMs
    Haoran Ou, Gelei Deng, Xuanye Zhang, Wenbo Guo +2

    As LLM-based agents perform increasingly complex tasks, Agent Skills have emerged as a flexible mechanism for extending their capabilities. An Agent Skill packages task-specific instructions with executable components and auxiliary resources to provide specialized functionalities. However, the growing adoption of third-party Skills introduces a new supply-chain attack surface. Malicious Skills can embed harmful behaviors that abuse agent privileges and compromise the agent execution environment or accessible resources. Although recent LLM-based malicious Skill auditing approaches have achieved

    agentagentic
  338. arxiv:2609.36875 · cs.CV
    S4VY: Segment Anything in Feed-Forward 4D Visual Geometry
    Jingdong Zhang, Xin Li, Jan Kautz, Wenping Wang +1

    Accurate instance segmentation in dynamic scenes is important for downstream applications such as robotics and autonomous driving. Existing Segment Anything models operate primarily on 2D image or video masks and preserve identity through sequential memory, while promptable 4D instance segmentation built upon feed-forward visual geometry remains underexplored. We introduce S4VY, a Segment Anything model built on feed-forward 4D visual geometry. From a set of RGB observations, S4VY transforms shared visual-geometric features into an exhaustive set of class-agnostic 4D instance masks through a s

    agentic
  339. arxiv:2609.36872 · cs.RO
    PreferenceFlow: Test-Time Guidance of Flow-Matching Robot Policies from Human Interventions
    Yiqi Tang, Diyuan Shi, Runze Li, Donglin Wang

    Flow-matching policies can represent complex robot behaviors but remain susceptible to local errors under distribution shift at deployment. Many reinforcement learning approaches to policy improvement require reward signals that are difficult to specify or obtain in real-world manipulation. We present PreferenceFlow, a framework for improving a pretrained flow policy at test time without environment rewards or updates to the base policy. Human intervention chunks are paired with robot chunks generated from the same initial conditioning state to train a preference model. During infer- ence, we

    manipulationfranka
  340. arxiv:2609.36870 · cs.RO
    VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness
    Hang Li, Mingxin Zhang, Zihan Wu, Yang Tian +8

    Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without distorting the trajectory shape. Another key limitation is that the resulting policies often lack precise object-level 3D geometry awareness, limiting object grounding and object shape awareness critical for precise manipulation. To bridge these gaps, we propose VidAct, an efficient

    manipulationsim-to-real
  341. arxiv:2609.36867 · cs.AI
    State Trace Rationale As Auxiliary Task in Reinforcement Learning
    Muhammad U. Nasir, Alex Vogt, Steven D. James, Julian Togelius

    We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations

    agent
  342. arxiv:2609.36864 · cs.LG
    Where the Model Changes Its Mind: Hindsight-Divergence Localization for Efficient Reinforcement Learning with Verifiable Rewards
    Fanchao Chen, Hengyu Fu, Shivaram Venkataraman, Jiantao Jiao

    Group-relative methods for reinforcement learning with verifiable rewards (RLVR) learn from differences in rollout outcomes. Independently sampling complete trajectories is costly and does not explicitly explore the decision space at critical positions. Feedback on a completed trajectory can reveal which earlier choices the policy reconsiders, suggesting where to sample alternative continuations. We introduce Hindsight-Divergence Localization (HDL), which uses hindsight-induced changes in token log-likelihoods to select branch points. HDL generates a small number of complete root trajectories

    agent
  343. arxiv:2609.36860 · cs.AI
    IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence
    Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang +13

    We present IronLLM-0.6B, a 654M-parameter language model designed for efficient on-device inference. IronLLM-0.6B combines a hybrid attention architecture with X-MTP, a lightweight shared-KV multi-token prediction design that eliminates per-depth KV-cache replay and employs a lightweight verification head for rollback-free drafting, achieving a 1.48x decoding speedup. The model is pretrained on approximately 6.2 trillion tokens using a quality-oriented data pipeline and is further post-trained with Multi-Domain On-Policy Distillation to integrate capabilities from domain-specialized teachers.

    embodied
  344. arxiv:2609.36858 · eess.SY
    Receding Horizon Control and Dissipativity - Optimal Control, Games and Uncertainty
    Sophie Hall, Jonas Schießl, Sergio Grammatico, Timm Faulwasser

    We provide a tutorial overview of receding horizon control across deterministic, stochastic, multi-agent, and game-theoretic settings. We contrast the minimizer of an optimal control problem with the equilibrium of a game in terms of cost and constraint handling to motivate the difference of MPC schemes involving multiple agents/players. Interestingly, dissipativity theory and the turnpike property have turned out to be a unifying thread and system-theoretic backbone across all these schemes. This paper is the first to give an overview of results, summarizing fundamental insights, drawing para

    multi-agent
  345. arxiv:2609.36855 · cs.AI
    When Upstream Messages Override Correct Answers: A Controlled Study of Multi-Agent LLM Collaboration
    Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen +6

    Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful information or an incorrect answer that causes a downstream agent to override a correct answer supported by its own evidence. Prior work has not clearly separated the benefits of communication from the damage caused by incorrect messages. We study this problem with controlled experiments across five benchmarks and five receivers, keeping the downstream task and evidence fixed while comparing answers under three conditions: no message, the upstream

    agentmulti-agentbenchmark
  346. arxiv:2609.36852 · cs.CV
    Socialality Anchors: Towards Group-bounded Trajectory Prediction
    Ziqian Zou, Conghao Wong, Qinmu Peng, Xinge You

    Trajectory prediction is a key component for understanding human behavior patterns in dynamic scenes. Researchers have devoted substantial efforts to modeling social interactions, especially group-wise interactions, since group membership often reflects shared intention, coordinated motion, and stable mutual adaptation, thus providing a persistent and semantically meaningful social prior for forecasting. However, existing group modeling methods may rely on a fixed threshold and infer groups mainly from agents' relative positions within the observation window, overlooking the fact that grouping

    benchmark
  347. arxiv:2609.36851 · cs.CV
    RoXDrive: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving via Action-Faithful Rollouts
    Hongbin Lin, Chaoda Zheng, Yiming Yang, Xiangyu Li +10

    End-to-end autonomous driving policies are commonly trained via imitation learning on logged demonstrations without observing the consequences of their own actions, leading to causal confusion in closed-loop real-world deployment. To address this issue, reinforcement learning (RL) post-training offers a promising alternative by leveraging world models as interactive training environments to enable future scene generation for policy improvement. Nevertheless, existing approaches either rely on reconstruction-based simulators, offering limited counterfactual interaction, or adopt synthetic simul

    sim-to-realworld modelpost-trainingevaluator
  348. arxiv:2609.36850 · cs.CV
    Rethinking Multimodal Fake News Detection in the Generative AI Era
    Wenbin Shen, Guoxuan Qin, Guangxu Yao, Baodong Wang +2

    Generative content is increasingly entering the production and dissemination of news, transforming fake news from manually fabricated or simply manipulated material into complex forms in which native and generated content jointly participate. Existing multimodal fake news detection research primarily focuses on veracity assessment and rarely characterizes how generativity differences affect the reliability of evidence. In contrast, AIGC detection primarily determines whether content is generated or modified by generative models, but it does not by itself establish whether the underlying news e

    benchmark
  349. arxiv:2609.36845 · cs.LG
    DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning
    Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen, Xianbin Cao +3

    Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollo

    world modelagentic
  350. arxiv:2609.36843 · cs.LG
    RolloutFaith: Auditing Persistent Internal Interventions in Visual World Model
    Junchi Yao, Ziyi Wang, Youling Huang, Lijie Hu

    Interpretability methods such as probes, activation patches and learned editors are designed to reveal or modify a model's current computation. World models pose a harder requirement: because their predictions become inputs to later predictions, a useful internal correction must survive after editing stops. We therefore propose RolloutFaith, a framework that measures semantic improvement both in the prediction produced at intervention time and over later autonomous predictions under fixed events, actions, noise, and information budgets. We evaluate ten fitted editors on three world models acro

    world modeldreamerv3memory
  351. arxiv:2609.36838 · cs.CV
    On-Policy Visual Evidence Distillation
    Shaohang Wei, Feifan Song, Guangyue Peng, Wenhao Yu +8

    Visual agents solve problems by interleaving reasoning with image operations, and on-policy distillation (OPD) provides guidance from a strong teacher on student-generated interaction trajectories. However, image operations change the evidence available for subsequent reasoning, so local errors in evidence acquisition (Acquire), reading (Read), or answer grounding (Ground) can propagate through the trajectory and lead to incorrect answers. Existing multimodal OPD methods primarily construct or contrast auxiliary views of the original image to strengthen supervision, without explicitly modeling

    benchmark
  352. arxiv:2609.36837 · cs.LG
    You Cannot Recover What Was Never Measured: Quantifying the Information Ceiling of Ultra-Low-Field MRI Super-Resolution
    Prathamesh Pradeep Khole, Shreya Handa, Utkarsh Gupta, Razvan Marinescu

    Generative super-resolution models can turn portable 64 mT MRI into images that look like 3T scans, and the field evaluates them with PSNR, SSIM, and pixelwise uncertainty, most often on pairs built by synthetically degrading high-field images. Prior work acknowledges that these models hallucinate and that the problem is ill posed, but to our knowledge no study measures how much information about the individual subject the real low-field scan actually contains. We measure it. Using paired 64 mT and 3T scans of the same subjects from three public datasets, and a measurement protocol validated o

    benchmark
  353. arxiv:2609.36835 · cs.AI
    ARC-KV: Amortizing Anchor Search for Reconstruction-Based KV Cache Compaction
    Zheyu Shen, Guanhua Wang, Dezhan Tu, Mengchi Zhang +4

    Long-context large language model inference is bottlenecked by KV caches that grow linearly with sequence length. This burden is especially severe for long, reusable context prefixes, whose cache must serve many downstream queries. Reconstruction-based methods such as Attention Matching achieve strong downstream task performance with compact KV caches. However, iterative anchor search dominates the compaction cost of OMP-based Attention Matching. This motivates our selective amortization principle of learning a reusable anchor-selection policy across contexts while retaining context-specific r

    long-context
  354. arxiv:2609.36830 · cs.AI
    Where Does Staleness Accumulate? Pool Aware Effective Staleness Control for Asynchronous RL in LLM Post-Training
    Chenliang Li, Neiwen Ling, Zijun Wei, Alfredo Garcia

    Fully asynchronous reinforcement learning (RL) improves resource utilization in large language model post-training by overlapping rollout generation with policy optimization, but it also introduces policy lag as trajectories are generated and queued while the trainer continues to update. We study how this lag accumulates over a trajectory's lifetime and how it can be controlled without sacrificing the wall-clock benefits of asynchronous execution. We decompose trajectory staleness into Generation Staleness, accumulated before rollout completion, and Waiting Staleness, accumulated after a compl

    post-trainingbenchmark
  355. arxiv:2609.36829 · cs.LG
    The Default Trap: Rethinking Plan Evaluation in Tool-Using LLM Agents
    Xueqi Li, Jingjie Ning, Yibo Kong

    An executor can respond strongly to a change in a supplied plan's priority while showing a small change in the same information-selection probability when a default-aligned whole plan is removed. We call the risk of interpreting the latter as weak responsiveness to alternative priorities the default trap. We compare paired plans that prioritize different information targets with a shared no-plan reference. An accounting identity relates these distinct behavioral contrasts. Across 3,200 decision windows on 160 selected Retail, Airline, and AgentDojo tasks, switching priorities strongly redirect

    llm agent
  356. arxiv:2609.36828 · cs.AI
    Calibrate the Decisions That Change the Future: On-Policy Post-Training Quantization for Multimodal Large Language Models
    Wenxiao Fan, Jingling Fu, Lichen Ma, Yu He +5

    Post-training quantization (PTQ) lowers deployment cost for multimodal large language models, but calibration typically reconstructs fixed sequences with local objectives. This overlooks autoregressive feedback: a quantization-induced token change redirects the prefix and changes future states. Yet on-policy coverage alone is insufficient because many decision mismatches barely affect future generation. We propose OnPTQ, an on-policy framework that calibrates on trajectories visited by the current quantized policy. On shared prefixes, OnPTQ identifies quantization-eroded boundaries, evaluates

    action-conditionedpost-training
  357. arxiv:2609.36822 · cs.CV
    RED: Reconstruction Evolution Dynamics for Generalizable AI-Generated Image Detection
    Wenpeng Mu, Junshan Jin, Tanfeng Sun, Xinghao Jiang +1

    The rapid evolution of image generators calls for forensic cues that generalize beyond known generation mechanisms. Existing detectors often rely on static image representations or endpoint reconstruction discrepancies, leaving the evolution of intermediate reconstruction stages underexplored. We observe that the relative token predictability of real and generated images can reverse across reconstruction scales, suggesting that intermediate stages may expose forensic evidence overlooked by endpoint comparisons. Motivated by this observation, we propose RED (Reconstruction Evolution Dynamics),

    benchmark
  358. arxiv:2609.36820 · cs.LG
    CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
    Wenbin Hu, Huihao Jing, Haochen Shi, Yuxuan Liu +2

    Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa, allowing update magnitudes to adapt to reward dependence. However, correlated rewards with large sc

    agenttool calling
  359. arxiv:2609.36816 · cs.LG
    Towards Better Training Signal: Advantage Clipped Policy Optimization
    Ruichuan Huang, Jinghan Liu, Congliang Chen

    Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy data substantially limits training efficiency. Reusing off-policy data through importance sampling (IS) can improve efficiency but introduce considerable instability. Hence, algorithms such as PPO and GRPO widely adopt IS-ratio clipping to stabilize training. However, training stability and gradient estimate are mainly determined by the product of IS ratio and advantage. To further stabilize training, we propose ACPO, which clips the product

    post-trainingbenchmark
  360. arxiv:2609.36812 · cs.RO
    Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
    Junhyun Ha, Juho Lee, Byungwoo Park

    Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To

    diffusion policy
  361. arxiv:2609.36810 · cs.CV
    MeteoVerse: Unified Weather-Controllable Video World Model
    Renlong Wu, Guanqiao Wang, Xuan Shang, Yin Hanming +4

    Video world models aim to predict future content from an observed scene while following prescribed camera motion. Real-world scene evolution is determined not only by changes in viewpoint and object dynamics, but also by environmental conditions such as weather, which can substantially alter scene appearance and visibility. Modeling such realistic weather evolution is challenging because the required weather modification depends jointly on the observed and desired weather states. Depending on their relation, the model may need to preserve, introduce, or remove a weather effect. Existing video

    world model
  362. arxiv:2609.36809 · cs.AI
    Geometry-Conditioned Fixed-Scaffold Encoders for Time-Warp Robust Sequence Retrieval
    Cassandra Yang, Yufan Tang

    Embedding-based retrieval is attractive for long sequence collections because each item can be encoded once and searched by nearest-neighbor ranking. The difficulty is that the objects being indexed are often observed under a noncanonical clock: cardiac cycles stretch with rate, speech changes with tempo, and sensor traces reach comparable states at different speeds. This paper studies a specific source of instability in patch-based encoders for this regime. If patch boundaries are chosen from signal geometry, then the tokenization can change under the same temporal deformation that the repres

    evaluation protocol
  363. arxiv:2609.36808 · cs.RO
    Spotter: Let the Embodied Model Lead, and the VLM Reflect for It
    Long Li, Qichao Zhao, Yue Yang, Fan Xu +5

    Current embodied models do not respond to their own failures, although what just went wrong could inform a small adjustment on the next attempt, the kind of reflection behind the gains of thinking in language models. We test whether they can repair a known error, which requires producing a correction and judging whether it is right. Stopped at a failure and allowed to retry, they seldom repair it through their own randomness or from a language description of the error, and best-of-N selection cannot pick the successful candidate after a failure. We attribute this to training only on successful

    embodiedrobotwin
  364. arxiv:2609.36806 · cs.AI
    CAD-Native Transformer Operators for AI-Aided Engineering
    Daniel Leibovici, Nikola Borislavov Kovachki, Dawon Ahn, Ruben Ohana +7

    Modern engineering systems, from automobiles to aircraft, are designed by using precise, continuous parametric computer-aided design (CAD) models. Evaluating design changes through numerical simulation requires meshing the continuous geometry, a computationally expensive and often brittle process that can require manual intervention and replaces the continuous representation with a discrete approximation. Most neural surrogates accelerate the simulation, but inherit this representation gap by relying on meshes, point clouds, voxels, or other sampled approximations of geometry. We introduce CAN

    benchmark
  365. arxiv:2609.36805 · cs.AI
    UpliftMem: Learning Set-Level Uplift for Agent Memory Retrieval
    Mengkun Liang, Haoran Qiang, Guannan Liu, Junjie Wu

    Large language model (LLM) agents reuse external memory to guide new tasks, but effective retrieval requires learning which memory sets improve execution. Such learning relies on costly outcome feedback: ordinary retrieval observes only executed sets, while evaluating alternatives requires additional rollouts. We introduce \textsc{UpliftMem}, which learns memory retrieval from set-level execution uplift relative to the same executor without memory. A theoretical analysis of how retrieval preferences restrict feedback coverage motivates targeted probing of alternative memory sets. Probe selecti

    memoryexternal memoryagent memoryagent
  366. arxiv:2609.36803 · cs.CV
    EGSD: Event-Grounded Self-Distillation for Streaming Video Understanding
    Yuwei Miao, Xuesheng Zhang, Wenhao Zou, Jixia Zhang +4

    Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the questi

    memorybenchmark
  367. arxiv:2609.36801 · cs.CV
    Scene Retargeting: Learning Object Placement with Analogical Transfer
    Minkwan Kim, Junho Kim, Seungmin Lee, Changwoon Choi +1

    Interactive simulations of embodied AI or spatial computing applications build on realistic 3D scenes that support daily activities. However, sparse, irregular layout structures impose scene-specific physical constraints, making it hard to define a generalizable framework for generating similar functional context. We formalize Scene Retargeting as stably transferring the semantically coherent spatial organization across layouts, rather than relying on textual descriptions or pairwise relationships. Our cluster-wise transfer flexibly handles mismatched object instances and adapts to distinctive

    embodied
  368. arxiv:2609.36800 · cs.AI
    AI as a Compiler: Compiling Triton kernels without the Triton compiler
    François Costa, Charly Castes, Thomas Bourgeat, Azalia Mirhoseini

    Compiler backends are expensive to build and maintain as programming models, workloads, and accelerators evolve. We investigate whether large language models can replace the conventional optimizing and lowering pipeline, a process that we call AI lowering. We study AI lowering from Triton to NVIDIA PTX: an LLM agent translates Triton kernels directly into PTX. We build an environment that evaluates candidate PTX, and an agentic harness in which an LLM translates Triton kernels into PTX. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels from recent ML papers, AI lo

    memoryagentllm agentagentic
  369. arxiv:2609.36798 · cs.LG
    Seeing What Should Be Heard: Diagnosing and Repairing Cross-Modal Shortcuts in Omni-Modal LLMs
    Yueran Ma, Ronghao Lin

    Omni-modal large language models (LLMs) are expected to answer a question using the modality it explicitly refers to. However, existing training paradigms rarely verify whether models actually follow this modality, because multimodal inputs from the same sample often provide redundant evidence for the same answer. In this work, we uncover a pervasive cross-modal shortcut in omni-modal LLMs: when asked an audio-related question, models rely on the image as much as on the audio, and sometimes even more. To systematically diagnose this behavior, we introduce the Factorized Modality Diagnostic, wh

    post-trainingbenchmark
  370. arxiv:2609.36789 · cs.MA
    GitHarness: Git Init Your Harness Working Memory for Perpetual User Requirements
    Zhibang Yang, Xinke Jiang, Yuxuan Liu, Mingyu Zhang +8

    LLM-based agents increasingly collaborate with users on long-horizon tasks, accumulating evidence, code, and drafts through extensive search, reasoning, and execution. As users inspect these results, they may supply missing information requirement completion, introduce new requirements requirement elicitation, or revise existing ones requirement shift. These changes often affect only part of the accumulated work, yet agents may carry forward obsolete information or turn local revisions into global rewrites. Existing approaches clarify current intent without determining how prior work should ch

    memoryagentagenticbenchmark
  371. arxiv:2609.36788 · cs.LG
    Harnessing Large Language Models to Compile Task-Relevant Context into Bayesian Optimisation
    Zhongwei Yu, Sourabh Roy, Bin Cao, Xue Yan +2

    Incorporating rich task-relevant context, such as domain knowledge and external observations, is a key capability yet remains challenging for Bayesian optimisation (BO). Recently, practitioners have started to use large language models (LLMs) to generate and execute BO programs through coding harnesses. In such emerging practices, the posterior belief is shaped not only by Bayesian inference but also by LLM-generated model and data artefacts, offering a flexible route for task context to enter BO as executable code. To study whether and how LLMs can be harnessed to compile diverse contextual s

    benchmark
  372. arxiv:2609.36787 · cs.MA
    Regularized policy gradient with learned mixtures of Gaussians for games with continuous actions
    Ondřej Kubíček, Viliam Lisý, Tuomas Sandholm

    Most successes of superhuman game-playing algorithms are in games with discrete actions, yet in auctions, robotics, sports, or trading, actions are nearly continuous. Prior techniques either rely on expert-designed discretizations or are sample inefficient. We present a scalable policy-gradient algorithm for large sequential games with continuous or mixed discrete and continuous actions. It combines magnetic mirror descent with a mixture of Gaussians reparametrization, trained via self-play. We show that it approximates equilibrium in games where gradient descent fails. In sequential games, it

    self-play
  373. arxiv:2609.36785 · cs.RO
    TaRL: Learning General and Physical Rewards from Tactile Demonstrations
    Po-Yi Wu, Dao-Jan Chang, Shang-Ya Hsiao, Hong-Ming Chen +2

    Contact-rich manipulation requires robots to sequence precise contacts, maintain stable grasps, and apply directed forces. Reinforcement learning (RL) can acquire such behaviors automatically, but its performance hinges on reward design: sparse rewards reduce the learning efficiency, while dense rewards are hard to specify. Visual reward learning addresses this by inferring rewards from action-free demonstrations. Because it conditions only on visual observations, it fails to capture rewards beyond visual goals. We propose Tactile Reward Learning (TaRL), a framework that learns rewards from ta

    embodiedmanipulationtactilegrasp
  374. arxiv:2609.36784 · cs.RO
    Scale-Invariant Manipulability Shape Tracking Across Heterogeneous Manipulators
    Geunwoo Kwon, Dong-gyu Lee, Kai Li, Soonwoong Hwang +1

    When transferring manipulability across systems with different sizes and kinematic structures, matching absolute ellipsoid scale may be unnecessary when the goal is to reproduce orientation and semi-axis length ratios. Full-matrix tracking, however, penalizes both shape and absolute-scale differences, even when only shape matching is required. We therefore propose a scale-invariant manipulability shape-tracking method that treats matrices differing only by a positive scalar factor as equivalent and uses their unit-determinant representatives. We derive the differential of the unit-determinant

    manipulator
  375. arxiv:2609.36782 · cs.CV
    Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning
    Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu +1

    While multimodal large language models (MLLMs) have demonstrated exceptional capabilities in objective understanding tasks, their performance in affective reasoning still falls significantly short of human standards. We attribute it to a central capability gap: MLLMs are difficult to reliably distinguish semantically proximal emotions based on fine-grained visual evidence, which could be decoupled as two limitations: 1) Insufficient Attribution. The global reasoning paradigm of conventional MLLMs severely dilutes fine-grained emotion cues, where subtle emotional states are usually implicitly e

    benchmark
  376. arxiv:2609.36779 · cs.RO
    DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients
    Xiaotian Zhang, Yusheng Wang, Naoya Kagawa, Noritaka Takamura +3

    Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations

    manipulationgrasp
  377. arxiv:2609.36777 · cs.AI
    Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes
    Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang +5

    Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning fro

    agentbenchmark
  378. arxiv:2609.36776 · cs.CV
    Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention
    Cheng Ye, Weidong Chen, Peipei Song, Zhendong Mao

    Emotional Video Captioning aims to generate factually accurate and emotionally empathetic descriptions. While recent methods have recognized the importance of visual causes to guide emotion perception and caption generation, they fundamentally rely on simple attention matching, which inevitably suffers from {causal redundancy and spurious correlations} in co-occurrence bias (e.g., misclassifying ``sadness'' as ``joy'' on a sunny beach), leading to severe shortcut learning from confusing backgrounds. Furthermore, existing evaluations fail to verify whether models have genuinely mastered causal

    benchmark
  379. arxiv:2609.36775 · cs.CV
    DSPO: Diversity-aware Subjective Policy Optimization for Robust Emotional Reasoning
    Cheng Ye, Weidong Chen, Bingyan Xu, Zhendong Mao

    Reinforcement Learning has significantly advanced the complex reasoning capabilities of MLLMs. However, prevailing RL algorithms suffer a severe failure in emotion reasoning tasks. These methods heavily rely on deterministic hard-label supervision and point-wise isolated evaluation, creating a fundamental gap with the inherently subjective and continuously distributed nature of human emotions. Furthermore, unlike explicit physical objects, emotional states are deeply implicit within visual cues. This abstract nature exacerbates visual hallucinations in MLLMs, leading to plausible yet ungrounde

    benchmark
  380. arxiv:2609.36774 · cs.RO
    LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks
    Zeming Wei, Jianheng Ye, Xinshuai Song, Sirui Chen +2

    Vision-language-action (VLA) models struggle to reuse recurring interactions in unseen tasks. Our diagnostic study reveals that reliable task completion does not imply consistent execution of constituent atomic actions across task contexts. We present LexiconVLA, a retrievable atomic-action lexicon for cross-task reuse. Global and detail codebooks capture shared interaction structure and fine-grained execution variation, respectively, preserving both reusable patterns and execution details. Visual-Atomic Action Alignment couples trajectory reconstruction from visual state changes with visual o

    vision-language-action
  381. arxiv:2609.36770 · cs.LG
    Emergent Specialization in Populations of Self-Supervised Collaborative Vision Experts Without a Shared Gate or Cross-Agent Gradients
    Aram Davtyan, Pablo Acuaviva, Sebastian Stapf, Paolo Favaro

    Can a population of neural networks develop a useful division of labor without a shared gate or gradients between agents? We study a setting where each network has its own weights, trains independently on the same heterogeneous data, and can ask another agent for help through a forward pass. Unlike mixtures of experts, where a jointly trained gate assigns inputs to experts, specialization here must emerge without central control. We test this in a small scale proxy for predictive visual pretraining. Initially identical agents are finetuned on an unlabeled mixture of six visual domains using ma

    agent
  382. arxiv:2609.36765 · cs.LG
    Graph-Spectral Flow Matching for Multivariate Time Series Anomaly Detection
    Zepeng Zhang, Jhony H. Giraldo, Wenbin Wang, Olga Fink

    Multivariate time series anomaly detection typically relies on evaluating discrepancies between observations and outputs produced by models trained on normal data. An alternative perspective is to characterize the distribution of normal data through the generative dynamics, i.e., the velocity field, of flow matching models. However, standard flow matching typically adopts linear probability paths that overlook dependencies among variables, leading to a misalignment with the structured data distribution. To address this issue, we propose GRASP, a flow matching framework with a graph-spectral pa

    graspbenchmark
  383. arxiv:2609.36760 · cs.LG
    QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching
    Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang +4

    Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memory still scales linearly with context length and batch size. In this work, we establish a systematic model of MLA's dual-path quantization errors, characterizing their distinct effects on attention-output distortion and explaining the pronounced amplification of RoPE-path errors. Guided by this analysis, we introduce QuantMLA, a function-aligned framework for low-bit dual-path quantization. We derive path-specific transformation spaces that prese

    memorybenchmark
  384. arxiv:2609.36757 · cs.CV
    FastVR: Efficient Streaming Video Restoration with One-Step Diffusion
    Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng +1

    Diffusion-based video restoration recovers realistic details, but its practical deployment is limited by two efficiency bottlenecks: costly VAE encoding and decoding, and the quadratic cost of full self-attention in diffusion transformers (DiTs). This paper presents FastVR, a streaming video restoration framework built on a one-step diffusion model, which delivers strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. To improve inference efficiency, FastVR combines a lightweight VAE with chunk-wise causal attention, which substantially

    benchmark
  385. arxiv:2609.36755 · cs.CV
    Drag as Evidence: Motion-Grounded Latent Recomposition for Drag-Based Editing
    Xinyu Pu, Hongsong Wang, Jie Gui, Pan Zhou

    Modern image editors excel at semantic manipulation and visual synthesis, yet remain limited in precise spatial control, motivating the development of drag-based editing. However, existing drag-based methods often struggle to balance drag accuracy with natural, plausible, and intent-aligned generation. We propose MoRe-Drag, a motion-grounded drag-based editing method. Our key insight is to treat pixel-space warping as coarse motion evidence, and to inject this evidence into the generative sampling trajectory. Specifically, MoRe-Drag performs region-aware latent recomposition over refinement, i

    manipulation
  386. arxiv:2609.36753 · cs.RO
    Degeneracy-Orthogonal Geometric Constraints for LiDAR SLAM
    Minseo Kim, Yina Kim, Jinhwa Hwang, Alex Junho Lee

    Autonomous robot navigation relies on simultaneous localization and mapping (SLAM) to estimate motion and maintain an accurate pose within an environment. However, in axially uniform corridors such as long tunnels and pipelines, LiDAR odometry is fundamentally limited by unconstrained drift along the feature-weak travel direction. This structural degeneracy cannot be resolved by local scan matching alone. To address this challenge, we propose the Degeneracy-orthogonal Contour Offset Descriptor (DeCOD), a structure-aligned geometric descriptor for cross-sectional landmarks. Cross-sectional boun

    benchmark
  387. arxiv:2609.36752 · cs.LG
    cktFormer: Transformer-Based Approach for Automated Analog Circuit Design
    Pasindu Dodampegama, Praveen Wijesinghe, Naveen Basnayake, Keshawa Jayasundara +1

    Circuit design is a complex and iterative process that requires expertise in electronic engineering. It involves selecting components while meeting performance constraints, such as power efficiency, cost-effectiveness, and signal integrity. However, manual design is time-consuming and prone to errors. Although other stages of the manufacturing pipeline have benefited from AI-driven optimizations, circuit design remains a bottleneck, limiting overall productivity. Generative AI and machine learning offer the potential to automate and improve this stage, boosting efficiency and accuracy. To addr

    benchmark
  388. arxiv:2609.36750 · cs.LG
    Group-Marginalized Self-Rewarding RL Drives Zero-Label Self-Evolving
    Yiming Wang, Yikang Liu, Qingyuan Tian, Xingyu Chen +3

    Self-rewarding reinforcement learning (RL) enables large language models (LLMs) to self-evolve without human labels. Existing ensemble-based methods construct reward references from rollout groups and assign rewards accordingly. However, a response's reward representation also depends on its randomly sampled group context, i.e., the other responses in its group. Using only one group-context realization may miss desired reward signals and provide unreliable guidance for policy optimization. To address this issue, we propose Group-Marginalized Advantage Estimation (GMAE), which aggregates reward

    self-evolvingbenchmark
  389. arxiv:2609.36746 · cs.AI
    EASE: Behavior-Adaptive Skill Curation for Self-Evolving Agents
    Zhen Xiong, Qiaoyu Tan

    Agent skills provide a lightweight mechanism for self-evolving agents to accumulate reusable procedural knowledge without updating model parameters. However, existing learned skill curators typically optimize curation without explicitly modeling downstream executor behavior. We show that this can cause systematic cross-executor degradation: curators trained with different executors perform best when paired with their own training executor, indicating that effective skill curation is executor-dependent. We formulate behavior-adaptive skill curation and introduce EASE, a framework that learns a

    agentself-evolving
  390. arxiv:2609.36742 · cs.AI
    SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation
    Zhenrui Yue, Huimin Zeng, Yueqi Wang, Yaokun Liu +7

    Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for improving large language models (LLMs) on various tasks, yet its sparse outcome rewards lack token-level credit assignment for intermediate steps. To address this, on-policy self-distillation (OPSD) leverages a self-teacher with privileged context to provide additional dense learning signals. However, because the self-teacher is often overconfident and imposes excessive penalties on long reasoning trajectories, OPSD frequently struggles in practice. To mitigate this, we propose self-instructing policy opti

    benchmark
  391. arxiv:2609.36741 · cs.AI
    Distinguish or Homogenize: Last-Chance Policy Identification and Risk-Budgeted Recovery under Irreversible Resource Depletion
    Yibo Guo, Xiaodan Wang

    Under irreversible resource depletion, an agent can spend resources to distinguish among latent fault models, or to change the system state so that the remaining models admit a common acceptable continuation--at which point further diagnosis becomes unnecessary. This distinguish-or-homogenize principle identifies a path that existing frameworks for identification, planning, and diagnosis do not make explicit: prior formulations treat the mapping from fault models to acceptable policies as a given, whereas LCPI makes it a function of the agent's own actions. We formalize this principle through

    agent
  392. arxiv:2609.36739 · cs.AI
    Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change
    Bravish Ghosh

    Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension ru

    memorymulti-agentbenchmark
  393. arxiv:2609.36734 · cs.CL
    Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
    Ayan Sengupta, Vaibhav Seth, Tanmoy Chakraborty

    Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adap

    benchmark
  394. arxiv:2609.36730 · cs.AI
    Can Agents Design Libraries for Agents?
    Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen +3

    Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-vali

    agentbenchmark
  395. arxiv:2609.36726 · cs.AI
    Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes
    Kargi Chauhan

    Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8

    agentbenchmark
  396. arxiv:2609.36722 · cs.CL
    ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation
    Xinghao Chen, Junnan Dong, Cai Ke, Chak Tou Leong +7

    Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mism

    memorybenchmark
  397. arxiv:2609.36720 · cs.RO
    T$^2$Mem: Learning Test-Time Memory for Robotics
    Yize Liu, Huang Huang, Yining Hong, Zijian Du +3

    Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use historical information from action demonstrations alone, without external memory support. We introduce T$^2$Mem, a framework that develops this capability within a pretrained vision-language-action policy, without external reasoning models or memory-specific annotations. T$^2$Mem

    vision-language-actionmanipulationmemoryexternal memorybenchmark
  398. arxiv:2609.36707 · cs.CL
    LAURA: Knowledge Distillation for Interpretable Ambiguous Clause Identification in Legal Contracts
    Amrita Singh, Aditya Joshi, Jiaojiao Jiang, Hye-young Paik

    Legal contracts contain ambiguities that expose enterprises to financial and legal risks. Some ambiguities allow flexible interpretation without triggering disputes, while others lead to significant legal conflicts. This makes identification alone insufficient, and interpretable rationale analysis essential. We propose LAURA, a post-training framework for interpretable ambiguous clause identification. LAURA leverages knowledge distillation with an IRAC-Unlearning prompting technique to transfer knowledge from a teacher LLM to an open-weight student model (<=1B parameters), which is then traine

    post-training
  399. arxiv:2609.36704 · cs.LG
    When Is Coarse Supervision Worth It? Cost-Aware Learning under Unknown Aggregation
    Jianyu Xu, Smriti Jha, Aarti Singh, Bryan Wilder

    Modern learning systems often acquire supervision at multiple resolutions, trading annotation cost against information content. We study cost-aware two-resolution learning, where expensive fine labels reveal a vector response and cheaper coarse labels reveal a scalar aggregate formed with unknown weights, while the target remains the full response. The challenge is that unknown aggregation changes which directions coarse data can identify, so the value of coarse supervision depends jointly on cost, noise, and identification. We characterize this information geometry and develop an estimate-and

    benchmark
  400. arxiv:2609.36700 · cs.AI
    Lost in Conversation or Lost in Translation? Diagnosing Multi-Turn Degradation in RAG
    Pranav Handa, Ariful Azad

    When conversing with large language models (LLMs), users often begin with a simple question and build towards a multi-hop question through follow-up turns. Retrieval-augmented generation (RAG) and its graph-based variant (GraphRAG) have become the dominant approaches for grounding LLM responses in external evidence, yet both are evaluated almost exclusively on single-turn, fully specified queries. We systematically investigate this evaluation mismatch through a large-scale simulation study. Building on prior work on multi-turn LLM evaluation, we transform questions from multi-hop question answ

    retrieval-augmentedragbenchmark
  401. arxiv:2609.36698 · cs.LG
    Learned Queries and Keys Are All You Need: Replacing the Value Projection with Structured Transforms
    Ene Meco, Emadeldeen Hamdan, A. Enis Cetin

    To reduce the number of parameters and cache memory requirements of transformers we introduce dual-headed transformers instead of three heads. We studied Walsh-Hadamard Transform (WHT), Discrete Cosine Transform (DCT), Discrete Fourier Transform, filterbank based Shearlet Transform, and Multiplication-Avoiding (MA) operators to construct dual heads. We combine spatial patches and their orthogonal transforms (or Shearlet and MA operators) in a structure similar to the attention block. We obtained better results than triple headed transformers in ImageNet. Extensive simulation examples are prese

    memory
  402. arxiv:2609.36695 · cs.LG
    Know Thyself, Teach Thyself: Internal Information Flow for Selective Self-Distillation
    Rui Wang, Ruijie Wang, Bo Chen, Jiangxuan Long +1

    Self-distillation turns knowledge distillation into a closed learning loop and offers a path toward recursive self-improvement. Without an external teacher, however, the model must determine both what information can improve its supervision and which induced changes should be learned. Existing methods typically improve teacher-generated data or select training examples in isolation, leaving the information transferred between these stages unmeasured. We introduce InFlow, a retrieval-guided on-policy self-distillation framework that models this process as potential-to-realized information flow.

    self-improvement
  403. arxiv:2609.36691 · cs.CL
    Video2Skill: From Streaming Experience to Reusable Embodied Skills
    Jianshu Zhang, Ce Zhang, Xiyuan Yang, Chenwei Xu +5

    Manipulation behaviors vary widely across objects and scenes, but they share a small set of reusable skills, and planning with these skills helps embodied agents generalize to new tasks. Yet an agent can only plan with skills it knows. Recovering skills from observed experience, the inverse of planning, builds this knowledge over time and yields skill data for training future agents. Vision-Language Models (VLMs) describe individual manipulation events well, but can they organize a stream of events into reusable skills? We formulate this problem as Streaming Embodied Skill Discovery (SESD): a

    embodiedmanipulationagentembodied agentbenchmark
  404. arxiv:2609.36690 · physics.optics
    Co-design of Silicon Microring Modulator beyond 200 Gb/s per Lane: Device Physics, Operating Point, and Compact Models for Scale-Up and Scale-Out Optical I/O
    Zhihong Huang, Yuan Yuan, Yiwei Peng, Samuel Palermo +2

    Optical input/output (I/O) supports high-bandwidth communication between processors in artificial intelligence (AI) systems. Depletion-mode silicon microring modulators offer compact footprints, wavelength multiplexing and femtojoule-scale junction switching energy per bit. However, as lane rates increase to 200~Gb/s and beyond, a digital signal processor (DSP) performing equalization and forward-error correction can consume up to half of the optical module power. We analyze the junction and cavity physics of silicon microring modulators, relating modulation efficiency, capacitance, optical lo

    ring modulatormicroring
  405. arxiv:2609.36689 · cs.LG
    CHAIN: Calibrated LLM Forecasting via Causal-Temporal Hypergraph Inference
    Wenjin Liu, Chenxi Wang, Yue Lu, Zhe Cui +1

    Large language models have achieved significant progress in event forecasting, yet their probability outputs exhibit systematic calibration bias that varies heterogeneously across different domains and question types, undermining the trustworthiness of probabilistic outputs for decision-making under uncertainty. However, existing calibration methods typically correct probability outputs after prediction is complete, without modeling the structural sources of bias within the prediction process itself. To address this challenge, we decompose probabilistic prediction over causal-temporal hypergra

    benchmark
  406. arxiv:2609.36686 · cs.LG
    Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition
    Hada Melino Muhammad, Luan Pham, Laure Barrière, Sachin Shetty +2

    Identifying the root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top@k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes, retrieval failure, where the true cause is never considered, and reranking failure, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments

    benchmark
  407. arxiv:2609.36685 · cs.CV
    When Semantics Matter: Reliability-Aware Semantic-Rhythm Control for Co-Speech Gesture Generation
    Zhirui Xing, Long Ye, Kaige Li, Ziyi Xu +1

    Co-speech gesture generation aims to synthesize natural gestures that are both temporally synchronized with speech and semantically consistent with the spoken content. Although recent methods can generate rhythmically plausible motions, they often rely heavily on acoustic prosody while underutilizing textual semantics, especially when semantic annotations are incomplete, noisy, or unavailable. Consequently, the generated gestures may follow speech rhythm while failing to express the intended semantics. To address this problem, we propose a reliability-aware semantic-rhythm control framework fo

    benchmark
  408. arxiv:2609.36684 · cs.CL
    ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
    Jianshu Zhang, Keliang Wu, Chengxuan Qian, Xiyuan Yang +5

    Embodied agents now take on ever longer tasks. For long tasks, knowing only whether a task finally succeeds or fails says little; the steps along the way matter. Progress Reward Models (PRMs) score how far a task has come at every step, and serve as dense rewards, verifiers and monitors. Yet in long tasks the current frame alone often cannot tell how far the task has come, because progress depends on what happened before. We call this problem context-dependent progress estimation. Existing benchmarks on progress estimation mostly focus on short tasks whose progress can be read from the current

    embodiedmanipulationautonomous agentagenticembodied agentbenchmark
  409. arxiv:2609.36683 · cs.LG
    MARCO: Multi-Round Agentic Reinforcement for Conditional Molecular Optimization
    Shicheng Fang, Yuxin Wang, Zhuo Yang, Xiaohu Xu +5

    Molecular optimization is inherently iterative: a candidate is proposed, evaluated against several objectives, and revised while preserving a relationship to the source molecule. Most instruction-following models instead emit one edited molecule, forcing validity, property improvement, and similarity control into a single response. We introduce MARCO, an evaluator-grounded reinforcement-learning framework that trains molecular editors on bounded proposal--feedback--revision trajectories. MARCO aggregates shaped turn rewards into an undiscounted trajectory return for group-relative policy optim

    agenticbenchmarkevaluator
  410. arxiv:2609.36680 · cs.CV
    Reprogramming Vision-Language Models via Structured Prompt Reparameterization
    Zizhao Li, Chengyi Cai, Mohammed Yaqoob Ansari, Feng Liu +2

    Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framewor

    benchmark
  411. arxiv:2609.36679 · cs.AI
    MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development
    Xin Yu, Lizhu Zhang, Jiamu Bai, Yanhong Wu +6

    Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagn

    tool use
  412. arxiv:2609.36677 · cs.CV
    ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking
    Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen +2

    Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This repr

    world model
  413. arxiv:2609.36676 · cs.RO
    Kinematic Nonlinear Spatio-Temporal Trajectory Warping for Contact-Rich Dexterous Manipulation Demonstrations
    Hyojae Park, Arjun S. Lakshmipathy, Nancy S. Pollard

    We present a straightforward but effective method for repurposing existing contact-rich dexterous manipulation demonstrations. Starting from inputs of hand and object trajectories, our method outputs high-quality nonlinear trajectory warps that account for intermediate waypoints, environmental barriers, temporal shifts, and varied start/end configurations. Foundational to our method is the utilization of contact distributions, which we show allows us to reliably compute complex and high-dimensional dexterous hand trajectories following a simple object-centric warp specification pipeline. We ev

    manipulationdexterousmanipulator
  414. arxiv:2609.36675 · cs.CL
    Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement
    Ziqi Zhao, Fanqing Meng, Haocheng Lu, Lingxiao Du +3

    Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model's own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth.

    memoryagentmulti-agentagent frameworkself-improvementleaderboard
  415. arxiv:2609.36670 · cs.AI
    FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation
    Song-Li Wu, Weinan Gan, Zhaocheng Du, Xianquan Wang +1

    A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients

    benchmark
  416. arxiv:2609.36659 · cs.LG
    On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
    Shufan Shen, Zhongni Hou, Junshu Sun, Yufei Zhang +4

    The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent direction

    post-training
  417. arxiv:2609.36657 · cs.LG
    Constitutional adapters: Inference-time interventions for misalignment and misuse
    Adam S. Lowet, Mark Kurzeja

    Training models to act in accordance with an explicitly defined set of principles, or "constitution," has shown promise as a robust and transparent mechanism for AI alignment. However, the generality and flexibility of such methods remain unclear. Here, we show that constitution-consistent behavior can be distilled from synthetic corpora into lightweight objects (low-rank adapters and steering vectors). Despite never seeing a harmful request or jailbreak during training, such objects increase jailbreak defense success and measured alignment -- particularly at long context lengths and against m

    long context
  418. arxiv:2609.36655 · cs.CV
    Not Every Correction Helps: Gain-Guided Continual Test-Time Adaptation
    Youjia Zhang, Huiling Liu, Soyun Choi, Jaehong Yoon +1

    Continual test-time adaptation (CTTA) adapts a source model to an unlabeled test stream whose distribution may change over time. Existing TTA methods often assess prediction reliability using confidence or entropy, which primarily reflect the model's self-certainty for the current sample. In CTTA, accumulated target observations can provide complementary evidence for correcting the source prediction, but this history may become misaligned as the target distribution changes. The key question is therefore not how much the correction differs from the source prediction, but whether and how strongl

    benchmarkevaluator
  419. arxiv:2609.36654 · cs.LG
    Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
    Ruiyi Ding, Jie Li, Kang He, Ziyan Liu +3

    Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share an E4M3 block scale. Choosing that scale is difficult in GPTQ because quantizing one column updates those that follow, so evaluating a block independently can misestimate its final reconstruction error. Large models pose a second challenge: full-precision weights, calibration ac

    memorybenchmark
  420. arxiv:2609.36652 · cs.AI
    RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation
    Zixuan Yang, Yiqun Chen, Qi Liu, Wei Yang +5

    Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local f

    benchmark
  421. arxiv:2609.36651 · cs.CV
    FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
    FangZhi Zhong, Xuerui Qiu, Yuqi Pan, Ya Liu +3

    Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines compressed low-DPI global views with selective region enhancement, integrating enhanced views into on

    memorylong-context
  422. arxiv:2609.36648 · cs.CV
    VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training
    Yuanwei Hu, Bo Peng, Yuheng Jia, Xinting Hu +2

    Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics to complement visual representations. Despite the rapid proliferation of LaIC methods, it remains unclear how much LaIC has actually advanced image clustering, as existing studies generally suffer from major limitations, including inconsistent experimental settings, inadequate dataset selection, and limited evaluation dimensions. To address this gap, we introduce VLM4Cluster, a comprehensive benchmark for image clustering in the era of pre-trai

    benchmark
  423. arxiv:2609.36645 · cs.RO
    Where Predictive Supervision Goes Shapes What VLA Policies Learn
    Hanseul Kim, Jewon Yeom, Youngjoon Jeong, Minsoo Jo +1

    Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through control

    vision-language-actionvlavla policy
  424. arxiv:2609.36644 · cs.CV
    OCA: ODE-Driven Cross-Attention for Image-to-Point-Cloud Registration
    Pei An, Jiaqi Yang, Yulong Wang, Siwen Quan +1

    Cross-attention is a crucial component in learning-based image-to-point-cloud (I2P) registration. Although existing cross-attention mechanisms have achieved promising progress, attention ambiguity remains a fundamental challenge that hinders the learning of discriminative 2D-3D correspondences. To address this problem, we revisit cross-attention and establish ordinary differential equations (ODEs) to model the ideal I2P feature interaction. Based on this formulation, we develop an ODE-driven cross-attention (OCA) module that refines feature representations and attention matrices through ODEs.

    benchmark
  425. arxiv:2609.36642 · cs.LG
    PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning
    Muyang Li, Jie Yang, Zhengyu Fang, Junchao Zhu +4

    Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specif

    agentic
  426. arxiv:2609.36641 · cs.LG
    Inducing Process Supervision from Outcome-Only Reinforcement Learning
    Shengda Fan, Xin Cong, Zhong Zhang, Haotian Chen +1

    Process reward models (PRMs) have become a key component for LLMs, as their step-level feedback supports both post-training and test-time reasoning. However, training strong PRMs remains costly: human step annotation is difficult to scale, while Monte Carlo estimation is computationally expensive and can drift from the intrinsic correctness of steps. To get effective PRMs at low cost, we introduce TIPS (Thinking-Induced Process Supervision), an outcome-only reinforcement learning (RL) framework for training generative PRMs. In TIPS, the model generates a chain-of-thought (CoT) followed by step

    agentagent benchmarkpost-trainingbenchmark
  427. arxiv:2609.36640 · cs.RO
    Foundation-Model-Guided Topology-Aware Semantic Risk Fields for Manipulation
    Giung Lee, Weihang Guo, Lydia E. Kavraki

    Robot motion planning in everyday environments must satisfy hard geometric constraints while accounting for context-dependent semantic risk. We present a foundation-model-guided, topology-aware semantic risk field that extends manipulation safety beyond collision avoidance. For each manipulated-object/scene-object pair, a foundation model provides six directional risk weights and a pair-specific spatial decay scale. The method combines these priors with voxelized 3D scene geometry using topology-aware shielding and geodesic spatial decay. A GPU-parallel backend batches object-level distance an

    manipulation
  428. arxiv:2609.36639 · cs.LG
    A Digital Simulation Toolkit for Physics-Based Generation of Realistic Experimental Scanning Tunneling Microscopy Images
    Huanhuan Zhao, Laxmi Bhurtel, Connor Vernachio, Fahmy Paiziah +2

    Scanning Tunneling Microscopy (STM) is a widely used tool for characterizing surfaces of materials at the atomic scale, playing a crucial role in discoveries across condensed matter physics and materials science. Despite its extreme spatial resolution, STM is one of the most sensitive microscopy techniques and is highly prone to noise. While existing unsupervised denoising methods are very cheap to train, these are primarily focused on removing the noise with minimal recovery of key physical information. While supervised methods can offer superior performance, the major bottleneck is that a la

    benchmark
  429. arxiv:2609.36638 · cs.LG
    PE-OPSD: Internalizing Prompt Enhancement into Flow-matching Models via On-Policy Self-Distillation
    Mingfeng Lin, Chengfei Cai, Lin Xu, Chengqian Ma +2

    Text-to-image users often provide concise and underspecified prompts, whereas generative models benefit from detailed textual conditions for reliable instruction following. Existing systems bridge this gap with Prompt Enhancers (PEs) that rewrite raw prompts at inference time, introducing additional latency and leaving prompt elaboration external to the generator. We instead view enhanced prompts as privileged training information and ask whether their benefits can be internalized. We propose Prompt-Enhanced On-Policy Self-Distillation (PE-OPSD) for text-to-image flow-matching models. During t

    benchmark
  430. arxiv:2609.36635 · cs.AI
    WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses
    Haomin Qi, Xiangzhe Xu, Yiming Huang, Jingbo Shang +1

    Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each project, and retains cases exposed by a constructi

    agentagent frameworkbenchmark
  431. arxiv:2609.36630 · cs.AI
    Distilling Agentic Systems: A Roadmap across Models, Artifacts, and Harnesses
    Ziluowen Luo, Senzhang Wang, Chaozhuo Li, Jun Yin +8

    Modern agents increasingly rely on memories, tools, and execution logic, so their competence extends beyond model parameters. This shift exposes a limitation of conventional knowledge distillation, which asks how a student model imitates a teacher model. We define Agent Distillation as the persistent transfer of task-solving knowledge from a teacher agent to a student agent. Our study organizes the field by where transferred knowledge is retained: within the model, as artifacts, through the execution harness, or across substrates. This perspective separates transfer evidence from its outcome a

    agentagenticevaluation framework
  432. arxiv:2609.36628 · cs.CV
    Beyond Binary Preferences: Graded Preference Optimization for Limb-Motion Captioning
    Yanan Wang, Tingsong Li, Kaixun Jiang, Chongyang Zhong +2

    Vision-Language Models (VLMs) can generate rich video captions, yet often misidentify which person performs an action or which limb is involved, particularly across camera cuts. Improving these details requires evaluation and training that distinguish missing information from incorrect assertions. We introduce FlexBench, a benchmark spanning 3,105 shots and 18,161 evaluation queries, with human-verified identities and systematic per-person coverage of fine-grained limb actions and states. Its reference-derived checklists support automated assessment of complete captions in their person and sho

    benchmark
  433. arxiv:2609.36626 · cs.AI
    Semantic Projection for Continual Self-Evolution of Language Agents
    Ziyu Liu, Jun Chen, Lixu Wang

    Language-model agents increasingly rely on persistent natural-language skills to adapt beyond their frozen model parameters. When a shared skill is repeatedly revised from a non-stationary, heterogeneous task stream, however, improvements for new tasks can overwrite procedures needed for earlier ones. In continual learning, Orthogonal Gradient Descent (OGD) addresses analogous interference by projecting a new-task gradient onto a subspace that locally preserves prior predictions. Natural-language skill revisions, however, have neither gradients nor a canonical vector space in which such a proj

    agent benchmarkbenchmark
  434. arxiv:2609.36621 · physics.optics
    Attosecond circular-dichroism spectroscopy of hole ring currents
    Guangru Bai, Zhihui Lyu, Jinlei Liu, Jing Zhao +1

    Ultrafast ionization of atoms by circularly polarized few-cycle laser pulses generates a hole ring currents, offering a route for ultrafast manipulation of magnetism. Here we explore the subcycle formation of these currents whose circulation direction is determined by the driving-field helicity, whereas their magnitude is governed by the quantum coherence of the residual ion. We show that the correlated ion-photoelectron dynamics can be probed with attosecond transient-absorption circular dichroism. Using the Wigner-Eckart theorem, we derive a linear relation between state-resolved dichroic ab

    manipulation
  435. arxiv:2609.36620 · cs.AI
    Neural Structural Reasoner: A Brain-inspired Architecture for Reasoning over Structured Knowledge
    Zixing Jia, Yuhang Pan, Ni Ji

    Structural reasoning, the ability to recognize and make inferences over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowled

    benchmark
  436. arxiv:2609.36617 · cs.CL
    Generating Edit-Inducing Questions for AI Research Manuscripts
    Sebastian Joseph, Zichao Wang, Jennifer Healey, Alexa Siu +2

    We study the ability of LLMs to generate edit-inducing questions whose answer will improve a paper draft. On a dataset of paired submission and camera-ready papers from ICLR and NeurIPS, we compare the helpfulness of questions from GPT models with or without full paper context to that of human reviewers. GPT produces more edit-inducing questions and its questions are associated with more extensive edits and cover a broader range of edited content compared to questions from reviewers. However, a much smaller percentage of the GPT questions are edit-inducing. Our analyses confirm that automated

    long context
  437. arxiv:2609.36615 · cs.LG
    CI-PINN: Causal Integral Physics-Informed Neural Network for Solving Evolution Equations
    Xiaodong Feng, Ziyu Sun, Tao Tang, Xiaoliang Wan +1

    Physics-informed neural networks (PINNs) solve partial differential equations (PDEs) by incorporating governing physical laws into the training loss. For evolution equations, however, their conventional pointwise space--time representation does not explicitly encode temporal dependence, which can hinder accurate prediction. To mitigate this limitation, this work proposes a novel neural architecture termed a causal integral neural network (CinNet). The core module of CinNet is a Volterra-type causal integral term, which aggregates historical features to encode temporal dependence, thereby incor

    benchmark
  438. arxiv:2609.36612 · cs.LG
    Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it
    Hadi Reisizadeh, Jiajun Ruan, Sijia Liu, Mingyi Hong

    Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, dr

    benchmark
  439. arxiv:2609.36611 · cs.AI
    Multi-Channel Mitigation of Source-Trust Shortcuts in Fact-Checking RL Agents
    Jianchang Su, Yiwei Yang, Wei Zhang

    Retrieval-augmented fact-checkers often receive a reliability label, such as HIGH or LOW trust, for each evidence source. These labels should adjust the model's confidence and its decision to search for more evidence, while the verdict should follow the evidence content. We introduce TrustSwap, a counterfactual test that swaps, lowers, or removes source labels while keeping every evidence text fixed, and measures its three output channels (the verdict, the confidence, and the search decision) separately. Across untrained and RL-trained models at two scales, three datasets, and two prompts, con

    retrieval-augmented
  440. arxiv:2609.36609 · cs.LG
    Best Practices in EEG Analysis: Preprocessing, Modeling, and Machine Learning
    Parsa Razmara, Woojae Jeong, Aditya Kommineni, Raymundo Cassani +2

    Electroencephalography (EEG) analysis requires careful choices in preprocessing, statistical modeling, and machine learning because EEG signals are highly susceptible to artifacts, volume conduction, low signal-to-noise ratio, and substantial inter-subject variability. This chapter provides a practical and methodological guide to modern EEG analysis, spanning EEG preprocessing, artifact removal, filtering, bad-channel detection and interpolation, re-referencing, independent component analysis (ICA), and preprocessing of simultaneous EEG-fMRI recordings. We review major approaches for computati

    benchmark
  441. arxiv:2609.36608 · cs.LG
    Act First, Reason Later: Accelerating On-Policy Distillation for Multi-Turn Agents via Reference-Conditioned Inverse Dynamics
    Zubin Zheng, Jiahao Wu, Shaofeng Zhang, Zhirui Zhang +2

    On-policy distillation (OPD) trains multi-turn language agents with dense teacher supervision on student-generated responses. However, standard think-then-act rollouts require lengthy reasoning before each short action, delaying environment transitions and experience collection. Generating actions directly reduces this delay but can degrade rollout quality. To address this, we propose ActFirst-OPD, an act-first, reason-later training framework that decouples environment interaction from full-response generation. The student infers and executes actions through reference-conditioned inverse dyna

    agentbenchmark
  442. arxiv:2609.36605 · cs.RO
    RoboChrono: A Real Robot Benchmark for Streaming Task Understanding
    Yuzhou Wu, Longteng Fan, Zimeng Li, Yu Wanchan +21

    Understanding ongoing robot manipulation requires models to interpret visual observations in relation to interaction history and task progress. We introduce RoboChrono, a benchmark for streaming task understanding comprising 39 scenarios and 34,713 evaluation instances, constructed from real robot executions and complementary bare-hand human recordings. The benchmark evaluates seven tasks grouped into recognition, alignment, and temporal grounding, covering action understanding and anticipation, visual correspondence, temporal ordering, and action localization. Zero-shot evaluation of 18 visio

    manipulationbenchmark
  443. arxiv:2609.36602 · cs.RO
    OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
    Guillaume Besset, Erwann Carn, Timothée Carecchio, Valentin Tordjman-Levavasseur +4

    Transferring human motion to humanoid robots requires adapting the demonstrated motion to the robot morphology while preserving interactions with the environment. This is particularly challenging for loco-manipulation tasks, where contacts with the ground and manipulated objects must remain consistent despite differences in body proportions. Yet, skeletal motion alone does not fully describe these interactions, and fixing object trajectories limits the adaptation to a new embodiment. In this paper, we introduce OTR ETARGET, a unified approach to jointly retarget robot and multi-object motion f

    manipulationhumanoid
  444. arxiv:2609.36601 · cs.AI
    SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
    Miteto Wei, Xiaohan Wang, Zehao Chen, Jiajun Chai +6

    On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-proba

    benchmark
  445. arxiv:2609.36599 · cs.LG
    Scaling Video Generation for Reasoning: At What Cost?
    Weihang Guo, Xiaoyu Wu, Yifei Wang, Niloofar Mireshghallah +1

    We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scali

    benchmark
  446. arxiv:2609.36598 · cs.CV
    Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation
    Ziying Zhang, Litao Li, Junchao Liao, Tianyi Zeng +3

    A video can exhibit convincing motion and photorealism yet fail immediately when visual text collapses. Unlike generic scene content, visual text is unforgiving in video generation: minor stroke corruption, temporal instability, or editing errors instantly break legibility and realism. Existing benchmarks overlook this challenge by treating text as incidental or using static OCR metrics that ignore temporal dynamics. We introduce VidScribe, a unified diagnostic benchmark spanning four generation regimes: writing from language (T2V), transferring text identity from a reference (R2V), sustaining

    benchmark
  447. arxiv:2609.36596 · cs.RO
    HACo: Learning Haptic Active Compliance for Force-Aware Dexterous Manipulation
    Naisheng Ye, Yinzhe Zhou, Junkai Zhao, Yuhang Lu +4

    Contact-rich dexterous manipulation requires policies that translate physical feedback into motion commands while regulating interaction loads across evolving multi-contact interactions. This requires haptic observations of contact state and action supervision showing how commands should adapt. Existing policies often overlook complementary fingertip tactile and joint-torque feedback, while common action targets either encode excessive loading or omit motion constrained by the object. We introduce HACo, a Haptic Active Compliance policy that learns force-regulating actions directly from haptic

    manipulationdexterousteleoperationtactilebenchmark
  448. arxiv:2609.36595 · cs.RO
    Simple Agentic Memory for Generalist Robot Policies
    Yuyou Zhang, Yunbei Zhang, Miao Li, Janet Wang +3

    Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM), a training-free memory layer for frozen generalist robot policies. From the task instruction, SimpleARM specifies what to monitor; frozen perceptual tools maintain compact typed state online; structured access retrieves that state only when a proposed subgoal depends on history;

    manipulationmemoryagenticbenchmark
  449. arxiv:2609.36593 · cs.AI
    Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise
    Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du +1

    Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execu

    agentichierarchical agent
  450. arxiv:2609.36590 · cs.LG
    SEED: Self-Speculative Decoding via Implicit Encoder-Decoder
    Hankun Lin, Patrick Pynadath, Ruqi Zhang

    Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose self-speculative encoder-de

    benchmark
  451. arxiv:2609.36588 · cs.RO
    Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning
    Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo +11

    We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data

    vision-language-actionvlamanipulationrobotwinfrankamulti-agent
  452. arxiv:2609.36587 · cs.LG
    Learned Reporting Preferences in RLVR Can Conflict with the Current Request
    Yupeng Chang, Wenxuan Zhang, Yuan Wu

    Reinforcement learning with verifiable rewards (RLVR) has become a prominent approach for improving language-model performance on reasoning tasks using automatically checked answers. Yet convention-matched evaluation cannot reveal whether reinforcing one reporting convention reduces adherence to a different request that the initial policy already follows. To test this, we train matched policies under two reporting conventions and evaluate each policy under both current requests, using the same initial policy as a shared reference. We complement this crossed design with controlled interventions

    post-training
  453. arxiv:2609.36582 · cs.RO
    Inferring Soil Friction Angle from Robot Foot-Ground Force Histories: A Bayesian Inverse Approach to Proprioceptive Soil Sensing
    Dawei Xu, Zhijie Wang

    Foot-ground interaction signals recorded by quadruped robots may enable spatially distributed, in situ characterization of soil strength. As a first step, we test whether the internal friction angle $φ$ of cohesionless soil can be identified from the force history of a simplified rotating leg. A two-dimensional continuum model implemented with the material point method, benchmarked against measured rotating-leg force histories, generates the training data, and two Gaussian-process surrogates support Bayesian inversion of the full histories. In matched-model experiments, the framework recovers

    quadrupedworld modelbenchmark
  454. arxiv:2609.36581 · cs.AI
    MemEvo: Automatic Discovery of Streaming Video Memory Mechanisms
    Guohong Liu, Jialei Ye, Shanhui Zhao, Yunxin Liu +1

    Query-agnostic streaming video understanding requires vision-language models to continuously compress an indefinitely growing visual stream into a bounded memory before future queries are known. The performance depends critically on the memory mechanism--what observations to preserve, how to represent and consolidate them, and what information to retrieve when a query eventually arrives. Rather than designing a single memory architecture by hand, we formulate memory design as a search problem over executable memory programs. We introduce a lightweight domain-specific language that expresses me

    memorymemory architecture
  455. arxiv:2609.36580 · cs.AI
    SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
    Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin +11

    LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard

    agentllm agentself-evolving
  456. arxiv:2609.36577 · cs.AI
    Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
    Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang +3

    Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we propose Long-Term Memory-Guided Audio Enhancement (LTM-AE) to improve selective target perception by refining the audio representations of ALLMs without training. LTM-AE extracts representations in hidden states from separate clean reference recordings as long-term memory for each category, guiding

    memory
  457. arxiv:2609.36576 · cs.AI
    Divide and Inject: Can Agents Reconstruct an Indirect Prompt Injection from Fragments?
    Michael Lee, Zhipeng Wei, Yue Dong, N. Benjamin Erichson

    Agentic systems are now being widely used to orchestrate tools and reason over long contexts. However, the improving capabilities of the large language models powering these agents also create new attack surfaces for indirect prompt injection. In particular, an attacker may not need to place a complete malicious instruction in retrieved content if the agent can reconstruct the objective from incomplete fragments distributed across a long context. In this work, we introduce adaptive long-context prompt injection (AdaLCPI), which combines long-context fragmentation with adaptive search. AdaLCPI

    long-contextlong contextagentagentic
  458. arxiv:2609.36575 · cs.RO
    EquivDP3: A SIM(3)-Invariant Point-Cloud Encoder for Data-Efficient Humanoid Loco-Manipulation
    Abu Hanif Muhammad Syarubany, Chang D. Yoo

    Visuomotor policies for humanoid loco-manipulation must generalize across object poses and lighting from only a handful of demonstrations. 3D Diffusion Policy (DP3) conditions a diffusion-based action generator on point-cloud features, but its PointNet-style encoder has no built-in equivariance to the rotations, translations, and scalings (SIM(3)) that manipulation tasks respect. EquiBot closed this gap for wheeled manipulators with a SIM(3)-equivariant Vector Neuron Network (VNN) encoder. We extend this to a substantially more complex embodiment, the 43-joint Unitree G1 humanoid, and propose

    manipulationhumanoiddiffusion policymanipulatorbenchmark
  459. arxiv:2609.36572 · cs.AI
    Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning
    Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun +1

    Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We i

    benchmark
  460. arxiv:2609.36570 · cs.AI
    CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
    Mark Russinovich

    Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxi

    manipulationagentllm agentbenchmark
  461. arxiv:2609.36562 · cs.LG
    ThinkingGuard: Decoding Implicit Hazards via Step-by-Step Risk Attribution in Multimodal Large Language Models
    Ruochen Zhang, Yao Huang, Yitong Sun, Jiahe Xie +4

    While Multimodal Large Language Models (MLLMs) are increasingly deployed in safety-critical domains, their reliability is threatened by multimodal implicit risks. Unlike explicit threats, these hazards emerge when individually benign text and neutral visual entities logically converge to induce unsafe outputs. Current detection methods fail to address this because they overlook the underlying risk activation mechanisms that govern cross-modal risk activation, leading to single-modality shortcut learning and hallucinated rationalizations. To bridge this gap, we first construct TriggerBench, the

    benchmark
  462. arxiv:2609.36560 · cs.CV
    FM-ReID: Selective Competitive Token Routing for Object Re-Identification
    Zhiqi Li, Xiaowei Zhou, Zeyuan Sun, Feng Gao +1

    Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, a

    benchmark
  463. arxiv:2609.36559 · cs.LG
    HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation
    Yansong Liu, Rui Liu, Yuan Zuo, Hongwei Zhao +5

    Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over l

    memoryknowledge graphbenchmark
  464. arxiv:2609.36558 · cs.RO
    A robust single-sensing-element tactile sensor for concurrent pressure and tackiness detection with real-time signal decoupling capability
    Ying Yang, Mingwei Gu, Jia-Sen Xie, Xingyu Ma +7

    Integrating tackiness sensation into the artificial skin of humanoid robots significantly enhances their cognitive and operational capabilities. However existing tactile sensors face challenges in decoupling of the multimodal signal and stability. Here we present a surface-soft tactile sensor that incorporates a Hall effect sensor and a soft magnetic composite within a robust elastic framework. The sensor surface indents under pressure and bulges prominently when retracted from sticky surfaces dynamically altering the Hall sensor-magnet distance. This generates whole-process-traceable and base

    humanoidtactile
  465. arxiv:2609.36556 · cs.AI
    MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems
    Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira +3

    Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for divers

    multi-agentagent systembenchmark
  466. arxiv:2609.36553 · cs.RO
    Learning to Explore Hidden Kinematics for Articulated Object Manipulation
    Ruiyao Liu, Boshu Lei, Zhuoyang Pan, Kostas Daniilidis

    The kinematics of an articulated object is often ambiguous from vision alone. Interaction resolves the ambiguity, and active perception methods exploit this by searching for the single action that most sharpens a belief over the kinematic parameters at each step. Such greedy search cannot be extended over a horizon without forward models of the contact and inertial dynamics, which are themselves unknown. We instead amortize action selection into training. We maintain a belief distribution over joint type and parameters, initialized from a generative prior and updated by Bayesian filtering on t

    manipulationbenchmark
  467. arxiv:2609.36550 · cs.CL
    Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment
    Josepha Michiko Leo, Hyun-seok Min, Yehoon Jang, Irvan Zidny +2

    Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tested where "correct" has a definable meaning. Patent claim amendment supplies that signal: the examiner names the attacked limitation and cites prior art, providing per-case ground truth. We release three artifacts: (i) a corpus of 7,385 USPTO prosecution cases with XML-aligned pre/post claims, rejection, and cited prior art; (ii) a seven-probe battery comparing random and structural-match retrieval as two policies under a fixed prompt scaffold;

    retrieval-augmented
  468. arxiv:2609.36546 · cs.LG
    Interactive-Policy Distillation with Bidirectional Propose-and-Verify
    Shutong Wu, Xiwen Chen, Brendan Rappazzo, Daiheng Zhang +3

    On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as propo

    benchmark
  469. arxiv:2609.36540 · cs.RO
    Reactive Real-Time Flow Policies via Asynchronous Distribution Alignment
    Moritz Zoellner, Reece O'Mahoney, Ioannis Havoutis, Rohan Paleja

    Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit t

    vision-language-actionvlalibero
  470. arxiv:2609.36531 · cs.CV
    Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models
    Estela Monserrat Arriaga Santana, Julian Rosas Scull, Ehécatl Sacamch'en Núñez Rico, Hugo Jair Escalante

    Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three

    world model
  471. arxiv:2609.36530 · cs.RO
    Trajectory-Level Mode Guidance for Controllable Diffusion-Based Multi-Robot Motion Planning
    Tianyou Yu, Shengze Cai, Chao Xu

    Motion planning often admits multiple feasible solutions, making multimodal generation valuable, particularly for flexible multi-robot coordination. Diffusion models naturally learn such trajectory distributions, yet incorporating coarse and partial trajectory priors without restricting generation remains challenging. Such priors indicate a desirable region of the solution space rather than a single solution, motivating conditioned generation that preserves multimodality. In this paper, we guide trajectory generation in the clean trajectory space and progressively incorporate trajectory priors

    multi-agent
  472. arxiv:2609.36529 · cs.LG
    Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
    Oliver Sieberling, Bharat Runwal, David Jin, Ryan Chin +2

    Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propos

    memorylong-context
  473. arxiv:2609.36526 · cs.LG
    Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations
    Guanghui Min, Liang Wu, Mingjia Shi, Yinhan He +3

    Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that

    context compressionagentbenchmark
  474. arxiv:2609.36525 · cs.AI
    Reliability Testing of Medical Model Performance under Distributed Deployment
    Yifei Wang, Xiaohan Zhang, Youtao Ding, Tianlin Li +3

    Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testin

    benchmark
  475. arxiv:2609.36521 · cs.LG
    PDE-OBS: Controlled Evaluation Across Observation Patterns
    Ruichen Xu, Siyao Wang, Fang Wan, Jiacheng Qiu +6

    Physical-field reconstruction and forecasting depend on both measurement density and spatial layout, yet evaluation under a single observation pattern does not characterize performance when that pattern changes. We introduce PDE-OBS, an integrated benchmarking platform spanning numerical data generation, model training, and inference and evaluation under varying observation conditions. It combines 560,000 fields and trajectories from seven partial differential equation families with configurable observation operators and seven adapted baseline methods for stationary reconstruction and short-ho

    benchmarkevaluation protocol
  476. arxiv:2609.36518 · cs.RO
    LIBERO-MAX: Do Robot Policies Adapt When the World Changes?
    Yunbei Zhang, Zijian Jin, Yuanzhe Liu, Janet Wang +13

    Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequ

    liberobenchmark
  477. arxiv:2609.36515 · cs.AI
    Large-scale factor analysis shows machine intelligence is only partially interpretable
    Faiz Ghifari Haznitrama, Afrizal Hasbi Azizy, Faeyza Rishad Ardi

    A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyze

    benchmark
  478. arxiv:2609.36505 · cs.LG
    BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning
    Quan Xiao, Mingda Liu, Gaowen Liu, Katsuki Fujisawa +1

    Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning

    retrieval-augmentedagenticbenchmark
  479. arxiv:2609.36502 · cs.AI
    Towards Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions
    Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Ashish Gurung, Ishan Miglani +4

    Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new affordances. An example of this is remote tutoring programs, where human tutors support students who use learning systems while video conferencing. Toward better platform-general modeling of learning, we introduce an AI-driven multimodal transcription system that processes screen-recording videos into unified screenplay-style transcripts c

    online learning
  480. arxiv:2609.36500 · cs.AI
    InterBias-SV: Compound Conditions in Speaker Verification
    Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal

    Speaker verification systems encounter combinations of noise, channel distortion, and changes in speech. Evaluating each condition separately does not establish whether their effects add. InterBias-SV organises this question around a four-term comparison: joint error, two marginal errors, and a common reference. Its results artefact contains 4,068 scored records across 17 experiments, 12 encoder labels, and six speech corpora, totalling 12 million trial evaluations. Three experiment families contain the same-corpus terms needed to compute additive contrasts. For labels assigned to speaker-trai

    benchmark
  481. arxiv:2609.36495 · cs.RO
    Closed-Form Cartesian Forward Kinetostatics for Spatial Multi-Segment Tendon-Driven Continuum Robots
    Ke Wu, Fangju Yang, Xiaohui Zhang, Junda He +3

    Forward kinetostatics of spatial tendon-driven continuum robots typically requires a nonlinear equilibrium solve for each actuation input. This paper develops a force-to-Cartesian-configuration model with a closed-form solution in quadratures for spatial multi-segment robots under tendon actuation. The Cartesian backbone centerline and accumulated material twist serve as generalized coordinates, from which the strain measures and tendon geometry are derived. Variational equilibrium yields explicit axial and bending relations and establishes zero equilibrium material twist within the proposed m

    benchmark
  482. arxiv:2609.36492 · cs.LG
    Benchmarking Vision-Language Models on Synapse Detection and Proofreading in Connectomics
    Yicong Li, Junjie Wang, Leander Lauenburg, Ella Hugie +4

    We benchmarked vision-language models (VLMs) on the decisions annotators take when inspecting electron microscopy images in connectomics: synapse detection (presence and polarity) and proofreading (split errors and merge errors). For synapse detection, we evaluated 19 open and 2 closed models across various architectures and sizes under zero-shot, four-shot in-context learning and LoRA settings, against specialist models, on datasets constructed by us using public resources. For proofreading, we evaluated 3 open and 2 closed models on the ConnectomeBench2 dataset, with cross-species transfer f

    benchmark
  483. arxiv:2609.36490 · cs.LG
    LLMs Learn to Evade Latent Monitors from Prior Feedback Alone
    Hugo Lyons Keenan, Christopher Leckie, Sarah Erfani

    Latent space monitors aim to detect undesired behaviors in LLM agents by inspecting an agent's internal activations rather than its outputs. However, interactive monitoring creates a feedback channel where each verdict the monitor delivers leaks information to the model about how its internal states are being evaluated. We ask whether an agent can infer the monitor's decision rule from this feedback and then selectively edit its activations to evade detection. Unlike prior evasion attacks, the model is never explicitly told what the monitor detects. Surprisingly, off-the-shelf models already p

    agentllm agentbenchmark
  484. arxiv:2609.36488 · cs.LG
    AdaptArena: Evaluating Test-Time Personalization of Web Agents
    Dongchan Shin, Xing Han Lù, Jiaqi Deng, Jay Gala +6

    Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent preferences from implicit signals. Despite its importance for deployment, this problem setting is largely underexplored in existing benchmarks. To address this gap, we introduce AdaptArena, a benchmark for evaluating test-time personalization of web agents via implicit preference infere

    llm agentbenchmark
  485. arxiv:2609.36482 · cs.RO
    DQ-MPCC: Dual-Quaternion MPCC for Quadrotor Racing
    Bryan S. Guevara, Luis F. Recalde, Guanrui Li, Tiago Nascimento

    Quadrotor racing demands aggressive attitude and progress control while passing through every gate, and conventional quadrotor MPCC formulations state the prediction model in inertial coordinates and the attitude error in the body frame. We present a Dual-Quaternion Model Predictive Contouring Control (DQ-MPCC) for quadrotor racing in which the pose is a unit dual quaternion and the contouring errors are projected onto the tangent space of the dual quaternion manifold, expressed in the desired body frame: the same rigid-body dynamics as the conventional model, in unified pose-twist coordinates

    sim-to-real
  486. arxiv:2609.36479 · cs.LG
    Quantum Computing for Network Security Classification: Near-Term Classification and Long-Term Memory Efficiency
    Yuqing Li, Poonam Bala Nehru, Yunpeng Zhang, Danindu Gammanpilage +3

    Quantum computing has already been explored in several network-security applications. However, how quantum computing may contribute to network-security classification in both the near term and the longer term has not been systematically discussed. This paper studies this question through two complementary experiments. First, we evaluate near-term quantum-kernel support vector machines (SVMs) on practical network-security classification tasks and compare them with classical SVM baselines on KDD Cup 1999, CICIDS2017, and BoT-IoT. Across these runs, quantum kernels are competitive. They can match

    memory
  487. arxiv:2609.36478 · cs.AI
    Learning to Harvest Without Collapse in a Regenerative Commons: A Lagrangian Framework
    Jose Tupayachi, Xueping Li, Soham Das

    The tragedy of the commons poses a multi-agent safety problem: reward-seeking agents can deplete a shared resource, and cooperation among its users does not itself specify how much must be preserved. We make preservation an explicit requirement by formulating a regenerative commons as a constrained Markov game or a constrained multi-agent MDP with a designer-specified depletion budget. We develop a nonstationary Lagrangian framework that constructs a policy sequence from solutions of unconstrained games or cooperative control problems. Extending earlier time-average constructions, we introduce

    multi-agent
  488. arxiv:2609.36472 · cs.LG
    DisCoMBO: Steering Expert-in-the-Loop Black Box Optimization via Distributional Conformance
    Jonas Seng, Bennet Wittelsbach, Kristian Kersting

    Sequential Model-Based Optimization (SMBO) traditionally relies on Bayesian or ensembling surrogates for uncertainty quantification. While historically treated as fully data-driven, SMBO increasingly integrates external domain expertise to accelerate discovery. To overcome the opaque guidance and diminished integration fidelity of standard acquisition re-weighting, Probabilistic Circuits (PCs) have emerged as a generative surrogate alternative, enabling direct knowledge injection via conditional sampling. However, these generative routines lack the formal exploration-exploitation semantics req

    benchmark
  489. arxiv:2609.36471 · cs.RO
    Staircase Policy: Streaming Inference for World-Action Models with Large Action Chunks
    Guoheng Sun, Chen Chen, Jin Wang, Ang Li +1

    World-Action Models (WAMs) improve robotic manipulation by conditioning action generation on predicted future observations, but future prediction adds further inference overhead to already expensive iterative action generation. Action chunking can amortize this cost over multiple actions, yet performance degrades over long execution horizons because later actions remain conditioned on stale observations. We introduce STAIRCASE POLICY, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggere

    vlamanipulationaction chunkinglibero
  490. arxiv:2609.36461 · cs.AI
    Rethinking Reasoning Paths as Phase-Structured Trajectories
    Zhenghao He, Guangzhi Xiong, Sanchit Sinha, Bohan Liu +2

    Large language models often improve problem-solving performance by generating multi-step reasoning paths, yet how to analyze the hidden states along these paths remains unclear. Existing approaches typically assign each intermediate state the final-answer correctness label and train probes across heterogeneous questions. We argue that this protocol obscures reasoning dynamics in two ways: (1) correctness prediction can exploit question-level variation rather than path quality, and (2) states aligned by absolute step indices may correspond to different functional phases of reasoning. In this wo

    benchmark
  491. arxiv:2609.36457 · cs.CL
    Memory Consolidation Flattens the Temporal Shape of User Facts
    Sugam Panthi, Muhaiminul Yeamin, Siyan Luo, Rabab Abdelfattah

    Long-term memory systems turn conversations into short stored notes. A note can keep a user fact while losing evidence about whether the fact still holds. For example, "I am driving a Peugeot" can become "The user drives a Peugeot," which drops the cue that the activity is ongoing. We call this aspectual flattening and measure it with LAPSE, a benchmark of matched user statements that differ only in temporal form. We find that memory writers flatten aspect selectively. Three writer models flattened the progressive statement but kept its simple-present match in 244 of 381 pairs, never the rever

    memorybenchmark
  492. arxiv:2609.36454 · cs.CV
    DynamicHOI: Coupled Dynamics for Physics-aware HOI Reconstruction
    Wenliang Guo, Zhanbo Huang, Yu Kong

    We study hand-object interaction (HOI) reconstruction from monocular RGB videos, where partial observations can produce visually plausible yet mechanically inconsistent trajectories. Existing methods mainly enforce visual and geometric agreement, leaving the underlying interaction dynamics insufficiently constrained. We propose DynamicHOI, a physics-aware HOI reconstruction framework combining geometry-grounded diffusion refinement with coupled hand-object dynamics. Geometry spatially grounds visual evidence for trajectory refinement, while articulated inverse dynamics and Newton-Euler dynamic

    manipulation
  493. arxiv:2609.36453 · cs.LG
    Channel-Dependent State Space Model for Multivariate Time Series Forecasting
    Yu-Cheng Wu, Fan-Keng Sun, Li-Chun Lu, Duane S. Boning

    Multivariate time series forecasting (MTSF) is critical across many real-world domains. Existing deep learning approaches fall into two paradigms with distinct limitations: channel-independent (CI) methods unconditionally ignore cross-variable dependencies and model only temporal dynamics, while channel-dependent (CD) methods consider both but typically rely on architectural compromises to mitigate overfitting and computational overhead. We therefore propose Chameleon, a specialized CD state space model (SSM) that enables data-dependent, fine-grained interactions across variables while scaling

    memorybenchmark
  494. arxiv:2609.36452 · cs.AI
    Reliable Parallel Decoding in Masked Diffusion Language Models
    Zhenghao He, Bohan Liu, Guangzhi Xiong, Aidong Zhang

    Masked diffusion language models (MDLMs) can generate text efficiently by predicting multiple masked tokens in parallel, but predictions from the same forward pass are not necessarily reliable when committed together. We study when parallel commitment is reliable. Our diagnostics show that confidence alone does not determine a reliable commitment order: confident predictions near the end of the sequence can fix an answer before its supporting computations are established, and downstream predictions become less reliable as the uncertainty of their upstream context grows. At the same time, a sin

    benchmark
  495. arxiv:2609.36440 · cs.CV
    DARE to Mitigate Hallucination: Dual-path Auto-Regressive-aware Editing
    Jae-Ho Lee, Jeong-Eun Lee, Gyeong-Moon Park

    Large vision-language models (LVLMs) have recently achieved remarkable progress across multimodal tasks, yet object hallucination remains a persistent challenge where models generate descriptions inconsistent with the visual input. Recent work mitigates hallucinations through training-free representation editing, typically by constructing hallucination-related directions from teacher-forcing (TF) contrasts between hallucinated and truthful responses. However, LVLMs operate through autoregressive (AR) decoding during generation, raising the question of whether TF-based analysis fully reflects t

    benchmark
  496. arxiv:2609.36438 · cs.RO
    World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving
    Jieyuan Pei, Meiyi Lu, Sining Ang, Yubo Zhao +11

    Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a t

    manipulationworld modelbenchmark
  497. arxiv:2609.36435 · cs.CL
    MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
    Jingxuan Wu, Yuzhe Yang, Yiqiao Huang, Chengzhi Liu +5

    An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acting on it, and the two are usually optimized as if they were. Keeping the information as text makes the reader's input grow with the retained history, while compressing it into a fixed number of latent vectors bounds the interface but is typically trained to reconstruct text or imitate reference answers, both of which are scored on sequences the reader never

    memorylong-context
  498. arxiv:2609.36432 · cs.CV
    Temporal-Aware Fusion for Robust Outdoor LiDAR Localization
    Minghang Zhu, Zhijing Wang, Yuxin Guo, Chen Liu +4

    LiDAR relocalization aims to estimate the global 6-DoF pose of a sensor in the environment. However, existing regression-based approaches often encounter limitations in dynamic or ambiguous scenarios, as they typically prioritize single-frame inference, leaving the potential of spatio-temporal consistency across scans not fully explored. In this paper, we propose a Temporal-aware Localization framework (TempLoc) designed to enhance the robustness of outdoor localization by effectively modeling sequential consistency. Specifically, a Global Coordinate Estimation module is first introduced to pr

    benchmark
  499. arxiv:2609.36416 · cs.RO
    FineART: Fine-grained Annotated Robotic Trajectory Dataset and Vision-Language-Action Model for Bimanual Manipulation
    Jade Choghari, Pepijn Kooijmans, Mansi Agarwal, Yusuf Umut Ciftci +7

    Robots operating in real-world environments must execute complex, multi-step bimanual tasks over long horizons rather than single, isolated actions. Current manipulation datasets struggle to support this capability: although single-arm datasets reach hundreds of thousands of trajectories, they typically provide only one high-level instruction per episode while the rare bimanual effort that does label subtasks annotates only a fraction of its hours. We present FineART, a densely annotated bimanual manipulation dataset of 40,543 episodes, 1,718 hours, and 533,913 subtasks across 151 tasks. We al

    vision-language-actionmanipulation
  500. arxiv:2609.36413 · cs.RO
    One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions
    Bang Du, Yichen Xie, Shuqi Zhao, Yuxin Chen +2

    A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read o

    liberorobotwinworld modelbenchmark
  501. arxiv:2609.36406 · cs.AI
    From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles
    Jiayuan Chen, Botao Yu, Tianyu Liu, Thai-Hoang Pham +2

    Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA predi

    memorymulti-agentagenticagent frameworkbenchmark
  502. arxiv:2609.36398 · cs.AI
    Where Should Physics Enter a Molecular Crystal Generator?
    Haocheng Tang, Junmei Wang, Wengong Jin

    Generative models make molecular crystal structure prediction fast, but their samples still exhibit geometric and packing violations. Physics can be introduced during training, post-training, or inference, yet these choices are rarely compared with the generator and physical signal held fixed. We introduce CrystAF, an all-atom crystal flow-map generation model, and use it with the UMA interatomic potential to systematically study where physics should enter. Post-training learns physical preferences directly into CrystAF, improving molecular validity and crystal packing while leaving sampling u

    post-training
  503. arxiv:2609.36393 · cs.LG
    Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
    Muhang Tian, Sherry Yang

    Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG),

    agenticself-improvement
  504. arxiv:2609.36392 · cs.AI
    ARCagent: An Adaptive Retrieval Calibration Agent for Clinical Question Answering
    Yuyan Chen

    In diseases where clinical guidelines are incomplete, contested, or mutually contradictory, knowledge completeness and dynamic conflict-aware synthesis are two safety-critical properties that standard Retrieval-Augmented Generation systems do not provide. Therefore, we present \sysname, an adaptive retrieval calibration clinical question-answering agent for ME/CFS, a disease where diagnostic frameworks coexist and major guidelines actively contradict each other on treatment. ARCagent contributes three components. First, a 1,706-chunk, 10-source knowledge base with a structured inter-guideline

    retrieval-augmentedagentbenchmarkllm-as-judge
  505. arxiv:2609.36386 · cs.CV
    Stealth Is a Relation, Not a Property: How Event Representations Create Blind Spots for Timing Attacks in Event-Based Perception
    Shoaib Ahmed Dipu, Md. Shaown Miah, Kamrul Hasan, Sayeed Shafayet Chowdhury

    An event camera produces an asynchronous stream, but what is visible in that stream depends on how a downstream consumer, such as a model or detector, processes time. The same timestamp change may leave a coarse temporal representation unchanged while changing the response of a model that preserves finer timing. We characterize this dependence as observer-relative stealth. For recorded event streams, retiming an event within its protected accumulation window leaves the accumulated integer tensor exactly unchanged. We use this exact blind space to construct Null, a gradient-guided timestamp-ret

    event camera
  506. arxiv:2609.36380 · cs.CV
    LEGO-Anything: Coding Agents for 3D Scene Reconstruction
    Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang +6

    A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact v

    agentbenchmark
  507. arxiv:2609.36376 · cs.AI
    Quantization Enables Private Dense Retrieval against Malicious Service Providers
    Louis Tremblay Thibault, Sofiane Azogagh, Marc-Olivier Killijian, Ulrich Aïvodji

    Dense retrieval, the key component of Retrieval Augmented Generation (RAG), retrieves the most relevant documents by comparing dense vector representations of queries and passages from a large corpus. In privacy-sensitive applications, the server observes the query and controls which evidence is returned, creating both confidentiality and integrity risks. We formulate private dense retrieval as providing query privacy and retrieval integrity against a malicious server, and develop a two-round cryptographic protocol that provides both guarantees. Our protocol reduces private and verifiable retr

    retrieval augmentedrag
  508. arxiv:2609.36373 · cs.AI
    Audience-Bound Persistent Memory: Authorization Across the Memory Lifecycle
    Sibo Liu

    A personal language agent that acts for its owner across private and shared conversations can learn a fact from one audience and later place it in the context it assembles for another. We study authorization before context across the whole memory lifecycle. Each memory item carries the audience present when it was recorded; derived items are partitioned by audience, receive the intersection of their sources' audiences, or are suppressed; an audience widens only by an explicit, object-specific grant; and an item enters a model attempt only when every current viewer belongs to one of its authori

    memorypersistent memoryagent
  509. arxiv:2609.36366 · cs.LG
    Cross-attention encoding models reveal dynamic spatiotemporal routing across human higher visual cortex
    Iishaan Inabathini, Margaret M. Henderson

    Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used deep neural network (DNN) models to build stimulus-computable fMRI encoding models that predict single-voxel responses to complex natural videos. However, the majority of video-computable encoding models predict responses using simple linear mappings from model tokens, overlooking the spatiotemporal structure shared by video representations and neural responses. Recent cross-attention encoding models address this limitation for static images, enabl

    v-jepa
  510. arxiv:2609.36365 · cs.AI
    Engineering Simplicity: Simple Mechanism Interfaces Steer LLM Agents
    Kehang Zhu, Anand Shah, David Parkes

    Can interaction formats and textual scaffolds help large language model (LLM) agents make better decisions, and do better decisions come with better explanations? We study these questions in auctions and matching, multi-agent environments with explicit rules and known optimal strategies. These settings let us vary how a decision problem is presented while retaining a benchmark for evaluating behavior. Drawing on human-motivated theories of simplicity, we compare interfaces that elicit a complete bid or ranking with sequential interfaces that make safe choices easier to identify. We then hold t

    llm agentmulti-agentbenchmark
  511. arxiv:2609.36364 · cs.LG
    Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation
    Xiaoyu Wu, Weihang Guo, Yifei Wang, Xinze Feng +2

    Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Conte

    memorylong-context
  512. arxiv:2609.36359 · cs.AI
    Better Nearest Neighbor Graph Indices via (Efficient) LLM-Guided Pruning
    Fangzhou Wu, Haike Xu, Sandeep Silwal

    Graph-based approximate nearest neighbor search (ANNS) is widely used for large-scale semantic search. Its indices are constructed primarily based on geometric relationships among embeddings of an input dataset (e.g., documents or images), rather than explicitly optimizing for semantic relevance. However, when using these indices for downstream query retrieval, performance is evaluated based on the semantic relevance of the retrieved results to the query. This creates a fundamental "geometry-semantic" mismatch between how the indices are constructed and how their retrieval results are evaluate

    benchmark
  513. arxiv:2609.36354 · cs.LG
    Explainability from Training with Applications to TCR-Epitope Prediction
    Jiarui Li, Zixiang Yin, Samuel Landry, Zhengming Ding +1

    Deep learning models have achieved strong performance in artificial intelligence for science, yet their black-box nature limits our understanding of how they learn scientific tasks. Existing methods for interpretability provide limited insight into how models organize evidence and evolve during learning. We introduce explainability from training (EFT), a model-agnostic paradigm that traces model interpretation during training to explain why models rely on specific features and how they organize these features as predictive evidence. We apply EFT to four state-of-the-art T cell receptor (TCR)-e

    benchmark
  514. arxiv:2609.36352 · cs.RO
    StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
    Ziyi Yin, Sangmin Woo, Kang Zhou, Sungyeon Kim +3

    Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate sup

    vision-language-actionvlamanipulationgr00tliberopost-training
  515. arxiv:2609.36350 · physics.optics
    Asymmetric imaging: Reciprocity bounds and manifestations
    Romil Audhkhasi, Anna Wirth-Singh, Rose Johnson, Maksym Zhelyeznyakov +5

    Recent demonstrations of unidirectional imaging using diffractive multilayers have brought a concept akin to optical isolation to imaging, yet the fundamental limits of such asymmetry have remained unknown. Here we develop a general theory of asymmetric imaging in linear, passive, reciprocal optical systems and derive a fundamental bound on the achievable unidirectionality. Reciprocity requires the forward and reverse transmission operators to have identical singular values, a wave-optical analogue of étendue equality, with a simple consequence: perfect asymmetry is attainable only for sets of

    benchmark
  516. arxiv:2609.36344 · cs.CL
    DeepRewind: Predicting and Repairing Premature Commitments in Deep Research Agents
    Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Lele Wang, Peter West +1

    Deep-research agents conduct long-horizon investigations through iterative search, evidence evaluation, belief revision, and synthesis. However, they may commit to claims before sufficient evidence is available, causing later reasoning to reinforce an incorrect interpretation. We introduce DeepRewind, an additive control layer for reversible deep research that represents the agent's evolving epistemic state as a typed graph of sources, evidence, claims, hypotheses, assumptions, commitments, plans, and drafts. Before accepting an intermediate conclusion, a prompt-based world model predicts its

    world model
  517. arxiv:2609.36340 · cs.AI
    ThuRunel: Dynamic Decoupling for Structured Advisory Dialogue
    Yuyan Chen

    High-stakes advisory domains such as medical aesthetics, legal consultation, and educational planning exhibit a two-phase structure. The early phase requires empathetic elicitation and emotional support, and the late phase requires authoritative specialist judgment. Neither fully automated agents nor human junior consultants adequately address this structure at scale. We formalize the core design challenge as dynamic decoupling, asking how an AI advisory agent should decide what to ask, when to stop, what to resolve autonomously, and what to forward to the specialist. We present ThuRunel, an a

    agent
  518. arxiv:2609.36333 · cs.RO
    ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
    Ke Fang, Yupu Yao, Lu Cheng

    Latent world models rely on representation geometry for planning, yet regularizing the latent marginal alone does not determine the state-to-state relationships used for action selection. We show that this can cause planning-relevant novelty structure to be weakened as representations are transformed into the final latent used by the planner. We introduce Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution. ATLAS transfers normalized pairwise structure from an informative encoder repr

    world model
  519. arxiv:2609.36326 · cs.AI
    PILLAR: Private Inverted-Index Lexical Lookup for Augmented Retrieval
    Truong Son Nguyen, Daniel Blackley, Ni Trieu, Evgenios M. Kornaropoulos

    Retrieval-augmented generation (RAG) hands the user's query to whoever hosts the corpus. We propose PILLAR, a Privacy-Preserving RAG (PPRAG) system based on Private Information Retrieval (PIR) in which a client utilizes the k documents most similar to their query from a server-held and publicly known corpus to respond to their query, while the server learns nothing about the query, either its terms or its access pattern. Prior PPRAG constructions rely on dense retrieval alone, translating approximate nearest-neighbor search into many query-dependent rounds of PIR, and pay for it in both latenc

    retrieval-augmentedrag
  520. arxiv:2609.36323 · cs.AI
    Towards an AI Software Factory for Data Systems
    Anna Pavlenko, Bogdan Crivat, Brandon Haynes, Carlo Curino +22

    AI-assisted coding tools deliver significant acceleration of coding, but only limited impact across the end-to-end software development lifecycle (SDLC)--an Amdahl's law effect! In this paper, we discuss our progress towards building an AI SW Factory that accelerates all the stages of SDLC-Targeting, Coding, Reviewing, and Ops. The AI SW Factory produces a metadata exhaust that enables self-improvement by fine-tuning model weights and updating our World Model (a rich data substrate). We focus on Data Systems and the important class of Evolutionary Coding Tasks (i.e., those with a measurable ob

    world modelagenticself-improvement
  521. arxiv:2609.36322 · cs.LG
    Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
    Xingyu Zhu, Pu, Yi, Ziheng Cheng +5

    Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. We uncover a systematic asymmetry in models using such compression: the same information can be easy to retrieve at one phase and difficult at another. We call this periodic variation in retrieval performance phase sensitivity. In large open-weight models with such compression,

    memorylong-contextbenchmark
  522. arxiv:2609.36319 · cs.AI
    StateTape: Action-Conditioned Evidence Lifecycle Modeling for Long-Horizon Coding Agents
    Ziyang Yu, Liang Zhao, Bowen Zhu, Hasibul Haque

    Despite the recent success of coding agents built on large language models, it remains challenging to run them over long horizons, since every observation is appended to the context and the context grows with each one. History-based maintenance is a common remedy, which masks or summarizes old observations, or prunes what a model reads as useless, and bounds the context at little cost. However, it decides from the text of the history alone and sees nothing of how the code is connected. Since a coding agent edits code many times over a single task, and each write can change what code elsewhere

    action-conditionedagentbenchmark
  523. arxiv:2609.36316 · cs.LG
    Training LLMs to Verbalize Evaluation Awareness
    Usman Anwar, Sahar Abdelnabi, David Krueger

    Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model's spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained wi

    agentic
  524. arxiv:2609.36315 · cs.LG
    PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States
    Arya Kondur, Giosue Migliorini, Cameron Schmitt, Francesco Immorlano +13

    Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requires integrating meteorological conditions, fuels, vegetation, and topography at spatial and temporal resolutions suitable for both physical simulation and data-driven approaches. However, existing datasets often lack the resolution and coverage needed to capture these interacting controls. The PyroStack dataset addresses this gap by provi

    benchmark
  525. arxiv:2609.36314 · cs.LG
    Fractional State Space Transition for Long Sequence Modeling
    Ivan Kobyzev, Abbas Ghaddar, Ali Nasiri-Sarvi, Lifeng Shang +1

    State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum

    memorylong-context
  526. arxiv:2609.36308 · cs.AI
    CheatBench: Measuring Reward Gaming in AI Agents
    Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika +9

    Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowl

    ai agentbenchmark
  527. arxiv:2609.36305 · cs.RO
    Bilinear World Models: Learning Representations with Structured Dynamics for Efficient Control
    Antonio Pariente, Ignacio Boero, Nikolai Matni, Alejandro Ribeiro

    World models jointly learn latent representations and dynamics that predict how high-dimensional observations evolve under actions. In this work, we propose a JEPA-style world model in which, rather than learning arbitrary latent dynamics, we restrict them to follow a bilinear parameterization. This structure enables efficient planning and control while shifting the modeling burden onto the encoder, encouraging richer representations that expose the controllable geometry of the system. In particular, this structured parameterization allows us to structurally enforce action recoverability, ther

    world modellatent dynamics
  528. arxiv:2609.36303 · cs.LG
    HeurEvo: Agentic Evolution of Hybrid Solver-Augmented Heuristics for Time-Critical Mathematical Optimization
    Feijie Wu, Hugo Barbalho, Konstantina Mellou, Marco Molinaro +4

    Recent advances in agentic heuristic design use AI agents and execution feedback to automate algorithm discovery for challenging optimization problems. In many practical settings, high-quality solutions must be obtained under strict runtime constraints, motivating hybrid approaches that combine problem-specific heuristics with powerful mathematical programming solvers. However, existing approaches typically improve heuristic components within predefined procedures or tune solver configurations in isolation. This limits holistic adaptation of where to allocate computation, how to leverage solve

    agentai agentagenticbenchmark
  529. arxiv:2609.36301 · cs.LG
    MoRE: Scaling mixture of experts with hardware-aware low-rank routing
    Honam Wong, Surbhi Goel, Enric Boix-Adserà

    Mixture-of-Experts (MoE) layers are central to frontier language models, and recent architectures push toward more and smaller experts. In this regime, the standard linear router becomes a bottleneck: with $M$ experts and hidden dimension $h$, its per-token cost $Θ(Mh)$ dominates the MoE layer once $M$ is large. We introduce MoRE (Mixture of Rank-reduced-routed Experts), which factorizes the router weight matrix at rank $r$ and reduces the routing cost to $O((h + M)r)$. We prove that rank logarithmic in $M$ suffices for routing expressivity when the number of active experts is fixed, and is ne

    memorybenchmark
  530. arxiv:2609.36292 · cs.MA
    Fully Decentralized and Safety-Aware Multi-Agent Reinforcement Learning for Control on Networks
    Theodore Rogalski, Shirantha Welikala

    This paper develops a safe and fully decentralized multi-agent reinforcement learning (MARL) algorithm to solve a class of discrete-time control problems on networks, including the persistent monitoring problem. Fully decentralized control of agents, while offering numerous benefits, faces issues such as exponentially increasing sample complexity, lack of global information about the system, and challenges in coordinating between agents. To address these issues, this paper introduces a fully decentralized multi-agent reinforcement learning algorithm that integrates deep reinforcement learning

    multi-agent
  531. arxiv:2609.36281 · cs.LG
    GNA: Granular Neighbor Assembly for Retrieval-Augmented Multivariate Time-Series Forecasting
    Vincent Uhse

    Deep forecasters predict from a fixed-length lookback window, and lengthening it gives diminishing returns at a growing cost. Retrieval augmentation instead shows the model how similar past situations continued. Retrieving a whole past window gives every variate the continuation of the same past moment. In multivariate series, however, the best past match differs from variate to variate. We present GNA (Granular Neighbor Assembly), a retrieval layer for forecasting backbones that assembles neighbors at two granularities: whole past windows, which keep the variates coherent, and per-variate nei

    retrieval-augmentedbenchmark
  532. arxiv:2609.36278 · cs.AI
    Illusory Truth or Mere Exposure? Model-Dependent Repetition Effects in LLM-Based Social Media Simulations
    Azza Bouleimen, Nicolò Pagan, Anikó Hannák

    Generative agent-based models (GABMs) are increasingly used to simulate social media dynamics, including misinformation spread. For such social simulations to be valid proxies of human behavior, LLM agents should replicate established human cognitive biases, among them the Illusory Truth Effect (ITE), where repeated exposure to a claim increases its perceived truth value. We investigate whether and how the ITE manifests across four LLMs (Gemma-3-4b-it, Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and GPT-5-nano) in a social media simulation context. We propose a two-phase within-context experim

    manipulationllm agent
  533. arxiv:2609.36264 · cs.AI
    OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models
    Liner Xiang, Wenbo Zhang, Hengrui Cai

    Reliable evaluation of large language models (LLMs) is essential for their development and deployment, yet is often costly, risky, and difficult to perform safely online. We study off-policy evaluation for LLMs, where limited human-labeled data from a behavior model are used to evaluate a newer target LLM. This setting is challenging because labels are scarce, behavior--target distribution shift is common, and response likelihoods are often unavailable for black-box LLMs. We propose the Optimal Transport-based Robust Off-Policy Evaluation (OTROPE), a likelihood-free evaluation that performs di

    evaluatorpolicy evaluation
  534. arxiv:2609.36262 · cs.LG
    Understanding LLM Parameter Update Sparsity through the Lens of Fisher
    Yufan Zhang, Sagnik Mukherjee, Hao Peng

    Recent studies have observed that parameter changes during language-model post-training can be concentrated in a small subset of coordinates. This phenomenon has been reported in reinforcement learning, on-policy distillation, and supervised fine-tuning on near-policy data. Its recurrence across different post-training paradigms suggests shared structure in training dynamics. In this paper, we examine this pattern through the diagonal model Fisher, which measures the sensitivity of the model's output distribution to individual parameters and is independent of any particular reward or teacher s

    post-training
  535. arxiv:2609.36259 · cs.LG
    CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory
    Yixiao Chen, Shuojin Yang, Shi-Min Hu

    Linear RNNs offer linear-time sequence processing and constant-memory decoding, but their fixed-size recurrent states must accommodate all past key--value associations. Existing forgetting mechanisms and Delta Rule updates reduce interference by selectively clearing or correcting the state, yet earlier associations can still become difficult to retrieve. We introduce CyFA (Cyclic Flow Attention), a Linear RNN with relative-time-partitioned memory. At each step, a learned clock controls the cyclic transport applied jointly to the key and value states before the current key--value pair enters th

    memory
  536. arxiv:2609.36254 · cs.AI
    Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
    Xiangyu Zhou, Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov +1

    Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final

    benchmark
  537. arxiv:2609.36253 · cs.AI
    Population Fidelity: Evaluating Population Representativeness in LLMs
    Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu +4

    Large language models (LLMs) show considerable potential in simulating human attitudes and preferences. Prior work finds that LLM-generated responses can compress the range of attitudes found within populations and misrepresent particular subgroups in ways that vary across models and topics. We introduce Population Fidelity, an evaluation framework that distinguishes key conditions required for a set of LLM-generated responses to represent a population. It incorporates three dimensions: group-level accuracy, the amount of between-group variation, and the structure of that variation. We demonst

    evaluation framework
  538. arxiv:2609.36250 · cs.RO
    Action Chunking Proximal Policy Optimization with Feedback Correction
    Sanghyun Hahn, Jonghyun Choi

    Action chunking provides temporal abstraction in reinforcement learning by selecting short action sequences instead of individual actions, but many existing approaches face two limitations in high-dimensional robotic control. First, many rely on value functions over action chunks, which can be difficult to learn as action dimensionality and chunk length grow. Second, executing chunks open-loop removes within-chunk feedback, limiting reactivity in contact-rich tasks. We present Action Chunking PPO (ACPPO), a PPO extension that uses a chunked actor while retaining a standard state-value critic,

    manipulationdexterousaction chunking
  539. arxiv:2609.36246 · cs.CL
    Learning from Teacher Continuations at Student States
    Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu +7

    We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution

    agenticpost-training
  540. arxiv:2609.36243 · cs.LG
    Think Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention
    Mingqiu Liang, Dongdong Wang, Siyang Lu, Ting Huang +1

    Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gated rEstoration), a selective restoration framework that first assesses where restoration is needed and then uses this assessment to guide restoration candidate generation and pixel-level selection. Its encoder predicts patch-level repair probabilities from complementary appearanc

    evaluation protocol
  541. arxiv:2609.36241 · cs.RO
    Design and Validation of an Antagonistic Tendon-Driven Dexterous Robotic Hand with Bidirectional Operation
    Chunghyeon Lee, Hyukjun Kwon, Sungeon Kim, Saehyun Moon +1

    Dexterous robotic hands typically reproduce human hand morphology but inherit its one-sided grasping workspace, requiring wrist or arm reorientation to grasp from the opposite side. Existing reversible hands generally rely on non-anthropomorphic, soft, or task-specific finger arrangements, whereas conventional five-digit anthropomorphic hands remain designed primarily for palmar-side grasping. This paper presents an anthropomorphic, human-scale (200 mm length), lightweight (220 g), 3D-printed, 17-DoF robotic hand built on a bidirectional antagonistic tendon-routing mechanism, in which flexion/

    dexterousgrasp
  542. arxiv:2609.36238 · cs.RO
    ChronoSRL: Temporal Geometry for Self-Supervised Reinforcement Learning
    Nico Bohlinger, Jan Peters

    A goal that is close in space can be far away in time. Obstacles, terrain, and the agent's own capabilities determine how long it takes to get there. Yet, critics in contrastive and survival reinforcement learning do not measure the distances in their representation space in units of time. We therefore introduce ChronoSRL, which gives the critic's embeddings an explicit temporal geometry. The distance between state-action and goal embeddings is trained to match the time that the agent takes to reach the goal (goal-reaching time), while goals that were not reached, and goals from other trajecto

    quadrupedsim-to-realagentbenchmark
  543. arxiv:2609.36235 · cs.LG
    MERID: Multimodal Exploration via Recursive Self-Improvement Agents for Major Depression Analysis
    Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan +3

    Major depressive disorder (MDD) severely impacts daily activities and quality of life. Detecting MDD involves multimodal data, such as interview recordings and sensor measurements. This is particularly challenging, as these heterogeneous modalities often demand distinct, customized prediction pipelines. Existing efforts to address this challenge have explored both manually engineered multimodal architectures and agent-assisted pipeline development. Despite their progress, it remains challenging to autonomously revise pipelines based on experimental feedback and carry verified improvements forw

    self-improvementbenchmark
  544. arxiv:2609.36227 · cs.LG
    One-Step Next-Latent Prediction Is Not a World Model
    Shitong Wang, Zhongang Cai, Yuzhou Hong

    Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward i

    world model
  545. arxiv:2609.36224 · cs.CV
    Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models
    Wentao Zhou, Weijie Gan, Jiayun Wang

    Unified multimodal models (UMMs) combine image generation and visual understanding in a shared backbone. Since generation and understanding are inverse tasks, recent studies self-train UMMs by letting the two branches cooperatively supervise each other. We introduce MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains. MATE lets generation and understanding take turns to be challenger and solver. Given an image, the understanding

    self-playpost-trainingbenchmark
  546. arxiv:2609.36222 · cs.LG
    BASE: Batch-Aware Selection of Experts Using Predicted Removal Error for Efficient MoE Decoding
    Ali Abbasi, Justin Shi, Soheil Kolouri

    Large language models are increasingly expensive to serve. In large-scale serving systems, autoregressive decoding is often bottlenecked by transferring model weights from accelerator high-bandwidth memory into on-chip SRAM. Mixture-of-experts (MoE) models reduce computation by activating only a small subset of experts per token, but this sparsity does not translate directly to batched decoding. Different requests select different experts; therefore, the combined active set across many concurrent requests can span a substantial fraction of the expert pool and require significantly more expert

    memory
  547. arxiv:2609.36219 · cs.CV
    LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
    Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen +4

    Perspective taking is a fundamental component of spatial intelligence, requiring models interpret spatial relations from a specified viewpoint, such as that of another entity or an imagined observer. Despite the increasing spatial reasoning capabilities of Vision-Language Models (VLMs), they still struggle with perspective taking, often defaulting to the camera viewpoint when a query requires reasoning from a different perspective. We introduce Learning Reference Coordinate Frames for Perspective Taking (LeRF), a framework that trains VLMs to construct and use explicit reference frames for vie

    benchmark
  548. arxiv:2609.36218 · cs.LG
    CineSubBench: Evaluating LLMs on Long-Form Narrative and Cultural Understanding from Multilingual Movie Subtitles
    Mir Tafseer Nayeem, Susmoy Chakraborty, Davood Rafiei

    Large language models are increasingly evaluated in specialized domains such as law, medicine, software engineering, and cybersecurity, yet film remains comparatively underexplored despite requiring long-form narrative integration, multilingual interpretation, and culturally situated audience judgments. We introduce CineSubBench, a benchmark for evaluating long-context film understanding from multilingual movie subtitles. A subtitle track represents a film as thousands of short, temporally ordered utterances from which models must reconstruct characters, relationships, events, causal progressi

    long-contextbenchmark
  549. arxiv:2609.36202 · cs.LG
    FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
    Darshan Thaker, Lachlan Ewen MacDonald, René Vidal

    Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per deco

    benchmark
  550. arxiv:2609.36201 · cs.AI
    SCOUT: Synergizing Reasoning and Tool-Use for Computer-Use Safety
    Jianxing Chen, Xiao Yu, Shipra Agrawal, Zhou Yu

    Computer-use agents (CUAs), while capable of completing computer tasks in everyday and professional workflows, can cause unintended harm even under benign instructions and environments. However, detecting such harm remains challenging. First, it requires careful, task-specific reasoning: verifiers guided only by general safety criteria often overlook many important but subtle harmful behaviors. Second, it requires active investigation: past trajectory screenshots show what the agent did but not always what actually changed in the environment, so LLM-as-a-judge verifiers that rely on screenshot

    agentagentictool-usebenchmark
  551. arxiv:2609.36199 · cs.LG
    PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
    Vighnesh Subramaniam, Boris Katz, Brian Cheung, Chun-Liang Li +2

    Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion samplin

    benchmark
  552. arxiv:2609.36198 · physics.optics
    Mode-selective acousto-electric modulation of phonons in a silicon photonic platform
    Ruoyu Yuan, Yishu Zhou, Matthew J. Storey, Ryan O. Behunin +8

    Acousto-electric (AE) interactions enable electrical control of acoustic propagation through piezoelectric media. Bringing AE control onto integrated photonic platforms provides a powerful on-chip control mechanism to reconfigure both the acoustic propagation through piezoelectric media. Bringing AE control onto integrated photonic platforms provides a powerful on-chip control mechanism to reconfigure both the acoustic delay line response and the effective photon-phonon interaction by electrically tuning the phonon propagation. Here we report a mode-selective AE modulation effect in a scalable

    silicon photonic
  553. arxiv:2609.36190 · cs.AI
    FigAct: Turning Scientific Figures into Active Canvases for Explanation
    Shishi Xiao, Zichao Wang, Alexa Siu, David H. Laidlaw +1

    Scientific figures are designed to communicate information visually, yet MLLMs typically explain them by translating their visual content back into text. This requires readers to manually map the resulting explanations back to the figure. Inspired by how people present visual information, we introduce FigAct, a framework that transforms static scientific figures into question-conditioned visual presentations by acting directly on their existing graphical elements. Like a human presenter, FigAct generates a sequence of short narrations, grounds each narration in the corresponding visual evidenc

    benchmark
  554. arxiv:2609.36182 · cs.RO
    Test-Time Adaptation of Manipulation Policies Under Actuator Degradation
    Som Sagar, Ransalu Senanayake

    Robot manipulation policies are usually trained under the assumption that a commanded action produces the same motion as it did during training even after hours of operation. Real hardware violates this assumption as the motors gradually heat up, current saturates near contact, voltage sags under load, thus the same policy action can produce a weaker, delayed, or noisier motion. These conditions are already measured by onboard telemetry, such as joint temperature, motor current, and supply voltage, yet this signal is typically used only for logging or safety checks rather than policy adaptatio

    manipulation
  555. arxiv:2609.36178 · cs.AI
    Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning
    Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li +6

    Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment of trajectory-level advantages to all policy tokens fails to distinguish consequential decisions from less relevant ones, obscuring which intermediate decisions contributed to success. We introduce ProVer, a framework that targets potentially pivotal decisions for fine-grained credit assignment in agentic reinforcement learning. Given a rollout group, an agentic judge contrasts successful and failed trajectories to propose a segment potentially res

    agenticjudge model
  556. arxiv:2609.36173 · cs.LG
    Draft in Parallel, Condition Through Depth: Adjacent Causal Injection for Speculative Decoding
    Haohui Zhang, Keyu Chen, Haocheng Sun, Weibo Gu +3

    Parallel speculative drafting generates multiple candidates in one backbone pass, but independent token selection can produce inconsistent continuations that shorten the accepted prefix. Existing methods mostly leave conditional decoding to a lightweight module after the backbone, which limits the flow of predecessor information to successors. Our analysis of DFlash shows that early positions already form recoverable predictions in shallow layers, and that accurate adjacent predecessors help successors more when they enter earlier. We therefore propose DSpine, a drafter with causal conditionin

    benchmark
  557. arxiv:2609.36171 · cs.RO
    SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation
    He Zhu, Lusen Zhao, Kwan Man Cheng, Su Li +1

    Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop controllers, scripted skill sequences, or task-specific programs. We introduce SkillWeaver, an agentic framework that autonomously generates robot experience by exploring over Neural Interaction Skills (NIS): reusable, parameterized, closed-loop policies that expose learned physica

    manipulationteleoperationsim-to-realmemoryagentagentic
  558. arxiv:2609.36167 · cs.AI
    Adversarial Debiasing of Machine Learning Models for Enhanced Network Security against DDoS Attacks
    Aadith Sukumar, Isha Singh, Devershika Mohane, Ankit Mukherjee +3

    Distributed Denial of Service attacks are a growing threat to network infrastructure, and new techniques, including the use of generative AI, make them harder to detect. Traditional detection systems, such as rule based firewalls, often fail to identify these evolving attack patterns. In this study, we propose a new method for detecting DDoS attacks by combining synthetic data generation using Generative Adversarial Networks with a Random Forest classifier. The GAN generated data showed 80.3 percent cosine similarity to real traffic, which helped the model learn underlying traffic patterns mor

    benchmark
  559. arxiv:2609.36161 · cs.AI
    From Dead Code and Static Requirements to Working Engines: Software Revival with Coding Agents
    Tianyu Liu, Dingyuan Dai, Yufan Du, Zhen Yang

    Can coding agents restore software that no longer runs while preserving its underlying methods, and reconstruct industrial software engines from open specifications? Here we introduce ReviveBench, a benchmark with two task families evaluated by hidden verifiers calibrated against native execution environments, established engineering tools, or purpose-built reference implementations. The revival family comprises ten tasks involving dependency incompatibilities, deleted core modules, legacy builds, and a GPU-based foundation model. Every starting workspace fails verification, and the strongest

    benchmark
  560. arxiv:2609.36159 · cs.LG
    Principled Thoughts for Latent Recursive LLM Systems
    Fahd Seddik, Fatemeh Fard

    Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that t

    agentmulti-agentagent systembenchmark
  561. arxiv:2609.36157 · cs.LG
    Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs
    M. Saeid HaghighiFard, Sinem Coleri

    Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. This assumption limits their applicability to practical vehicular environments, where vehicles may perform heterogeneous but related perception tasks with different output spaces. This paper proposes encoder-sharing hierarchical multi-task federated learning (EN-HMTFL), which integrates cluster-based hierarchical federated learning with a globally shared encoder and vehicle-local decoders. EN-HMTFL enables vehicles performing different tasks to colla

    benchmark
  562. arxiv:2609.36151 · cs.RO
    KPI: A Promptable Kernel for Physical Interaction on Humanoids
    Yikai Wang, Honghao Zhu, Xiao Hu, Hao Zhang +4

    Humanoids now walk, balance and reach with remarkable generality: one whole-body tracking policy follows references from a human, or from an end-to-end policy. That generality travels in the trajectory, and a trajectory alone carries limited information about the interaction it should produce: at contact, the executing controller determines how the robot behaves. Single-task policies usually reach hard interactions by optimising trajectory and controller together in simulation; general stacks usually assume a preset or hand-chosen controller. We present KPI, a promptable kernel for physical in

    humanoidagentagentic
  563. arxiv:2609.36145 · cs.CV
    From Sharp Eyes to Expert Mind: Internalizing Expert Knowledge in MLLMs for Tampered Text Detection
    Kaiqing Lin, Songze Li, Shen Chen, Yunfei Guo +7

    Tampered Text Detection (TTD) is essential for safeguarding document authenticity in security-critical workflows. Existing expert models are effective at capturing subtle manipulation traces but often generalize poorly across diverse document domains, while Multimodal Large Language Models (MLLMs) offer stronger semantic understanding and transferability yet remain insensitive to fine-grained forensic artifacts. This complementarity motivates us to investigate how expert forensic perception can be internalized into an MLLM rather than merely accessed through an external module. We identify a f

    manipulationbenchmark
  564. arxiv:2609.36144 · cs.LG
    Learning Continuous Patient Trajectories from Electronic Health Records
    Silas Ruhrberg Estévez, Kara Liu, Christopher Chiu, Benjamin Atta Owusu +3

    Electronic health records provide irregular observations of latent patient states that evolve continuously over time. Recent autoregressive models condition on clinical histories to forecast future events as sequences of discrete observations. Conversely, multi-marginal flow matching provides a continuous-time formulation, but using multiple observations to supervise training paths does not itself give the learned dynamics access to preceding patient history. We introduce EHRFlow, a multi-marginal flow-matching framework that conditions on encoded patient history, thereby allowing future dynam

    benchmark
  565. arxiv:2609.36138 · cs.CL
    When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
    Jiayi Li, Ruizhe Li

    Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collateral damage they inflict. We present SAKIKO, an auditing framework that formalizes representation repair via directional error discovery, router-conditioned intervention, destination-resolved verification, and prospectively frozen statistical licensing. Across seven LLMs on When2Ca

    agentictool-use
  566. arxiv:2609.36136 · cs.CV
    Xiaomi-OCR-0 Technical Report
    Xin Chen, Anan Du, Feng Feng, Pei Fu +6

    Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-centric corpus using an automated data engine that combines expert consensus, render-based verification, and targeted synthesis. Starting from Qwen3.5-0.8B, our progressive training recipe combines Q-Mask-based text anchoring, continued pretraining, and mixed-task reinforcement le

    benchmark
  567. arxiv:2609.36131 · cs.CL
    A Character-Level Neural Approach to Sinhala Sandhi Splitting
    Yasas Ekanayaka, Deshan Sumanathilaka

    Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We present a character-level sequence-to-sequence study based on SandhiLex, using native Sinhala Unicode input and evaluating recurrent encoder-decoder models for affixational and more complex lexicalized, derivational, and etymological Sandhi. The central challenge is the hard subset lexicalized, deriv

    benchmark
  568. arxiv:2609.36130 · cs.AI
    Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term Agents
    Hongjun Liu, Chen Zhao

    Long-running LLM agents compress past interactions into persistent memories that may be reused as premises for later tasks. This creates a distinct derivation problem: whether the memory actually follows from what the interaction history supports. Relevant evidence may be scattered across earlier interactions, while compression can introduce relations or event status that the history never established. A valid memory may therefore appear unsupported because its citations omit relevant evidence, while individually supported facts may be composed into a stronger statement the history never estab

    memorypersistent memoryllm agent
  569. arxiv:2609.36129 · cs.AI
    Accessible, but Not Adopted: Increasing LLM Adoption among First-generation, Low-income (FGLI) College Students beyond Expanding Access
    Hyungsik Kim

    Large language models (LLMs) are increasingly positioned as a force to empower underserved communities, and significant efforts are being made to expand access. Yet, access alone does not equate to meaningful adoption. First, even if a system is accessible, it won't be adopted if users are not willing to adopt it. Second, even if an LLM system is superficially adopted, the heterogeneity of LLM tools means that LLM adoption can be further deepened. Closing this access-adoption gap is critical to ensuring that the full social potential of LLM is not only accessible but fully realised. Drawing on

    agentictool use
  570. arxiv:2609.36121 · cs.AI
    Render Before Reading: Visual Rendering as a Prompt Injection Defense
    Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran +1

    Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are more likely to follow adversarial instruction when they appear as text than when the same instruction is delivered through a non-textual channel (e.g., as an image). We hypothesize that this modality gap arises from text-centric instruction tuning, which teaches models to obey textual instructions while treating other moda

    benchmark
  571. arxiv:2609.36118 · cs.AI
    The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface
    Yuxiang Liu, Lizhi Yang, Fengze Xie, Aaron Ames +1

    Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that

    vision-language-actionvlamanipulationaction headliberoaction-conditioned
  572. arxiv:2609.36116 · cs.LG
    Dyad: Extending Large Language Models with Native Typed Decision-Making
    Yundaichuan Zhan, Weishi Wang, Wenbiao Liu, Daniel Dahlmeier +4

    We study how to build more capable general-purpose agents by extending large language models (LLMs) with native typed decision-making. We introduce Dyad, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM's internal state to yield a distribution over typed actions. By factorizing decision-making into representations of the evolving interaction state and environment-specific action semantics, Dyad introduces an inductive bias for learning reusable re

    post-training
  573. arxiv:2609.36115 · cs.LG
    Koa-action: Fast and Consistent Structured Decision Making with Generative LLMs
    Shenghong Dai, Shiva Kumar Pentyala, Yingchi Liu, Shubham Mehrotra +5

    Industry applications often demand low-latency classification, yet current large language model (LLM) approaches remain poorly suited for latency-critical applications. Existing prompting and constrained decoding produce verbose, multi-token outputs that require expensive token-by-token generation, while encoder-based models achieve faster inference but sacrifice task flexibility. We propose Koa-action, a framework for low-latency atomic actions -- fast, single-step decisions such as classification, semantic endpointing, Boolean checks, and scoring -- formulated as constrained generation with

    benchmark
  574. arxiv:2609.36107 · cs.RO
    Scouting the Dynamics Gap: Test-Time Policy Adaptation via Action-Outcome Feedback
    Yishu Li, Liyuan Geng, Xinyi Mao, Amber Li +1

    While pretrained robotic policies exhibit impressive capabilities in controlled environments, unobserved physical properties and dynamics require these policies to rapidly adapt during deployment. Existing test-time adaptation methods typically rely on sparse scalar rewards, failing to exploit the rich geometric and dynamic feedback from the environment during physical interaction. To address this challenge, we propose SCOUT, a dynamics-aware meta-learning framework that enables manipulation policies to rapidly adapt by continuously revising their internal beliefs about environment dynamics. O

    manipulationsim-to-realagentbenchmark
  575. arxiv:2609.36104 · cs.AI
    An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures
    Blaz Bertalanic, Carolina Fortuna

    Replacing one LLM agent with a collaborating team can raise accuracy, but whether scaling the team helps, and which architecture to scale, is unclear. Sweeping eight agent orchestration architectures across five instruction-tuned 7-9B models, five short-answer benchmarks, and an executable-code benchmark up to 30 calls, we find that the returns to team scaling are sharply task-dependent: from three to thirty calls accuracy rises by up to 17 points on the two arithmetic word-problem benchmarks (GSM8K, GSMHard) but by at most four on ARC, GPQA, and MMLU, for every architecture, a split the usual

    agentllm agentbenchmark
  576. arxiv:2609.36087 · cs.LG
    PHASE: A Physiology-Guided Hierarchical Foundation Model for Intracranial EEG
    Yipeng Zhang, Chenda Duan, Yuanyi Ding, Tianyi Wang +8

    Clinicians and neuroscientists have long analyzed intracranial electroencephalography (iEEG) through directly measurable physiological characteristics, which carry much of the information that downstream tasks depend on. Recent iEEG foundation models learn by reconstructing or predicting their inputs, which leaves the retention of these characteristics implicit. They are also evaluated mainly on cognitive decoding and a narrow clinical task, i.e., seizure detection. On a broad, clinically relevant benchmark such as Omni-iEEG, they remain below task-specific models when used frozen. We introduc

    benchmark
  577. arxiv:2609.36086 · cs.AI
    PADMÉ: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators
    Cheng Chang, Yining Mao, Peng Qi

    Language models are frequently employed to evaluate other language models. An LM evaluator scoring agentic behaviors across multiple criteria is valuable, provided that its decisions align with human judgment. We call the problem of evaluating this alignment Meta-Evaluation. Tackling it directly is difficult: collecting human data is expensive, absolute scoring is hard to align, and using an LM meta-evaluator recurses the question of trustworthiness. We adopt a reformulation of meta-evaluation as a preference judgment problem: rather than comparing human and LM evaluator scores of a trajectory

    agentagenticevaluator
  578. arxiv:2609.36082 · cs.AI
    GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis
    Ethan D. Frakes, Amy Kvien, Rishabh Kundu, Redad Mehdi +7

    We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothet

    benchmark
  579. arxiv:2609.36074 · cs.LG
    GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization
    Peng Xu, Nihar Koganti, Volodymyr Kindratenko, Xiaohui Chen

    Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. Nonnegative low-rank (NLR) matrix factorization for $K$-means is a scalable clustering method, which connects to semidefinite relaxations with optimal average-case exact recovery guarantees. However, a direct GPU implementation of NLR requires multiple large factor-sized buffers and substantial data movements that are essentially memory-bound. In this paper, we introduce GEM-KMeans, a spectrally normalized yet mathematicall

    memory
  580. arxiv:2609.36071 · cs.AI
    LongCat-DeepResearch Technical Report
    Meituan LongCat Team, He Zhu, Yue Xu, Wanli Wu +27

    We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled

    multi-agentpost-trainingbenchmark
  581. arxiv:2609.36066 · cs.RO
    AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search
    Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang +8

    Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizabil

    benchmarkevaluation framework
  582. arxiv:2609.36062 · cs.AI
    SMat-Attention: Structured Long-Context Sequence Modeling
    Emile Anand, Abdullah Ateyeh, Archer Wang, Marin Soljačić

    Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear attention obtains linear-time training and constant-time decoding by compressing history into a fixed-size state. In this work, we ask whether we can connect these regimes through a tunable notion of structure. To this end, we introduce Structured Matrix Attention (SMat-Attention) via a family of causal masks with structured long-range routing whose row supports have VC-dimension $d$. In our construction, $d=1$ recovers the standard causal mask, a

    long-context
  583. arxiv:2609.36059 · cs.AI
    Mnemon: Raw Records, Fast Judgments, Slow Thoughts
    Guangren Wang

    Long-term memory lets an LLM assistant use a history it can no longer reread, and most memory systems build it by rewriting conversations into facts, graphs or typed memories at write time. We argue that the work of memory divides, as thinking does, into two systems. Most of it is fast System 1 work: many small, independent yes/no judgments about records, such as whether a record is needed or no longer current, which a decision model makes by the dozen in a third of a second. Only a little is slow System 2 work: writing a few search queries, naming what the reply needs and composing the answer

    memoryagent
  584. arxiv:2609.36058 · cs.LG
    ABC: Advantage-Based Control Variates for Reinforcement Learning with Verifiable Rewards
    Hsiao-Ru Pan, Florent Draye, Bernhard Schölkopf

    Recent progress in reinforcement learning with verifiable rewards (RLVR) has highlighted the effectiveness of simple critic-free policy-gradient methods such as Group Relative Policy Optimization (GRPO). In contrast, actor-critic methods rely on learned value functions whose approximation error can introduce bias through commonly used advantage estimators such as temporal-difference error. Motivated by this observation, we revisit trajectory-level control variates through an advantage-value formulation, which we call Advantage-Based Control Variates (ABC). This formulation reveals that the cov

    online learning
  585. arxiv:2609.36057 · cs.LG
    Mirror-Score: Calibrated, Inference-only Scoring Exposes the Limits of Sequence-compatibility Ranking in D-peptide Design
    Jiada Li

    D-peptides combine protease resistance with high target specificity, but computational design of D-peptide binders remains immature. Mirror-Peptidizer introduced an in silico mirror-image screening pipeline using target reflection, backbone generation, and ProteinMPNN sequence design, but its raw ProteinMPNN negative log-likelihood (NLL) ranking was not validated against measured affinities, and only 4 of 9 tested MDM2 designs bound detectably. We introduce Mirror-Score, a calibrated, inference-only scoring framework for heterochiral D-peptide/L-protein complexes, and a public benchmark of 31

    benchmark
  586. arxiv:2609.36052 · cs.LG
    PowerZooJax: A JAX-based Power System Benchmark for Reinforcement Learning
    Zhanhua Pan, Xiao Liu, Zhilong Cao, Jianhong Wang +1

    Power system operation is a safety-critical sequential decision-making problem, making it a natural testbed for reinforcement learning (RL). However, existing RL environments for power systems are often narrow in scope and computationally limited by CPU-based simulation workflows, making large-scale evaluation difficult. We introduce PowerZooJax, a JAX-based benchmark suite for RL in power system operation. It provides five constrained Markov decision process tasks spanning generation, transmission, distribution, distributed energy resources, and data center microgrid. By rewriting power flow,

    benchmark
  587. arxiv:2609.36049 · cs.LG
    Improving scalable oversight with co-trained monitors
    Joseph H. Rudoler, Kevin Tan, Benedict Tessler, Timothy Kong +1

    Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. Fo

    online learning
  588. arxiv:2609.36043 · cs.AI
    SAGE: A Statistical Acceptance Gate for Self-Evolving Agents
    Yihao Wang, Linhan Xia, Rui Liu, Zhaofeng Zhang +5

    Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Seco

    tool-useself-evolvingbenchmark
  589. arxiv:2609.36039 · cs.LG
    Multi-Class, Multi-Tier Network Intrusion Detection: A Comprehensive and Reproducible Benchmark
    Yufeng Xin, Bryant Goseland, Mohamed Rahouti

    Machine learning (ML) and deep learning (DL) have dominated Intrusion Detection System (IDS) research in recent years. Unfortunately, many existing studies have produced inflated results and unreliable benchmarks due to critical oversights and mistakes in the ML and DL pipeline, from data collection and labeling to feature engineering and model training and evaluation. CIC-IDS2017 is a standard benchmark for network intrusion detection. Still, many published results on this dataset are difficult to compare due to labeling errors, inconsistent flow extraction, potential leakage, and performance

    benchmark
  590. arxiv:2609.36031 · cs.RO
    SAKI: Skill Assembly and Kinematic Imitation from Human Videos for Long-Horizon Mobile Manipulation
    Yijie Lu, James Zhao, Weiming Zhi

    Learning from human videos offers a promising route to acquiring diverse manipulation skills. Extending this capability beyond tabletop settings to long-horizon mobile manipulation requires adapting and composing demonstrated interactions across changing scenes and robot configurations. We present Skill Assembly and Kinematic Imitation (SAKI), a framework connecting human-video skill acquisition, cross-demonstration assembly and closed-loop whole-body execution. SAKI prepares reusable object-centric skills that preserve task-critical interactions while allowing transfer paths to adapt. Given a

    manipulationgripper
  591. arxiv:2609.36024 · cs.RO
    CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes
    Shuzhao Xie, Lelin Wang, Guying Lin, Zhi Wang +1

    Reconstructing simulation-ready 3D scenes from real-world observations enables robotics, gaming, and immersive applications, yet existing methods largely assume rigid objects. This leaves an important gap for deformables, whose simulation-ready geometry depends on dimensionality (curves, surfaces, or volumes) and whose behavior may require models beyond elasticity. We present CoDimRecon, an agentic framework that reconstructs editable scenes containing rigid, articulated, and deformable objects from multi-view RGB observations. Scene-level geometric priors ground scale and layout, while object

    agentagentic
  592. arxiv:2609.36012 · cs.RO
    In-Context Learning for Robots: Methods and Applications
    Haojian Huang, Zexi Li, Junhao Guo, Yehang Zhang +35

    General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clar

    manipulationmemoryself-improvement
  593. arxiv:2609.35965 · cs.RO
    Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method
    Yunzhe Xu, Zhe Liu

    Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). A verified four-stage crafting pipeline instantiates the task as MA

    memoryagentmulti-agentbenchmark
  594. arxiv:2609.35966 · cs.LG
    SIFARI: Self-Supervised Interferometric Fitting for Astronomical Radio Imaging
    Shunyuan Mao, Andrea Isella, Paris Perdikaris, Li-Ta Lo +1

    Radio-interferometric images are reconstructed from sparsely sampled visibilities, and CLEAN-based imaging can struggle with spatial filtering, complex morphologies, and uncertainty quantification. Alternative methods that fit visibilities directly can address some of these limitations but often require manual choices of image priors and model hyperparameters. We present SIFARI (Self-Supervised Interferometric Fitting for Astronomical Radio Imaging), a self-supervised neural network workflow that represents sky brightness as a continuous function of position and fits measured visibilities with

    benchmark
  595. arxiv:2609.35958 · cs.AI
    Solver Agent: an Agentic AI Framework for Theoretical Physics Computations Applied to F-theory Uplifts of O3-planes and S-folds
    Eliott Morgensztern, Cesar Fierro Cota, Alessandro Mininno

    We introduce Solver Agent, an AI framework based on large language models for calculations and proofs in mathematics and theoretical physics. The solution process is tracked through a persistent ledger that records assumptions, derivations, and computations. A central agent delegates tasks to specialized sub-agents, while independent agents verify both intermediate steps and the final result. This setup improves the traceability, reproducibility, and verification of computer-assisted calculations. Applying Solver Agent, we study global F-theory uplifts of Type IIB orientifolds and their S-fold

    agentagentic
  596. arxiv:2609.35955 · cs.CV
    HEIR: Learning Human-Entity Interactions with Functional Roles
    Di Wen, Wenhao Guo, Yuedong Tan, Yun Huang +14

    Understanding human-entity interactions requires recovering each person-action event's participants, roles, and shared identities. This structure can support embodied agents by clarifying who acts on which entities and how, informing anticipation and coordination in shared environments. Standard HOI metrics score individual links, leaving complete event composition undermeasured. We introduce HEIR (Human-Entity Interactions with Functional Roles), an image benchmark for complete grounded participant-role sets across object, interpersonal, and self-directed interactions. It contains 18,730 imag

    embodiedembodied agentbenchmark
  597. arxiv:2609.35954 · cs.LG
    ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
    Zhiwei Zhang, Huayu Deng, Fei Zhao, Jiayan Fu +3

    Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserve

    agenticpost-trainingbenchmark
  598. arxiv:2609.37500 · cs.LG
    REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse
    Yuxiao Yang, Shangzhe Li, Tianrun Yu, Kaixiang Zhao +2

    On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout efficiency by reusing each student rollout for multi-step learner updates. REVO addresses prefix-level and current-token policy mismatch through stabilized prefix weighting and one-step resampling from the current student, which enables repeated updates without regenerating full trajec

    benchmark
  599. arxiv:2609.35952 · cs.AI
    HEAR: Real Voices, Real Bias: A Large-Scale Human-Recorded, Demographically Diverse Benchmark for Audio Language Models
    Shen Yan, Duc Le, Irina-Elena Veliche

    We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real human audio samples from 843 demographically diverse participants. HEAR enables comprehensive evaluation through Multiple Choice Question Answering (MCQA) and open-ended long-form tasks. To our knowledge, this is the first large-scale voice benchmark grounded entirely in authentic human speech. We evaluate model behavior across both real-time speech-to-speech and speech-to-text architectures. Our results reveal that voice-conditioned bias is a model-

    benchmark
  600. arxiv:2609.35948 · cs.LG
    Intrinsic Associative Memory on Riemannian Manifolds: Curvature, Capacity, and Emergent Modes
    Krishnakumar Balasubramanian, Zhaoyang Shi

    Geometry does more than constrain an associative memory: curvature determines what it remembers and which states it creates. We develop intrinsic dense associative memories on Riemannian manifolds by casting memory as Epanechnikov kernel-density mode seeking. We compare geodesic and volume-corrected energies and show that curvature separates their behavior. We prove that geodesic memory always retains an isolated pattern, while corrected memory obeys a sharp Ricci-curvature threshold: positive curvature can erase memories in high dimensions, while negative curvature reinforces them. We derive

    memory
  601. arxiv:2609.35947 · cs.LG
    FluxLite: Inference-Time Proposal Control for Discrete Diffusion Models
    Yinuo Ren, Haoxuan Chen, Grant M. Rotskoff, Jiequn Han +1

    Many inference-time tasks for pretrained discrete diffusion models and diffusion language models reduce to drawing samples from a tilted version of the pretrained distribution. Feynman-Kac sequential Monte Carlo (SMC) makes this correction exact in principle, but its prescribed weights routinely degenerate when the proposal dynamics are misaligned with the tilt, capping the practical gains from additional particles. We introduce FluxLite, a lightweight, training-free proposal-control framework for discrete diffusion. On the sparse directed graph of pretrained reverse rates, any sparse jump-rat

    benchmark
  602. arxiv:2609.35943 · cs.CV
    HERO: Histology Encoder for Robust Representation in Oncology
    Zhi Li, Eghbal Amidi, Yating Cheng, Tyson Dawson +11

    Foundation models trained on large pathology image corpora now provide strong, transferable representations for computational pathology. Over the past few years a series of such models has been released, each trained on more slides than the last; on standard classification and segmentation benchmarks, the leading models are now separated by small margins. In clinical use, however, the foundation model is applied to images from hospitals, scanners, and staining protocols outside its training data. Encoders generally embed these acquisition factors alongside biological information, which may int

    benchmark
  603. arxiv:2609.35942 · cs.CL
    Question-Specific Knowledge Graphs for Efficient Visual Reasoning
    Ting-Chih Chen, Emile van Krieken, Shujian Yu, Filip Ilievski

    Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and implicit knowledge sufficient to support reasoning, without introducing spurious assumptions. Existing methods that leverage detailed image captions introduce visual details unrelated to the reasoning task, inflating input token counts and increasing computational cost. To address

    knowledge graphpost-trainingbenchmark
  604. arxiv:2609.35938 · cs.LG
    KernelOnet: An Interpretable Neural Operator Based on Kernel Functions
    Yuan Guo, Hanshu Chen, Qiang Xi, Timon Rabczuk +1

    This paper proposes an interpretable neural operator framework, the Kernel Operator Network (KernelOnet), which incorporates kernel functions explicitly into the neural operator architecture, so that the operator structure matches the kernel-expansion form used in boundary-type kernel-expansion methods. Unlike traditional neural operators such as DeepONet, which learn basis functions implicitly through deep networks, KernelOnet replaces the trunk network with explicit kernels and offers three complementary kernels: a data-driven learnable kernel, in which a neural network parameterizes a radia

    benchmark
  605. arxiv:2609.35937 · cs.AI
    PrivacySkills: How Privacy Guidance Shapes Source Selection in LLM Agents
    Lucas Biechy, Cédric Eichler, Héber H. Arcolezi, Nicolas Anciaux

    While prior work has documented privacy failures in LLM agents, it remains unclear how the presentation of privacy guidance influences their choice of information sources. We introduce PrivacySkills, a controlled framework for evaluating how agents choose among acquisition pathways that provide the same task-relevant value: consulting publicly available personal information, accessing confidential sources, or interacting with the user. The evaluation framework comprises 55 synthetic tasks spanning 11 categories of personal information, with 169 associated skills that describe the available acq

    llm agentevaluation framework
  606. arxiv:2609.35936 · cs.AI
    Embodied Semantic Communication for Collective Autonomous Agents: A Tutorial on Representation, Wireless Delivery, and Closed-Loop Coordination
    Yizheng Huang, Wensheng Lin, Lixin Li, Qinghe Du +2

    As autonomous systems and embodied intelligence enter the dynamic physical world, multi-agent collaboration calls for a paradigm shift in communication design. However, existing communication paradigms overlook that agents form action understanding from their own states, environmental observations, and collaboration relations through a process that evolves as a task unfolds. Consequently, reliable bit delivery, general semantic recovery, or single-task utility optimization alone cannot ensure that heterogeneous agents form coordinated actions compatible with their own conditions from shared in

    embodiedworld modelautonomous agentmulti-agent
  607. arxiv:2609.35935 · cs.RO
    Passive-Dynamic-Walking-Inspired Dynamics Guidance for Energy-Efficient Humanoid Locomotion
    Hyeonjin Choi, Joongheon Kim, Daekyum Kim

    Learning energy-efficient humanoid locomotion requires discovering mechanically economical gait coordination, not merely reducing actuator effort. Reinforcement learning promotes efficiency through effort-related reward penalties, which guide the step-to-step mechanics of walking only indirectly. This article proposes a framework inspired by passive dynamic walking (PDW) that temporarily creates slope-equivalent conditions favorable to economical gait discovery and removes all PDW-specific guidance before nominal-dynamics optimization. During early training, a tilted-gravity field assists sagi

    humanoid
  608. arxiv:2609.37494 · cs.LG
    Your Benchmark Is Not Saturated: Reviving Multiple-Choice Evaluation with Answer Pooling
    Mohamed Eltahir, Abobaker Ahmed, Nawaf Barebood, Hussain Bu Subayt +2

    Multiple-choice benchmarks are cheap to grade and are running out of room, and the standard remedy, writing harder items, is slow and repeated for every benchmark. A saturated benchmark still holds a harder task. Each question's wrong options are written for that question alone, so a model can score by eliminating a few options. We propose AnswerPool: take $N$ questions that share a context, pool all their options into one list, and ask the model to assign every question its answer. No item is written and no label changes. The chance of guessing a group right falls from $10^{-3}$ to $5\times10

    benchmark
  609. arxiv:2609.35932 · cs.LG
    Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
    Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang +2

    Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged mark

    agentllm agentagent benchmarkbenchmark
  610. arxiv:2609.35930 · cs.LG
    Bregman Consensus
    Andrei N. Soklakov

    Consider a community of agents who are seeking consensus on a set of parameters. The agents agree to use the same Bregman-type divergence to quantify disagreement between their individual estimates of the parameters but have varying confidence in each other's abilities. Each agent is happy to revise their estimate by moving to the weighted barycenter of all individual estimates with higher weights applied to more trusted agents. We show that such revisions naturally lead to an iterative algorithm which converges to a unique consensus estimate of the parameters. Furthermore, since the consensus

    agent
  611. arxiv:2609.38027 · cs.CL
    Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs
    Junning Shao, Siwei Wang, Zhixuan Fang

    In recent years, the performance of large language models (LLMs) on reasoning tasks has been remarkable, even surpassing human capabilities on various benchmarks. However, there remains a lack of clear understanding in the academic community regarding how the structure and internal parameters of LLMs progressively solve complex reasoning problems. In this study, we investigate the inference process of LLMs on cross-linguistic materials and propose the hypothesis that LLM layers exhibit a structured division of labor across conceptualization, reasoning, and textualization. Based on this hypothe

    benchmark
  612. arxiv:2609.35928 · cs.AI
    Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems
    Xavier Del Giudice, Alessio Palma, Matteo Migliarini, Fabio Galasso +1

    Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantly impairs cooperation. We show that when agents are aware of each other's model family, the group splits into clusters, where agents prefer interacting with others carrying their same label, although nothing in the task rewards or asks for such a split. We argue that the label itself causes this split, which we define as $\textit{factionalism}$. We show and measure this phenomenon in two cooperative games and on a reasoning benchmark, with nine to

    multi-agentbenchmark
  613. arxiv:2609.35926 · cs.LG
    Normative Loss Landscape Navigation: A Trajectory-Based Approach to Mitigating Forgetting in Incremental Learning
    Isabelle Aguilar, Zayn Andre Zainal, Luis Fernando Herbozo Contreras, Zhaojing Huang +1

    Continual learning models suffer from catastrophic forgetting when trained sequentially on non-stationary data distributions. Previously, this has been addressed through weight regularization. While preconditioning gradients offer a promising alternative to mitigate forgetting, current approaches are myopic. Conversely, standard regularization methods apply rigid, scalar Euclidean penalties that entirely ignore the underlying Riemannian geometry of the parameter space. To overcome this gap, we propose TMLN (Trajectory-Modulatory Landscape Navigation), a normative navigation policy that formali

    benchmark
  614. arxiv:2609.35150 · cs.CV
    Toward a Culturally Adapted Chinese Language Agent: A Wizard-of-Oz Study of Nonverbal Behavior in Chinese-German Intercultural Interaction
    Siddhant Jain, Anna Lea Reinwarth, Dimitra Tsovaltzi, Rafael Math +1

    Successful intercultural communication requires more than grammatical competence. It demands sensitivity to culturally embedded social norms whose violation triggers subtle but meaningful nonverbal responses. For German learners of Mandarin Chinese, acquiring this sensitivity is critical yet poorly supported by existing language-learning agents. We present a Wizard-of-Oz (WoZ) study design and supporting real-time system for collecting multimodal behavioral data from native Chinese speakers reacting to social norm violations by German learners. The system features a photorealistic MetaHuman av

    agent
  615. arxiv:2609.35924 · cs.AI
    Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives
    Hua, Xu, Dongxin Li, Gwen Yidou-Weng +3

    Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the number of unresolved positions. We introduce COFFEE, a plug-and-play framework that avoids this enumeration by separating sequence dependence from the

    benchmark
  616. arxiv:2609.35922 · cs.LG
    Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
    Bhavik Mangla

    A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct

    agent
  617. arxiv:2609.37501 · cs.LG
    Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance
    Dipankar Sarkar

    We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation validity, source grounding, schema compliance, escalation correctness, constitutional alignment, and unsafe-action rate. Signals are distinguished by their source of supervision: programmatic verifiers, task-level escalation labels, or AI-judge scores. A deterministic runtime supervisor blocks ungrounded answers and forces escalation, logging interventions. The same domain constitution informs evaluation, training rewards, and serving guardrails. Tas

    agentic
  618. arxiv:2609.35916 · cs.CV
    VehicleArena: A Realistic Urban Environment for Multi-Agent Driving
    Jie Yang, Jiajun Chen, Jiazheng Zhou, Mianqiu Huang +3

    Real-world embodied agents often pursue independent objectives within a shared physical environment, where their actions can alter the conditions faced by others. Existing benchmarks, however, typically assume shared goals or explicitly prescribed interaction protocols, leaving such emergent physical coupling underexplored. We introduce VehicleArena, a 3D urban-driving benchmark for studying independently operating agents in a dynamic shared world. In VehicleArena, LLM-controlled agents must fulfill evolving passenger requests while navigating complex traffic, and each agent's driving decision

    embodiedmulti-agentembodied agentbenchmark
  619. arxiv:2609.37473 · cs.LG
    KT-EGO: A Knowledge Transfer Assisted Efficient Global Optimization Algorithm for Solving High-Dimensional Expensive Black-Box Problems
    Qineng Wang, Liming Song, Yun Chen, Guangjian Ma +2

    Many engineering problems involve optimizing a high-dimensional expensive black-box (HEB) design space. To solve such problems efficiently, we propose a knowledge transfer assisted efficient global optimization (EGO) algorithm, labeled as KT-EGO, which extends the EGO algorithm for solving problems over higher dimensions (i.e., $d>20$). Specifically, the original design space is divided into several low-dimensional subset design spaces. More importantly, in order to extract information from the subset design spaces to accelerate the progress of full optimization, we propose a surrogate-based d

    benchmark
  620. arxiv:2609.35914 · cs.LG
    Agentic Federated Learning: Rule-Based Client and Server Agents for Adaptive Training
    Deepthy K. Bhaskar, VP Binu, B Minimol

    Federated Learning (FL) enables collaborative model training across distributed clients without sharing raw data, making it suitable for privacy-sensitive applications such as healthcare, finance, and edge intelligence. However, conventional FL approaches rely on static client participation and fixed aggregation strategies, which limits their effectiveness under non-IID data distributions, heterogeneous client behavior, and noisy or unreliable updates. To overcome these issuess, this paper proposes an Agentic Federated Learning (AFL) framework that integrates lightweight rule-based autonomous

    agentautonomous agentagentic
  621. arxiv:2609.35912 · cs.AI
    MMSkillRisk: Can Agents Stay Safe When Multimodal Skills Become Traps?
    Lingqi Jiang, Jialuo Chen, Jianan Ma, Xinhao Deng +7

    Agent skills are shareable packages of procedural instructions, tools, and examples. Multimodal skills additionally include visual references that agents retrieve and inspect during execution. Because these images guide actions, attackers can disguise malicious instructions as ordinary visual guidance within otherwise legitimate skills. Existing skill-security research primarily examines text-carried attacks or scanner detection, leaving the runtime effects of image-borne attacks insufficiently evaluated. We introduce MMSkillRisk, to our knowledge the first publicly available benchmark dedicat

    agentbenchmark
  622. arxiv:2609.35911 · cs.LG
    Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents
    Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu +4

    Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validati

    llm agent
  623. arxiv:2609.35910 · cs.LG
    The Decision Value of Perception Compute
    Hoang Pham Cong, Ho Viet Duc Luong

    Adaptive perception spends extra computation on inputs where perception is expected to improve. When perception feeds a downstream decision system, a better perception output need not produce a better decision. We define the decision value of perception compute as the change in downstream loss from escalating an input from a cheap to an expensive perception mode. Because this value can be negative, the allocation of perception compute should be judged against a budget-constrained decision oracle, with uniform full-fidelity inference as a baseline rather than an upper bound. We introduce DEEP (

    benchmark
  624. arxiv:2609.35909 · cs.AI
    Cheap to Hypothesize, Costly to Verify: The Defense Surface of Agentic Vulnerability Discovery
    Kaikai Zhang, Zihan Zhang, Yuchong Xie, Zesen Liu +3

    Autonomous LLM agents turn vulnerability discovery into a repository-scale search: they generate many vulnerability hypotheses but can verify only a subset under a finite budget. We show that autonomous vulnerability discovery exhibits a hypothesis-verification asymmetry, where verifying a candidate hypothesis through reachability analysis, execution, and proof-of-concept construction is substantially more expensive than forming it. Under a finite resource budget, this makes autonomous discovery a resource-bounded selective-verification process, further exposing verification effort as a unique

    agentllm agentagentic
  625. arxiv:2609.37453 · cs.AI
    Commitment Hierarchies under Intent Revision: A Belief-Revision Account of Salvage in Tool-Use Agents
    Spandan Ghose Chowdhury

    When a user changes their mind partway through a task, an agent that has already split the task into sub-goals and paid for tool calls must decide, per cached sub-result, whether to keep, patch, or discard it (salvage), restarting wastes valid work and continuing unchanged answers the old question. Our main finding is that salvage quality is a matter of role design rather than model capability: a language model asked the keep/patch/discard question one node at a time is unreliable, but asked to classify the revision once, with a deterministic layer propagating the decision, it reaches the cost

    agenttool-use
  626. arxiv:2609.35900 · cs.AI
    Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy
    Yuehui Wang, Xinyu Qi, Guirong Xue, Cheng Wang +3

    The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dependencies-data selection, calibration corrections, priors, and domain assumptions-implicit. This ambiguity complicates the evaluation of LLM-based agents, since a failure to reproduce a result may reflect either limitations of the agent or underspecification in the source. We present a framework that evaluates agents through end-to-end re

    agentllm agent
  627. arxiv:2609.35898 · cs.LG
    Wasserstein Causal Forests for Distribution-Valued Outcomes
    Hugo Gobato Souto

    This paper proposes Wasserstein Causal Forests (WCF) for settings in which each unit's outcome is itself a probability distribution. This study also defines finite-grid transformed average and conditional average treatment effects, including a reference-distance contrast that asks whether treatment moves unit-level distributions toward a prespecified benchmark. Simulations cover null effects, location and shape changes, limited overlap, equal-mean but different laws, heterogeneous effects, multimodality, and structural zeros. WCF is most accurate on the conditional-law metric in most reported

    benchmark
  628. arxiv:2609.35897 · cs.LG
    Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?
    Haomin Luo

    The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. While algorithm self-discovery has produced Disco103 that surpassed PPO to achieve SOTA benchmark performance -- its internal update machinery remains

    self-improvingself-improvementself-evolvingbenchmark
  629. arxiv:2609.37470 · cs.LG
    Probability Contracts: Accuracy, Coherence, and Decisions Across LLM Interfaces
    Han Chen, Yingrui Li

    A probability used for a decision should refer to the same event across equivalent requests. We introduce probability contracts, a benchmark connecting exact finite-world posteriors, validated event transformations, and failure-aware decision evaluation. Four model-interface configurations are evaluated on 1,000 worlds. Their assessments differ across accuracy, coherence, and decision loss: Kev has lower aggregate canonical posterior error than Jev, but larger complement and coarsening residuals, with accuracy ordering varying by stratum. Jev's Event and Choice interfaces induce different bina

    benchmark
  630. arxiv:2609.37469 · cs.LG
    Relevance Is Not Sufficient Evidence: Detecting Evidence Gaps Before Generation in RAG
    Suting Chen, Peichun Hua, Yunming Xiao

    Retrieval-augmented generation (RAG) grounds large language models in external sources, but retrieved passages often name the right entities without providing the facts needed to answer. Even when instructed to abstain, 12 generators answer 40.0-99.3% of insufficient-evidence questions. Training generators to abstain ties the decision to model weights, may reward answers recalled from parametric knowledge, and still requires a full generator call. Can sufficiency be judged from the question and evidence alone, before any answer exists? We identify pitfalls in constructing insufficient-evidence

    retrieval-augmentedragbenchmark
  631. arxiv:2609.37465 · cs.CV
    Event-Only Wingbeat Counting under Camera Motion: A Controlled MuJoCo Benchmark
    Zhang Nengbo

    Counting completed wingbeats requires identifying individual cycles, including during frequency changes and pauses; estimating a dominant frequency alone is insufficient. Camera motion further mixes target and background brightness changes in event observations. We present a controlled MuJoCo benchmark that separates motion training from event-only image translation compensation. The acquisition contains 324 streams from 24 independent scenes, three flapping geometries, two distances (1.5 and 3.0 m), and static, moderate-motion and stronger-motion views. Fifteen scenes are used for fitting, th

    benchmark
  632. arxiv:2609.37447 · cs.LG
    Engineering Efficient Self-Play Chess: Search, Replay, and Throughput Under Limited Compute
    Bertil Braun

    How strong can an AlphaZero-style chess system become under limited training compute when its entire learning loop is engineered for efficiency? We train from random initialization through searched self-play on a single eight-GPU node for 2.5 days. The resulting 6.32-million-parameter model reaches 3,251 benchmark Elo [3,206, 3,297] at 100,000 searches per move (estimated at under five seconds of thinking time) against a fixed-node Stockfish 13 ladder. The run ingests 3.25 million completed games, involves an estimated 100 billion search simulations, and makes 836.6 million training presentati

    self-playbenchmark
  633. arxiv:2609.37488 · cs.LG
    FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding
    Taiyo Sato, Takamasa Sanda, Keisuke Maeda, Takahiro Ogawa +2

    Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region suppor

    benchmark
  634. arxiv:2609.37487 · cs.LG
    Physics-Guided Flow-Map Matching for Precipitation Nowcasting
    Shunya Nagashima, Takumi Bannai, Makoto Misaizu, Keisuke Maeda +2

    Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and rare high-intensity structures that matter most. Deterministic models minimize a pixel loss and are driven toward the conditional mean, which blurs exactly those structures, while generative models that add a stochastic residual on top of a deterministic backbone inherit the same blur. We propose Physics-Guided Flow-Map Matching (PG-FMM),

    benchmark
  635. arxiv:2609.37455 · cs.RO
    Surgical Master Console Using General-Purpose Robot Arms and a Separable Articulated Distal Interface: Porcine In-Vivo Evaluation
    Minsung Kim, Seonho Shim, Dongho Yee, Younghoon Noh +8

    High-performance surgical master consoles offer intuitive articulated manipulation but are expensive and difficult to reproduce, while accessible commercial haptic devices lack built-in interfaces for surgical wrist articulation and continuous grasp. We present a laparoscopic master in which general-purpose robot manipulators provide the programmable base and surgical-specific interaction is concentrated in a separable distal adapter. Each 6-DoF arm carries a custom 2-DoF direct-drive wrist, continuous grasp sensing, and clutch-based workspace management; the complete bimanual hardware costs a

    manipulationmanipulatorgrasp
  636. arxiv:2609.35889 · cs.AI
    SINGED: Correct Outputs Do Not Certify Safe Execution in LLM Agents
    Xiaoyu Xu, Zi Liang, Minxin Du, Qipeng Xie +3

    Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the requested output but add an effect forbidden by the task contract. We introduce SINGED (Source Integrity and the Nonidentifiability Gap in Execution Decisions for LLM Agents), a controlled benchmark covering five primary and two held-out task families. It varies displayed rank,

    agentllm agentbenchmark
  637. arxiv:2609.35886 · cs.AI
    Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money
    Ankit Srivastava, Debjyoti Paul

    AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it should, and no check keyed on identity will see it. We present three artefacts for measuring and reducing that loss. First, a taxonomy of agentic commerce fraud that separates five observation levels (agent reasoning, wire, settlement rail, counterparty, principal) from the request-level and history-level evidence available

    agentai agentagenticbenchmark
  638. arxiv:2609.35885 · cs.MA
    Collective Regimes in Multi-Agent LLMs under Reasoning Effort and Communication Topology
    Machiko Hirota, Akshara Nadayanur Sathis Kanna, Ujwal Kumar, Phan Xuan Tan

    Multi-agent LLM systems are increasingly used for deliberation and evaluation, often under the assumption that greater peer interaction leads to more reliable consensus. Existing work largely evaluates these systems through final accuracy or aggregate agreement. However, such measures do not reveal how agreement is organized in the panel. In this paper, we study \(N=50\) stateless LLM agents that update their predictions from locally visible peers, and characterize their behavior using both global and local measurements of agreement. We identify three collective regimes: synchronised, twisted

    llm agentmulti-agent
  639. arxiv:2609.35880 · cs.LG
    From Static Policies to Adaptive Priors in Offline Reinforcement Learning
    Tianwei Ni, Vineet Jain, Akash Karthikeyan, Pierre-Luc Bacon

    Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. We argue that this formulation becomes incomplete when an offline-trained policy is subsequently updated through online interaction, as increasingly occurs in modern intelligent systems through test-time adaptation and online fine-tuning. This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and inste

    self-correction
  640. arxiv:2609.35879 · cs.AI
    CruxBench: A Benchmark of Information Discovery
    Hui Dai, Lina Piao, Nick Merrill, Nadja Flechner +2

    Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates bel

    benchmark
  641. arxiv:2609.37454 · cs.AI
    Governing the Edge: Automating Commercial Property and Casualty Insurance Underwriting via a Hybrid Local-Cloud Multi-Agent Framework
    Vivek Kumar Singh, Gautam Bhowmick

    Underwriters in commercial Property and Casualty (P&C) insurance spend 30 to 40% of their time on administrative work rather than risk judgment, and a single submission takes about 40 minutes by hand. We present Governing the Edge, a multi-agent framework for that layer, organized around a data-residency constraint: sensitive submission data must not leave the perimeter. Eleven agents and two deterministic control nodes, one a human-escalation interrupt, form a 13-node LangGraph workflow across two tiers. Agents touching raw submissions run locally on Gemma 2, each bound through a tool interfa

    agentmulti-agentagent frameworkbenchmark
  642. arxiv:2609.37457 · cs.AI
    VeriWeave Govern: Evidence-Gated Deterministic Runtime Governance for Enterprise AI Agents
    Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya

    Enterprise artificial-intelligence agents increasingly call tools, modify infrastructure, and process protected data, creating a need to separate action generation from action authorization. This article presents VeriWeave Govern, a deterministic runtime governance layer that evaluates structured agent actions against versioned policies, validates typed evidence, applies fixed deny > review > allow precedence, routes consequential actions to accountable human review, and records replayable tamper-evident audit state. GovernBench evaluates the design over 30 independent seeds and 60,000 oracle-

    agentai agent
  643. arxiv:2609.35875 · cs.AI
    Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models
    Leonardo Ferreira, Gardenia Liu, Kaden Zheng

    Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generatio

    multi-agentbenchmark
  644. arxiv:2609.35874 · cs.AI
    Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees
    Yaacov Pariente, Vadim Indelman

    Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planne

    policy evaluation
  645. arxiv:2609.37468 · cs.LG
    Backdoor in the Loop: Compromising Agentic Search via Malicious Retrievers
    Beining Xu, Peichun Hua, Yunming Xiao

    Agentic retrieval-augmented generation (RAG) interleaves reasoning with repeated retrieval, giving the retriever influence over both the evidence an agent observes and its subsequent search decisions. We study retriever backdoors that exploit this feedback loop and repurpose weak backdoor purification to conceal their presence. An attacker supplies a compromised retriever checkpoint while leaving the search agent and deployment corpus unchanged. Without corpus write access, the attacker can still suppress useful evidence, persistently retrieve a selected existing document, or steer the agent t

    retrieval-augmentedragagentagentic
  646. arxiv:2609.37495 · cs.CV
    MotionMaestro: Masked Tokenization for Unified Motion Generation
    Yun Chen, Munchurl Kim, Jeonghyeok Do

    Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different o

    embodied
  647. arxiv:2609.37496 · cs.CV
    GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation
    Jeonghyeok Do, Munchurl Kim

    Paired synthetic aperture radar (SAR) and electro-optical (EO) imagery is increasingly available across sensors, resolutions, and geographic regions. Yet existing SAR-to-EO image translation (SET) methods are typically trained on a single, limited-scale dataset, producing models specialized to particular sensing conditions. We introduce GeoSET, the first generalist model for SET, built around a single pretrained parent that is adapted to downstream datasets under a common protocol. We curate over 3 million high-quality SAR--EO pairs from a collection of more than 10 million SAR observations, s

    benchmark
  648. arxiv:2609.35868 · cs.AI
    Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
    Jinhao Zhang, Zeyu Liu, Zicheng Yan, Yunquan Zhang +2

    Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fl

    benchmark
  649. arxiv:2609.37474 · cs.AI
    Authority Before Utility: Non-Compensatory Control for Persistent LLM Memory
    Wesley Shu

    Persistent memory creates a control problem that retrieval relevance alone does not solve: a memory can remain highly useful after an update, deletion, or revocation makes it inadmissible for the current answer. We formalize this as a separation between utility and authority. A fixed finite penalty applied to an unnormalized utility score cannot guarantee exclusion under arbitrary positive-affine reparameterization of that score; by contrast, rank-normalized compensation is scale-invariant and therefore forms a stronger empirical comparator. Our prospectively frozen TIDE/LongMemEval primary wa

    memorypersistent memory
  650. arxiv:2609.37476 · cs.RO
    Learning Social Navigation from Internet Videos in the Policy State Space
    Jiaming Wang, Duc Thang Nguyen, Jizhuo Chen, Volodymyr Shcherbyna +3

    Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversabili

    benchmarkarena
  651. arxiv:2609.32268 · cs.LG
    Self-Reconstruction Dynamics for Autoencoder Reconstruction Refinement
    Hitoshi Iyatomi

    Standard autoencoder (AE) inference uses a single encoder-decoder pass, though the latent may not be optimal for each sample under a fixed decoder. We ask whether a trained AE can reveal information for improving its own reconstruction. Repeated application of a frozen AE to its reconstruction produces transient image- and latent-space trajectories, termed Self-Reconstruction Dynamics (SRD). Although this degrades fidelity in the AEs studied here, SRD contains sample-specific information for correcting the reconstruction. We propose SRD-guided Reconstruction Refinement (SRD-RR), which predicts

    latent dynamics
  652. arxiv:2609.32267 · cs.LG
    When Can Old Evaluations Certify a New Model? Label-Efficient Release Decisions under Evaluator Drift
    Joyanta Jyoti Mondal, Mridul Banik, Md. Shifatul Ahsan Apurba, Md Masud Al Mahmud

    Releasing a model update requires certifying that its current-population risk stays below a threshold. Trusted labels are expensive, while a cheap evaluator, such as an LLM judge, scores every example. Reusing evaluator errors from earlier audits is tempting, but when may such evidence replace current labels? It depends on the status of history. If the errors can change invisibly, no label-free test detects the change, and every valid, useful certifier must keep buying labels at a rate we characterize; if a bound on the change is assumed, label-free certification is valid at an explicit error

    evaluator
  653. arxiv:2609.32263 · cs.LG
    Representation Editing for Multimodal Test-Time Adaptation
    Longfei Huang, Xiangyu Wu, Yang Yang

    Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in

    benchmark
  654. arxiv:2609.32259 · cs.LG
    Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
    Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo +2

    Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires each receiver to prefill shared context already processed by the sender. Reusing the sender's key-value (KV) cache avoids this redundancy, but prefill-free transfer across model families must handle differences in tokenization, model depth, and KV representations. To address these issues, we propose \textit{HeteroFold}, a prefill-free cross-family KV cache transfer method that keeps both the sender and receiver frozen. HeteroFold aligns model structur

    long-contextagentmulti-agentagent benchmarkbenchmark
  655. arxiv:2609.32256 · cs.AI
    LAM: Efficient Lossy Agent Memory Framework With A Retrieval-Score Error Bound
    Baixi Sun, Le Chen, Anjir Ahmed Chowdhury, Xiaolong Ma +9

    Agent memory grows as agents read inputs, reason, and call tools. Longer histories increase inference cost and eventually exceed the context window. LLM-based summarization reduces this history but adds latency and provides no explicit bound on information loss. We propose LAM, a Lossy Agent Memory system with three components: a deterministic deduplication rule with a substitution bound on retrieval scores - a bound on score perturbation, not a certificate of unchanged ranking; a memory manager that preserves the cached prefix and overlaps compaction with inference; and a performance model th

    memoryagent memoryagent
  656. arxiv:2609.32255 · cs.LG
    Clarify the User or Verify the World? Uncertainty Routing for Proactive Agents
    Zhaofeng Li, Xuan Zhang, Xiaokui Xiao, Yang Deng

    Tool-using LLM agents must decide not only whether additional information is needed, but also which source can resolve the uncertainty. Existing proactive approaches often specialize in either user clarification or environment verification, without explicitly determining the appropriate information source for each decision. We formulate this problem as uncertainty routing among ACT, CLARIFY, and VERIFY, and propose PROUR, a proactive uncertainty routing framework. PROUR decomposes action uncertainty into disagreement across plausible user-goal interpretations, which signals user-side ambiguity

    llm agent
  657. arxiv:2609.32253 · cs.RO
    DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control
    Yaxing Lyu, Jingyi Li, Mingkun Xu, Yujie Wu

    Vision-language-action (VLA) models have achieved strong performance in language-conditioned manipulation, yet success under nominal evaluation does not necessarily translate into robust closed-loop behavior when executed actions are transiently corrupted. We introduce DS-VLA, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to address this limitation. Specifically, to enable modularized feature processing and temporal information integration, DS-VLA equips action neurons with multiple sparsely connected dendritic branches, each featuring h

    vision-language-actionvlaembodiedmanipulationopenvlagr00t
  658. arxiv:2609.32250 · cs.CV
    RoboSTAR: Next-Scale Autoregressive Sign Language Translation for Humanoid Robots
    Yujia Zeng, Chensheng Peng, Yuxin Chen, Alex Shao +2

    Sign-language interpretation in public communication relies on qualified professional interpreters and can be difficult to scale, motivating robotic signing as a complementary accessibility interface. We present RoBoSTAR, a text-conditioned sign language production (SLP) framework for generating human-centric sign motion that can be retargeted for robotic execution, with speech supported optionally through an external ASR front end. Conventional autoregressive approaches flatten motion into a single full-resolution token sequence, forcing long-range and local dependencies to be modeled at a un

    humanoid
  659. arxiv:2609.32248 · cs.LG
    Anytime-Valid LLM Leaderboards via Benchmark-weighted and Block-Factorized e-Processes
    Hongfu Gao, Songxin Zhang, Zejian Xie, Bingyi Jing +2

    Large language model (LLM) leaderboards compare model capabilities by ranking models according to their mean performance on fixed benchmarks. However, variability in evaluation outcomes across runs may produce unsupported claims of model superiority on the benchmark, a risk compounded by leaderboard updates. In this paper, we propose BB-EDGE (Benchmark-Weighted and Block-Factorized e-processes for Directed Graph Evaluation), a principled framework that represents an LLM leaderboard as a directed graph whose edges certify pairwise mean-performance advantages, with anytime-valid family-wise erro

    benchmarkleaderboard
  660. arxiv:2609.32245 · cs.AI
    AutoPDEBench: Benchmarking LLM Auto-Research for Neural PDE Solver Design
    Ruoyan Li, Wei Wang, Yizhou Sun

    Partial differential equations (PDEs) are essential for modeling complex physical systems, and neural solvers have recently emerged as powerful data-driven tools for numerically solving them. However, existing neural solvers struggle with domain-specific challenges, such as varying parameters and high-speed flows, necessitating specialized architectures. Manually designing these specialized solver architectures is a highly iterative, time-consuming process requiring deep expertise, creating a significant bottleneck in scientific discovery. We propose leveraging autonomous AI research agents to

    ai agentmulti-agentagenticbenchmark
  661. arxiv:2609.32239 · cs.RO
    Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot Manipulation
    Biprodip Pal, Kaushik Roy, Yanming Zhu, Brendan Tidd +2

    Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter aggregation destructive. We present FedDRMan, a federated subspace-guided distillation framework for heterogeneous robot manipulation. At each communication round, the server model provides a frozen teacher for local behavior cloning, while low-rank multimodal subspace and actio

    vision-language-actionmanipulationlibero
  662. arxiv:2609.32236 · cs.RO
    RoboFFT: Finetuning generative robot policy via online reinforcement learning with forward process
    Yu Li, Shenghe Hu, Yuhan Wang, Yaoxiang Pu +3

    Generative models, such as diffusion and flow-based models, have shown strong promise for robot policy learning by capturing complex and multimodal action distributions from demonstrations. However, policies trained solely with imitation learning often suffer from imperfect demonstrations and distributional shifts, while further improvement typically requires additional expert data. Reinforcement learning offers a natural solution through environment interaction, but effectively finetuning generative robot policies remains challenging due to the intractability of likelihood estimation. In this

    robot policybenchmark
  663. arxiv:2609.32228 · cs.LG
    CompassPlay: Rewarding the Proposer for Where It Moves the Solver
    Sophia Xiao Pu, Ximeng Sun, Jiang Liu, Jialian Wu +3

    In self-play, a proposer generates verifiable tasks to train a solver. Proposer rewards often depend on the solver's success rate, but equally difficult tasks can differ in their training value. We introduce CompassPlay, a self-play method that rewards the proposer through gradient alignment. The reward favors tasks whose solver loss gradients align with those of reference tasks representing the target capabilities. It draws on a first-order approximation to learning progress and scores each eligible task without additional solver training. Our experiments show gains in performance and trainin

    self-play
  664. arxiv:2609.32227 · cs.CL
    OptiArena: Can LLMs Improve Executable Algorithms under Fixed Resource Budgets?
    Wenjun Peng, Xinyu Wang

    Static QA and code-generation benchmarks only partially capture the role that large language models (LLMs) now play as coding agents and research tools. We introduce OptiArena, a budget-controlled testbed for studying whether LLMs can improve executable game-playing algorithms through five rounds of code edits within a fixed minimal scaffold and under bounded evaluator feedback and fixed resource budgets. The testbed uses two optimization regimes, surface obfuscation controls, calibrated references, held-out/stress splits, and diagnostics for degradation and exceptional failures, with LLM API

    benchmarkevaluator
  665. arxiv:2609.32226 · cs.AI
    Toward Agentic Optical Networks: A Vision of LLM Agent-Driven Autonomous Lifecycle Management
    Yao Zhang, Shengnan Li, Yuchen Song, Yidi Wang +10

    As optical networks continue to expand in scale, complexity, and service diversity, the implementation of automation has become essential for ensuring agility, efficiency, and reliability in lifecycle management (LCM) of optical networks. Large language model (LLM) Agent, distinguished by its progressively sophisticated capabilities in logical reasoning, adaptive decision-making, complex problem solving, and multi-task orchestration, presents great opportunities to advance network automation beyond traditional AI techniques. Nevertheless, the application of LLM Agent in optical networks remain

    agentllm agentmulti-agentagenticagent framework
  666. arxiv:2609.32225 · cs.AI
    LaMET-Agent: An Agent Framework for Large-Momentum Effective Theory Analysis
    Jinchen He, Xiangyu Jiang, Fei Yao, Dian-Jun Zhao

    Large-momentum effective theory (LaMET) provides a first-principles framework for computing the $x$ dependence of light-cone parton distributions from lattice QCD. Over the past decade, theoretical and numerical advances have established a mature multi-stage workflow for systematic calculation of parton physics, although its implementation still requires expert judgment and substantial repeated effort. We present lamet-agent, an open-source large language model (LLM) agent framework that organizes this workflow into an executable, reproducible, and inspectable analysis pipeline. The present re

    agentagent framework
  667. arxiv:2609.32224 · cs.AI
    RAO-Nav: Probing Omni-Language Models for Zero-shot Semantic Audio-Visual Navigation
    Qilang Ye, Meng Liu, Yu Zhou

    We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to ``hear'', ``see'', ``reason'', and ``act'' in the environment. To further elicit t

    embodiedagentembodied agentbenchmark
  668. arxiv:2609.32222 · cs.CV
    Geometry-Preserving Blind Watermarking for Raw 3D Point Clouds
    Rungui Zhou, Chuanzhi Zhou, Ruihuan Wang, Peng-Shuai Wang

    Raw 3D point clouds are a core geometric representation. Establishing their ownership is challenging because point sets are irregular, unstructured, and frequently altered by resampling and geometric preprocessing. We present a blind watermarking framework that operates directly on xyz coordinates and supports both object-level shapes and scene-scale scans. At verification time, the embedded message is recovered from the observed point cloud alone, without access to the original point cloud, color, normals, or mesh connectivity. The method jointly learns watermark embedding and extraction thro

    benchmark
  669. arxiv:2609.32220 · cs.AI
    A bilingual AI audiologist built through rubric-guided playbook induction outperforms human audiologists in a blinded evaluation of simulated cases
    Linkai Li, Changgeng Mo, Hanlin Yu, Congxi Lu +3

    Audiology consultation requires structured history-taking, audiometric interpretation and patient-centred communication, yet real-world case material is scarce. We present a bilingual AI audiologist pairing a general-purpose large language model with rubric-guided playbook induction, multimodal audiogram interpretation and retrieval-augmented grounding, without fine-tuning the language-model backbone. Using a 21-item rubric and an AI patient simulator, we induced a 19-rule consultation policy from 73 training cases (43 English, 30 Chinese) and evaluated the system on 58 independent simulated c

    retrieval-augmented
  670. arxiv:2609.32213 · cs.LG
    HM-ROUTER: Joint Model and Harness Routing for Agentic Systems
    Hao Mark Chen, Royson Lee, Yasuyuki Okoshi, Dimitris Anastasiou +2

    Agent performance depends on both the underlying model and the harness that manages its tool use and execution. Selecting a suitable pair requires accounting for their compatibility, yet training samples may cover only a subset of the growing combination space. We introduce HM-Router, a routing method that jointly selects a model and harness for each query. It learns separate model and harness representations shared across routes, with an interaction term inspired by canonical polyadic (CP) tensor decomposition to capture how their compatibility varies with the query. This sharing allows train

    agentagenticagent benchmarktool usebenchmark
  671. arxiv:2609.32211 · cs.AI
    Rank Confidence Sequences:Anytime-valid Leaderboards
    Hamed Khosravi, Xiaoming Huo

    Leaderboards rank models by their average scores on benchmark items, and they are consulted repeatedly while the evaluation is still running. Existing confidence intervals for a model's rank control their error rate only if they are computed once, after a number of items chosen in advance. If they are recomputed as results arrive, and the evaluation stops once they look decisive, their error rate exceeds its nominal level. Anytime-valid methods keep their guarantees at all sample sizes simultaneously and hence under any stopping rule. They exist for the accuracy of one model, for one pair of m

    benchmarkleaderboard
  672. arxiv:2609.32208 · cs.AI
    Witness: Discovery, Deciphering, and Epiphany in Interactive Puzzle Environments
    Guanghan Ning, Ping Liu, Linyi Li, Huangjie Zheng +5

    Automated science needs agents that can work out the rules of an unfamiliar environment by interacting with it. Interactive rule-discovery puzzles offer a controlled setting for studying this ability: an agent infers hidden rules through experimentation and uses what it has inferred to reach a stated goal. We ask what limits current language models on these puzzles and whether reinforcement learning (RL) improves performance on rules held out from training. To study both, we introduce WITNESS, a 2D grid-based puzzle environment with ground-truth ASCII observations and controlled access to rule

    agentagenticbenchmark
  673. arxiv:2609.32203 · cs.LG
    Kernel-Based Steering of CLIP with Vision-Language Model Preferences
    Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri, Mahnoosh Alizadeh +2

    Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that learns from elicited pairwise judgments without accessing teacher embeddings or collecting new human similarity annotations. ASK constructs positive semidefinite target kernels within small image groups and combines visual kernel matching with an image--text distributional anchor.

    benchmark
  674. arxiv:2609.32201 · cs.LG
    Instruct, Not Answer: Using Instruction Privileges in On-Policy Context Distillation
    Hantao Yu, Sandy Han, Udaya Ghai, Ferhat Erata +3

    On-Policy Context Distillation (OPCD) has recently emerged as a powerful technique for transferring context to student models and for self-improvement. In OPCD, the teacher is conditioned on privileged information, and the goal is to minimize the Kullback-Leibler (KL) divergence between the privileged teacher and the student, evaluated on student-generated tokens. Many existing studies show that using instance-specific gold answers or gold demonstrations as the default privilege can hurt training performance, especially out-of-distribution (OOD). In this work, we instead design general instruc

    self-improvement
  675. arxiv:2609.32198 · cs.AI
    Agents as Software: A Programming Languages Agenda for Agent Reliability
    Shraddha Barke, Adithya Murali

    AI agents increasingly resemble software systems: they call tools, remember facts, follow policies, delegate work, and take actions with real consequences. % Yet the ``program'' of an agent is scattered across prompts, tools, memories, workflows, and execution traces, making its behavior difficult to inspect through ordinary testing and debugging alone. % This essay argues that a programming-systems perspective offers a natural lens for making agents reliable. % We recast agents as programmable artifacts whose behavior can be specified over traces and state, checked before deployment, monitore

    agentai agent
  676. arxiv:2609.32196 · cs.LG
    The Judge Is Not Its Twin: Post-training makes a model's writing more predictable but barely moves its taste, as a judge, toward predictable writing
    Arman Nik Khah, Arvin Bahreini

    Language models are now routinely graded by other language models. If post-training makes a model's own writing more predictable, it may also teach the same model, acting as a judge, to reward predictable writing, so that progress on creativity would be invisible to automated evaluation. We follow two open model families, OLMo-2 and Zephyr (7B parameters each), through their public training stages and measure every stage twice, as a writer of short stories and as a judge of pairs of stories. As writers, the models drift as feared: each family's fully trained model finds the stories of its base

    post-training
  677. arxiv:2609.32193 · cs.CV
    Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling
    Hongyi Cai, Yi Herng Ong, Tingshiuan C. Wu, Lim Chiew Hui +3

    Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene evolves under candidate actions. World Action Models (WAM) attempt to address this limitation by predicting future states, but existing designs keep prediction and policy learning architecturally separate, connecting them only through the predicted output, whether through pixel space video generation or a latent forecasting module trained independently of the policy. We present Devol-ONE, a Mixture of Transformers architecture that unifies vision l

    vision-language-actionvision language actionliberorobotwinworld modelv-jepa
  678. arxiv:2609.32192 · cs.AI
    CoMemBench: Benchmarking Collaborative Memory Boundaries across Multi-Agent Workflow Topologies
    Sen Zhao, Ruiqi Kong, Zuyu Zhang, Lifeng Shen +4

    Multi-agent workflows require task-relevant information to be shared across agents, while irrelevant, stale, unverified, or incompatible information must remain isolated. We call this task-conditioned scope of information a collaborative memory boundary. Workflow topology determines which intermediate artifacts are applicable to which downstream workers and when they cease to be valid, thereby providing a structural stress dimension for sharing and isolation. Existing memory benchmarks primarily evaluate retention and retrieval, whereas multi-agent benchmarks emphasize coordination and end-to-

    memorymulti-agentagent benchmarkbenchmarkevaluator
  679. arxiv:2609.32188 · cs.CV
    Presence Is Not Faithfulness: Figurative Vehicle Intrusion in Text-to-Image Generation
    Xiaoyu Ma, Chen Yang, Hao Chen

    Text-to-image (TTI) models increasingly generate high-quality images from natural-language prompts, yet figurative language exposes a failure: a vehicle that should guide the depiction of a tenor may instead be rendered as a visible object. We call this failure Figurative Vehicle Intrusion: the intruding content is textually licensed, but it is assigned the wrong visual role, showing that visual presence is not always faithfulness and that presence-oriented evaluation can miss such errors. To study it systematically, we introduce Vehicle Intrusion and Semantic Tenor Assessment (VISTA), a multi

    benchmark
  680. arxiv:2609.32185 · cs.LG
    Contamination, Prior, or Evidence? Decomposing and Training Evidence Use in Whole-Slide Vision-Language Models
    Wenhao Zhang, Zhongliang Zhou, Shiyuan Zhang, Yiqing Yang +5

    Pathology vision-language models (VLMs) are conventionally evaluated by accuracy, but accuracy alone does not measure evidence use: it may conflate dataset contamination, prior knowledge, and image evidence. In a motivating study of lymph-node metastasis prediction, we found that most public pathology VLMs showed minimal differences when changing from feeding the models with whole-slide images, an annotated lesion, or no image at all. To better understand the specific features leveraged by these models, this paper presents two contributions aimed at disentangling these factors. First, we prese

    benchmark
  681. arxiv:2609.32183 · cs.CV
    Scalable In-Domain Self-Supervised Foundation Model for Dense Representation Transfer in High-Resolution Plant Imaging
    Junlin Guo, Sharmin Majumder, Isaac Lyngaas, John Lagergren +1

    High-resolution plant imaging enables detailed characterization of plant morphology, but dense scientific analysis remains limited by costly pixel-level annotations, large image pixel dimensions, and substantial variation in imaging conditions. This work proposes a scalable in-domain self-supervised pretrained foundation model for high-resolution, high-pixel-dimension multi-species plant imagery. A masked autoencoder with a ViT backbone is pretrained on more than 10 million multi-view plant image tiles using distributed training. Following scalable pretraining, the learned foundation-model rep

    benchmark
  682. arxiv:2609.32182 · cs.CV
    KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding
    Zihan Chen, Xuejian Rong, Xiaojuan Wang, Boqing Gong +3

    Vision-language models are increasingly used to understand long videos and continuous streams. However, dense visual tokens accumulate with video duration, making long-context inference prohibitively expensive. Existing training-free visual-token selection methods reduce this cost by retaining informative tokens, but may lose coherent event evidence and fail to distinguish detailed recent observations from long-range history. We propose \textbf{KeyRec}, a training-free framework for constructing bounded visual memory. During query-agnostic writing, KeyRec preserves fine-grained recent observat

    memorylong-contextbenchmark
  683. arxiv:2609.32180 · cs.CV
    Binaural Audio-Visual Instance Segmentation
    Saijun Wang, Guanfeng Tang, Hongbo Zhao, Zhicheng Lei +3

    Audio-visual segmentation (AVS) aims to segment sounding objects at the pixel level by integrating auditory and visual cues. However, existing methods are predominantly developed under the monaural setting and primarily rely on cross-modal semantic correspondence, which limits their ability to distinguish visually similar instances of the same semantic class. In contrast, humans naturally exploit binaural hearing, where interaural differences and direction-dependent acoustic filtering introduced by the head and pinnae provide physically grounded spatial cues for accurate sound source localizat

    benchmark
  684. arxiv:2609.32178 · cs.LG
    Analytic-Walk Rotary Positional Encodings for Graphs
    Jiaqing Xie, Yuxin Wang, Xipeng Qiu

    Rotary position encodings make attention sensitive to relative position, but extending them to graphs requires choosing how graph structure enters the rotation. Previous works assign each node a rotation from spectral coordinates, so the rotary factor between two nodes depends only on their endpoints and cannot distinguish the routes connecting them. We introduce \textit{Analytic-Walk Rotary Positional Encodings} (AW-RoPE), which place the rotations on edges and sum the transported features over all walks, so contributions along different routes can reinforce or cancel. An exact variant evalua

    benchmark
  685. arxiv:2609.32177 · cs.CV
    Federated 3D Gaussian Splatting for Large-Scale Scene Reconstruction at Wireless Edge
    Guanlin Wu, Chao Hu, Pu Chen, Juyong Zhang +3

    Three-dimensional (3D) Gaussian splatting (3D-GS) has emerged as a promising technique for large-scale scene reconstruction due to its high rendering efficiency and fidelity. However, the training of large-scale 3D-GS models at wireless edge faces various technical challenges including the limited communication, computation, and graphics processing unit (GPU) memory resources at edge devices, the structural inconsistency issue across local models hindering their effective aggregation, as well as privacy leakage risks associated with raw visual content and camera parameters. To address these ch

    memory
  686. arxiv:2609.32172 · cs.AI
    Noisy Test-Time Reinforcement Learning for Code LLMs
    Xikai Yang, Hieu Trung Nguyen, Dunyuan Xu, Yuzhi Zhao +3

    Large language models (LLMs) have demonstrated remarkable performance across various code-related tasks. However, unlike carefully curated datasets that are typically high-quality and error-free, real-world user instructions are often vague and error-prone, posing significant challenges to the robustness of code LLMs. Furthermore, robustness-oriented fine-tuning relies on paired clean-noisy samples, which are costly to curate and require sophisticated noisy simulation techniques. To address these challenges, we propose the Noisy Test-time Reinforcement Learning framework (NTRL-Code), which ena

    benchmark
  687. arxiv:2609.32170 · cs.LG
    Uncertainty-Aware Selection of Online Algorithms with Simulator Ensembles
    Yongyi Guo, Zifan Xu, Ziping Xu, Kelly W. Zhang

    The performance of online reinforcement learning depends critically on design choices, especially those that affect exploration. These choices are often selected by fitting a simulator to offline data, evaluating candidate algorithms in that simulator, and deploying the best-performing one. The simplest Plug-In selection rule simply selects the best performing algorithm on the fitted simulator, making evaluations unreliable when the offline data used to fit the simulator are limited. We investigate Uncertainty-Aware selection, which forms an ensemble of simulators---for example, obtained by bo

    sim-to-real
  688. arxiv:2609.32169 · cs.LG
    Uncertainty Quantification of Next Generation Reservoir Computing with Applications to Memory-Driven Dynamical Systems
    Livia Popa, Sumanta Basu, Martin T. Wells

    Nonlinear dynamical systems with memory arise across science and engineering, yet uncertainty quantification for efficient forecasting methods such as Next Generation Reservoir Computing (NGRC) remains underdeveloped. We study Bayesian ridge and conformal prediction intervals for NGRC and characterize when their uncertainty estimates agree or differ. In low dimensions, their asymptotic widths are governed by different summaries of the residual distribution, so agreement depends on residual shape rather than dimensionality alone. In high dimensions, regularization introduces a further tradeoff

    memory
  689. arxiv:2609.32166 · cs.AI
    PastForward: Faster On-Device GUI Agents via Computational Experience Reuse
    Taehwan Park, Changmin Lee, Hayeon Lee, Taesik Gong

    Running GUI agents on edge devices can keep sensitive screens and interaction histories local, but the computational cost of inference at every action step makes deployment challenging. Existing GUI agent systems either perform full vision-language model (VLM) inference at each action step or reuse coarse-grained knowledge matched to prior tasks. However, dynamic mobile environments and user tasks make it difficult to fully utilize prior task executions without additional fine-tuning or task-specific offline exploration. To address this challenge, we present PastForward, a system that accelera

    agentagent system
  690. arxiv:2609.32162 · cs.CV
    PruneForget: Joint Unlearning and Pruning of Vision Models
    Yu-Shan Tai, Amber Yijia Zheng, Raymond A. Yeh

    Machine unlearning and model pruning are increasingly coupled in the real world. Models must support unlearning requests, e.g., for safety concerns, while also meeting requirements in latency and memory budget. Until recently, existing works have studied each aspect as an independent problem, e.g., running unlearning and pruning sequentially. In this work, we show that unlearning and pruning are naturally aligned and should be solved jointly to be made aware of each other. Intuitively, parameters that encode information of the unlearned samples are natural pruning targets, as unlearning and pr

    memory
  691. arxiv:2609.32158 · cs.RO
    FutureRay: Control-Aligned Future Range for Agile Quadruped Navigation
    Tianhao Zang, Shanze Wang, Ziqian Wang, Liyou Luo +3

    Moving obstacles can block a previously clear route while a quadruped robot executes a motion command. We investigate whether predicting changing clearance improves navigation when motion selection accounts for the robot footprint and the time needed to react and brake. We present FutureRay, which predicts ranges across viewing directions and future times, together with encounter risk, from depth-derived range history and observable robot motion. Training emphasizes near-term clearance and penalizes errors that overstate available space. A local planner queries the same forecast for candidate

    quadruped
  692. arxiv:2609.32157 · cs.RO
    CausalDriveBench: Evaluating Causal Reasoning in Vision-Language-Action Models for Autonomous Driving
    Narendiran Chembu, Navvrat Rao, Shreedhar Shreeshail Kodate, Gayatri Srujana Banda +9

    Vision-Language-Action (VLA) models for autonomous driving produce natural-language reasoning alongside predicted trajectories, but whether this reasoning reflects the causal structure of the scene remains untested. We introduce CausalDriveBench, an evaluation framework grounded in Pearl's Causal Hierarchy (PCH) that tests causal reasoning in driving-specific VLAs through structured visual question answering (QA) and alternative-trajectory prediction. To this end, we construct causal scene graphs over nuScenes that distinguish causally active, dormant, and distractor entities, separating perce

    vision-language-actionvlascene graphpost-trainingbenchmarkevaluation framework
  693. arxiv:2609.32155 · cs.RO
    RecastVLA: From Past Interaction to Future Control with Adaptive Policy States
    Wenbo Li, Jun Yang, Yiteng Chen, Wei Zhang +1

    Sequential manipulation requires a robot to track what has already happened, even when the current scene no longer reveals it. Policies with explicit history representations make past interactions available as context for current decisions. We ask how action generation itself can form a persistent state for subsequent control. Building on action-side test-time training, RecastVLA maintains an adaptive policy state within a flow-matching vision-language-action policy. The state is represented by shared fast weights and remains fixed throughout action generation. Depth-specific interfaces read t

    vision-language-actionmanipulationliberorobotwinpersistent state
  694. arxiv:2609.32154 · cs.RO
    GAUGE: Planner-Conditioned Active Calibration of Opaque Quadruped Velocity Interfaces
    Tianhao Zang, Zihan Liu, Shanze Wang, Liyou Luo +3

    In this paper, we present a Goal-Aware Uncertainty-Guided Exploration (GAUGE) framework for planner-conditioned active calibration of opaque quadruped velocity interfaces. Commercial quadrupeds commonly expose planar-velocity commands, but the underlying locomotion controller remains inaccessible and can produce systematic discrepancies between commanded and realized motion. A navigation planner typically uses a structured subset of the command envelope. GAUGE maintains a Bayesian command-to-motion model and selects authorized trials according to their expected reduction of posterior epistemic

    quadruped
  695. arxiv:2609.32149 · cs.LG
    A Unified Optimism-Agnostic Framework for Linear Bandits over Spherical Action Sets
    Arda Güçlü, Subhonmesh Bose, John R. Birge

    Linear bandits model sequential decision-making problems with noisy rewards that are linear in the decision variable, where an agent must simultaneously learn about an unknown parameter that governs the mean rewards, while maximizing (expected) rewards over time. Two prominent algorithmic families--upper confidence bound (UCB) and Thompson sampling (TS)--achieve a balance of exploration (to estimate said parameter) and exploitation (utilization of knowledge about it) across time. The quality of estimation of that parameter depends on the eigenvalues of a design matrix. In this paper, we begin

    agent
  696. arxiv:2609.32146 · cs.LG
    Playing to Par: Reinforcement Learning for Provably Optimal Quadrilateral Block Decompositions
    Arjun Narayanan, Per-Olof Persson

    A quadrilateral block decomposition of a planar domain is judged by whether it is complete, whether its elements are well shaped, and how many of its vertices are irregular. The last has a provable floor: the discrete Gauss-Bonnet identity enforces a lower bound on the total vertex irregularity of any all-quadrilateral mesh of a given domain purely based on its topology and corner angles. We train a reinforcement learning agent to build decompositions that reach this bound, which we call par. It acts directly on the mesh's half-edge data structure through local edits, with a policy network who

    agent
  697. arxiv:2609.32143 · cs.LG
    Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting
    Hanxu Yang, Yuhuan Zhao, Xiaodong He, Zhao Kang

    Pre-training Graph Neural Networks (GNNs) via self-supervised learning has become a dominant paradigm, yet efficiently adapting frozen encoders remains a challenge. Graph prompting offers a parameter-efficient alternative to fine-tuning, but existing methods largely treat pre-trained models as opaque feature extractors, ignoring their internal spectral structure. In this work, we identify a systematic phenomenon in pre-trained GNNs, which we term spectral bias: optimization during pre-training disproportionately aligns representations with directions associated with large singular values, leav

    benchmark
  698. arxiv:2609.32134 · cs.CL
    Checking Leakage Witnesses versus Certifying Bounded Non-Leakage
    Chao Feng, Burkhard Stiller

    When a language-model audit finds no leak, what is needed to certify non-leakage? We study guarantees over a declared prompt domain under an executable leakage criterion and decoding rule. For general bounded polynomial-time evaluators, a supplied leaking execution is polynomial-time checkable, while leak existence is \NP-complete and deterministic certification is \coNP-complete. Exact stochastic certification is $\coNP^{\PP}$-complete at every fixed rational cutoff in $(0,1)$. Restricting the computation can change these bounds. For example, certification is in \coNP\ when all randomness is

    evaluator
  699. arxiv:2609.32129 · cs.RO
    Residual Denoising Enables Sample-Efficient Multi-Agent Coordination on Demand
    Dayi Dong, Maulik Bhatt, Aayushi Shrivastava, Lasse Peters +1

    Pretrained robot policies offer strong manipulation skills but are typically limited to single-agent settings, where a robot acts in isolation. In this work, we study how to adapt pretrained single-agent diffusion policies to multi-agent settings using minimal collaborative data, co-optimizing for two key objectives: high coordination performance and single-agent skill retention. To this end, we introduce ALTER, an adaptation method for coordination on demand: the adapted policy coordinates with other robots when deployed in a team while remaining capable of acting independently when operating

    manipulationmulti-agent
  700. arxiv:2609.32123 · cs.AI
    READ-Bench: Benchmarking Historical Instance Retrieval for Time-Series Diagnosis
    Gerardo Pastrana, Haojun Li, Dhruv Mehta, Anoushka Vyas +3

    Time-series diagnostic systems rarely rely on retrieving relevant historical cases, and when they do, retrieval is evaluated only indirectly through downstream prediction. We introduce READ-Bench, a benchmark for historical-case retrieval across 12 diagnostic datasets, centered on multivariate time series, that defines relevance by shared fault or event type rather than signal shape, so visually different traces of the same fault count as relevant while similar-looking traces of different faults do not. Treating retrieval as a base retriever followed by a reranker, we evaluate classical distan

    benchmark
  701. arxiv:2609.32122 · physics.app-ph
    In-Memory AM Demodulation Using an All-Silicon Independent-Dual-Gate Gain-Cell Memory with $<10^{-22}$ A Leakage Determined by Single-Electron Counting
    Katsuhiko Nishiguchi, Toshiaki Hayashi, Kensaku Chida, Takase Shimizu +2

    Ultra-low-leakage memories are attracting increasing attention for in-memory sensing and computing. However, achieving sufficiently long retention in silicon memories for analog signal processing remains challenging because of leakage through the access transistor. In this work, we demonstrate an all-silicon independent-dual-gate memory whose leakage current, inferred from single-electron counting statistics, is below $10^{-22}$ A, giving a measured retention time exceeding 1000 s. Owing to the extremely low leakage, the subthreshold nonlinearity of the access transistor can be exploited witho

    memory
  702. arxiv:2609.32119 · cs.CL
    Using LMs to Model the Effects of Context and Coreference during Sentence Comprehension
    Kohei Kajikawa, Lin Ai, Tatsuki Kuribayashi, Ethan Gotlieb Wilcox

    Language models (LMs) are often used as a tool to model human language processing. Recent studies suggest that severely restricting LMs' context window improves their fit to human psycholinguistic data by simulating human working memory constraints. However, it is possible that this strict memory-decay approach overlooks humans' reliance on long-range structural representations, such as discourse structre. In this work, we systematically vary the context window size of GPT-2 across four large-scale naturalistic English reading-time datasets and observe a U-shaped relationship: Although restric

    memory
  703. arxiv:2609.32114 · cs.AI
    Empowering Hybrid Attention Models on NPUs
    Yinyuan Zhang, Daliang Xu, Xiaolong Huang, Wangsong Yin +3

    Hybrid attention models have emerged as a crucial architecture for Large Language Models (LLMs) (e.g., the Qwen3.5 and Kimi series). Their memory and computational efficiency make them highly attractive for on-device inference, forming a promising synergy with edge Neural Processing Units (NPUs). However, naive execution of these hybrid models on edge NPUs fails to deliver these benefits, often bottlenecking the prefill stage due to severe memory-system inefficiencies and architectural mismatches within the linear attention (LA) layers. We present HA-NPU, the first system to enable efficient h

    memory
  704. arxiv:2609.32109 · cs.LG
    How Reusable Are Benchmarks with Richer Feedback?
    Youssef Allouah, John Duchi

    We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $Θ(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated t

    benchmark
  705. arxiv:2609.32108 · cs.RO
    SAMBAR: Selective Anchoring via Method of Multipliers for Balanced Knowledge Acquisition and Retention in Vision-Language-Action Models
    Aayushi Shrivastava, Xunlan Zhou, Hongrui Zhao, Ziyu Chen +1

    Vision-Language-Action (VLA) models leverage large-scale pretraining to ultimately achieve generalist manipulation. Deployed VLA policies must support continual learning to acquire new tasks over time. Teaching a VLA a new task generally requires finetuning it on demonstrations of that task. However, naively finetuning on downstream tasks causes the policy to forget earlier tasks and degrades generalist capabilities. This failure is known as catastrophic forgetting. Most continual learning methods counter it by replaying data from earlier tasks. However, the old task demonstrations are not alw

    vision-language-actionvlamanipulationliberobenchmark
  706. arxiv:2609.32103 · cs.LG
    LLM Unlearning Evaluation with TRIAGE
    Danial Ataee, Peter Triantafillou

    Large language models can memorize private or harmful information, motivating machine unlearning methods that remove targeted knowledge while preserving other capabilities. However, existing evaluations rely primarily on behavioral benchmarks, which assess \emph{whether} a model appears to forget but provide limited insight into \emph{how} unlearning changes the model or affects related knowledge. We introduce \textit{TRIAGE} (\textit{Tripartite Representation-internal Introspection for Adjacency Gap Evaluation}), a benchmark-agnostic evaluation framework for characterizing these changes. TRIA

    benchmarkevaluation framework
  707. arxiv:2609.32098 · cs.MA
    Robust Game-theoretic Motion Planning over Extended Time Horizons
    Bennet Outland, Vishala Arya

    This work presents a solution to nonconvex, game-theoretic motion planning problems subject to disturbances over long time horizons. The problem is posed as a partially-decoupled generalized Nash equilibrium problem, in which each agent's dynamics depend only on its own state and control, admitting fast solution methods for competitive multi-agent motion planning. An algorithm, WOLF, is developed that applies receding-horizon model predictive control to an open-loop differential games solver based on sequential convexification. In contrast to robust formulations that fix the uncertainty descri

    multi-agent
  708. arxiv:2609.32093 · cs.AI
    GameBoyWorlds: A Testbed for Self-Improvement in Embodied Video Games
    Dhananjay Ashok, Adam Shen, Aslan Huo Feng, Chinmay Khanna +9

    Powered by expert guidance, agents can operate in interactive environments; however, it is unclear whether they can learn autonomously from their own experience. To evaluate such self-improvement methods, we introduce GameBoyWorlds, a testbed for agentic self-improvement in video games. GameBoyWorlds-Execution evaluates task execution on a collection of 5 distinct game series. Agents are allowed access to dedicated training games but are provided no demonstrations, documentation, or rewards. Agents must ground themselves in the environment through self-directed exploration and by inferring act

    embodiedworld modelmemoryagenticself-improvingself-improvement
  709. arxiv:2609.32092 · cs.AI
    On Evaluating and Improving Conversational Agents in Production
    Kasra Hosseini, Wen-Sen Cheng, Marco-Andrea Buchmann, Emir Mulabegovic +1

    We present a framework for evaluating and improving a large-scale, multi-agent shopping assistant in production, and report lessons from its use. Offline evaluation of such a system faces three obstacles. (i) A logged conversation cannot be replayed against a modified system, because a different response changes every turn that follows. (ii) The unchanged system itself varies from run to run. Its LLM components are stochastic, and in product search the available products, their prices, and the customer's personalization signals change. (iii) Aggregate quality scores combine distinct behaviors,

    multi-agent
  710. arxiv:2609.32091 · cs.AI
    Memory as Middleware for Self-Improving AI Agents
    K. R. Jayaram, Vatche Isahagian, Vinod Muthusamy, Gegi Thomas +3

    AI agents are stateless across sessions by default and therefore operationally amnesic: each session begins with little durable knowledge of prior failures, repairs, preferences, or successful strategies. As a result, agents repeat the same mistakes and discard hard-won experience. The dominant fix is \emph{bespoke memory}---retrieval, persistence, and learning logic hand-wired into one agent and bound to one storage engine. This creates a fragmented landscape where memory cannot be swapped, shared, isolated, or reasoned about independently of the agent that owns it. We argue that this is a mi

    memoryagent memoryagentai agentself-improving
  711. arxiv:2609.32082 · cs.CL
    mu-bench: A Multilingual Utterance Transcription Benchmark
    Andrea Li, Soham Ray

    Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an AI banking agent in English, Spanish, Turkish, Vietnamese, and Mandarin, centered on form-field inputs such as names, email addresses, and confirmation codes. We release Utterance Error Rate (UER), an LLM judge of whether a transcript preserves meaning, calibrated against human

    agentbenchmarkleaderboard
  712. arxiv:2609.32081 · cs.AI
    Toward Interactive Understanding of Code APIs
    Dhananjay Ashok, Jesse Thomason, Jonathan May

    Empowered by advances in Language Model agents, systems have made substantial strides in code generation and understanding. However, these approaches often rely on read access to the relevant code, an assumption which does not hold when dealing with external APIs. In this work, we introduce the PAU (Python API Understanding) benchmark, where we provide models with black-box, API-level access to code snippets. Models must query the API with exploratory inputs and draw insights from the resulting outputs, with the goal of describing the snippet's true functionality. By treating the code snippets

    benchmark
  713. arxiv:2609.32079 · cs.RO
    Fiber-Normalized Manipulability and Determinant Proxies: Intrinsic Redundancy Optimization Across and Within Task Fibers
    Antonio Franchi, Mirko Mizzoni

    This work establishes that determinant-based manipulability is an exact objective for fixed-task redundancy optimization, despite its dependence on task coordinates and the choice of task-space metric used for volume measurement. On every regular task fiber, the determinant proxy, its representation in any task chart, and every metric-completed manipulability differ only by positive constants. They consequently induce the same complete ordering, constrained extrema, gradient directions, critical points, and local optimality classifications. For comparisons and trajectory optimization across ta

    manipulator
  714. arxiv:2609.32069 · cs.RO
    Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models
    Yuan Fang, Zechu Li, Haolei Tong, Puze Liu +1

    Vision--language--action (VLA) models provide strong priors for robotic manipulation but are typically deployed as frozen policies, unable to improve from their own failures. Real-world reinforcement learning (RL) offers a path to continued improvement, yet manual environment resets and task-success supervision hinder autonomous learning. We introduce \textbf{FIND}, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace. FIND reframes autonomous practice as a scene-condition

    vlavla modelmanipulationagentagenticself-improving
  715. arxiv:2609.32064 · cs.RO
    Grasp2Twist: Learning Bimanual Dexterous Jar Opening by Reinforcement Learning
    Mo Xu, Yunfu Deng, Jianuo Wang, Josiah Hanna +1

    This paper presents Grasp2Twist, a bimanual dexterous manipulation system that learns to grasp and twist open jar lids using reinforcement learning. Learning this task raises three challenges: learning a unified policy for a multi-stage task, sustaining lid twisting, and sim-to-real transfer. To address the first challenge, we introduce a continuous enclosure measure to guide grasp formation and a binary enclosure indicator to guide the grasp-to-twist transition for unified policy learning. We derive both from the geometric relationship between the object center and the convex hull formed by t

    manipulationdexteroussim-to-realgrasp
  716. arxiv:2609.32060 · cs.LG
    Depth Laws for the Precision Floor of Trained Neural Networks: Amplification, Residual Scaling, and a Quantization-Aware Training Paradox
    Ahmad S. Tarawneh

    How many bits does a network need before its accuracy collapses, and how does this grow with depth? We study the precision floor, the perturbation level or bit-width at which accuracy falls halfway to chance, in MLPs, CNNs, Vision Transformers and nine pretrained language models, under post-training quantization (PTQ) and quantization- or noise-aware training (QAT). (i) A first-order theory sets the floor through one full-precision quantity, the predictive amplification $G$: $η_c=Λ/G$, and $G^2$ grows linearly in depth at a rate proportional to the squared residual branch scale. (ii) The predi

    post-training
  717. arxiv:2609.32049 · cs.AI
    EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic Memory
    Bhavyateja Potineni, Lohit Giri, Anu Jain, Vadim Kutsyy +1

    As autonomous LLM agents are deployed across multi-session environments, conventional memory architectures suffer from Associative Blindness (inability to traverse multi-hop relational dependencies), Scaffolding Amnesia (temporal decay evicting core persona invariants), and Static Topology Stagnation (immutable graphs ignoring usage dynamics). Grounded in Complementary Learning Systems (CLS) principles, we propose EngramRAG, an adaptive memory architecture coupling a low-latency Waking State reflex with an asynchronous background Dreaming State consolidation cycle. EngramRAG introduces: (1) Us

    memorymemory architecturellm agentagenticbenchmark
  718. arxiv:2609.32048 · cs.LG
    Interactive Distributionally Robust Multi-Agent Learning with General Function Approximation
    Debamita Ghosh, George K. Atia, Yue Wang

    Model misspecification poses a fundamental challenge in multi-agent reinforcement learning, where transition uncertainty can be amplified by strategic interactions among agents. Distributionally robust Markov games (DRMGs) provide a principled framework for addressing such uncertainty, yet existing methods often rely on restrictive assumptions or scale poorly to large state and joint action spaces. We study online learning in general-sum DRMGs with general function approximation and $φ$-divergence uncertainty sets. We propose RoMEX-$φ$, a model-free framework that integrates equilibrium-based

    multi-agentonline learning
  719. arxiv:2609.32046 · cs.AI
    Receiver-Conditioned Latent Communication gives 94% CacheBack
    Maximillian Rossi, Prajwal Raghunath, Haoqing Xuan, Yusen Zhang +1

    Multi-agent systems distribute large contexts across agents that communicate to solve a task. Text messages are compact but require decoding and may omit evidence the receiving agent needs. Recent latent communication instead transfers KV caches. This avoids text generation and can improve accuracy and latency. However, a full KV cache grows linearly with both the context an individual agent processes, and the number of agents that coordinate together. This raises memory and context costs, often far exceeding available GPU resources and context window sizes. Our key observation is that agents

    memoryagentmulti-agentagent system
  720. arxiv:2609.32042 · cs.CL
    Quantization Thresholds Replicate, Failure Modes Do Not: A Three-Model Study of Agentic Tool Use in Polish from 8-bit to 2-bit
    Jakub Prejzner

    We ask how GGUF quantization affects agentic tool use in Polish and whether the effects generalize across models. We introduce PolAgentBench, a deterministic benchmark with Polish prompts and English tool schemas: a 67-task main suite (15 adversarial probes, 52 hard-tier tasks) and a 46-task arithmetic isolation ladder. Three models span two axes of variation: Bielik-11B-v3.0 and its pruned, distilled child Bielik-Minitron-7B-v3.0 isolate model compression, and Llama-PLLuM-8B adds a change of pretraining family. Each is measured at six precisions, Q8_0 to Q2_K. Only the collapse threshold repl

    agentictool usebenchmark
  721. arxiv:2609.32041 · cs.LG
    Amnesia by Design, Memory By Necessity: Persistent State for Document Intelligence
    Souhail Bakkali, Ayoub Merimi

    Modern Document AI reads contracts, extracts fields, reasons over tables, and grounds answers to page regions, then forgets everything. Processing an amendment the next day begins from scratch: no schema retained, no contradiction detected, no experience carried forward. This is a structural choice, not a scale failure: current systems are stateless functions. We call this the statelessness bottleneck. This bottleneck lies beyond parameter scaling, context extension, and retrieval augmentation: storage provides persistence and retrieval provides access, but neither consolidates observations in

    memorypersistent statebenchmark
  722. arxiv:2609.32036 · cs.CV
    ScreenHaystack: Finding Blind Zones in GUI Grounding
    Chenyue Li, Xiaoxiao Sun, Yubo Deng, Qinlin Zhao +2

    We introduce ScreenHaystack, a dynamic needle-in-a-haystack benchmark for evaluating spatial reliability in GUI grounding. Instead of testing each target at a fixed position, ScreenHaystack systematically relocates controlled target icons across high-resolution GUI backgrounds and measures whether models can localize them consistently. Using this benchmark, we find that leading GUI grounding models, including Qwen3-VL, UI-TARS, GTA, and UI-Venus, exhibit blind zones: spatial regions where grounding accuracy drops sharply despite fixed target appearance and instruction. These blind zones transf

    benchmark
  723. arxiv:2609.32035 · cs.AI
    Reasoning Concentrates Errors, and Self-Consistency Never Notices
    Asaad Althoubi

    Self-consistency assumes that independent samples disagree when a model is unsure, so agreement is evidence of correctness. Holding weights fixed and toggling only a reasoning mode, over five benchmarks and 74,944 samples, we show that reasoning concentrates a model's errors: the probability that two independently drawn wrong answers coincide rises in all ten dataset-scale comparisons (p = 0.00098), and in nine of nine after restricting both arms to the problems each gets wrong. Where the answer space is unbounded, reasoning cuts the distinct answers produced to 0.43-0.65 of the non-reasoning

    benchmark
  724. arxiv:2609.32027 · cs.RO
    Depth Any Seen: Which Surfaces and How Far?
    Xiaohao Xu, Xiaonan Huang

    When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-free Exact Multi-Bernoulli objective (ExactMB) learns depth and presence by marginalizing one-to-one assignments to complete, distinct targets. Our analysis shows that matching expected count can leave component-surface assignment unresolved. We extend real and synthetic layered-

    benchmark
  725. arxiv:2609.32021 · cs.AI
    SilentCall: Hidden Tool-Call Backdoors in Open-Weight Agents, and How to Catch Them
    Bhanu Pallakonda, Mikkel Hindsbo, Sina Ehsani, Prag Mishra

    Open-weight tool-calling agents are adopted on evidence of merit, usually benchmark scores and a record of reliable use. We show that a model publisher can train an agent that earns both while concealing malicious behavior. Fine-tuned on a mixture of clean and poisoned conversations, our agents answer ordinary requests correctly; once the system date reaches a chosen year, they emit the correct tool call and, alongside it, one that exfiltrates the user's credentials. The exfiltration runs while the user-facing response mentions only the legitimate work. We call this attack SilentCall. Under th

    agentbenchmark
  726. arxiv:2609.32020 · cs.AI
    A Benchmark for LLM's Understanding of Middle School and High School Science Topics
    Noah L. Schroeder, Yessy Eka Ambarwati, Yuji Zhang, ChengXiang Zhai

    Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level ps

    human-in-the-loopbenchmark
  727. arxiv:2609.32019 · cs.AI
    Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
    Valeriy Vyaltsev, Anton Andreychuk, Taisia Zlotnikova, Konstantin Yakovlev +2

    Decentralized multi-agent path finding (MAPF) with communication requires agents to reach individual goals without collisions under partial observability. Learnable policies trained on expert data provide an effective approach to this problem. However, when several coordinated joint actions are valid in the same context, independently sampling from per-agent distributions can recombine locally valid choices into incompatible joint actions. This failure can arise from the final sampling mechanism even when the per-agent action distributions are learned correctly. DMM (Decentralized Master-Mind)

    multi-agentiterative refinement
  728. arxiv:2609.32016 · cs.AI
    VoiceNet: Fine-Grained Voice Understanding Beyond Emotion at Scale
    Christoph Schuhmann, Robert Kaczmarczyk, Gollam Rabby, Felix Friedrich +6

    Expressive speech synthesis has outpaced expressive speech perception: systems now render fine-grained vocal performances that no public benchmark can score. Most benchmarks for this inverse problem stop at six to nine basic emotion categories, largely on acted speech. This paper introduces VoiceNet, a human-annotated representation-level benchmark for voice performance understanding on permissively-licensed in-the-wild speech. VoiceNet has two subsets: VoiceNet-Emo applies a 40-emotion taxonomy with three expert ratings per item, and VoiceNet-Ext, a preliminary subset, scores 57 talking-style

    benchmark
  729. arxiv:2609.32013 · cs.CV
    TriO: Tri-Modal Unsupervised Occupancy World Model for Anything Perception
    Quinlan Sykora, Sourav Biswas, Christopher Diehl, Andrew Cunningham +2

    We present TriO, a multi-modal unsupervised world model that predicts 4D occupancy, obstacle segmentation, flow and LiDAR. In contrast to prior work, TriO utilizes three distinct sensor modalities (camera, LiDAR, and RADAR) as both inputs and sources of self-supervision, eliminating the need for additional human annotations. Thanks to its novel supervision, the model is able to segment any occupancy from the drivable surface, overcoming the limitations of existing open-set methods in handling long-tail objects. TriO achieves state-of-the-art results in multiple 3D and 4D tasks, including occup

    world model4d occupancyoccupancy world

02 US SEMI · SEC 8-K FILINGS

2 items

scanned: NVDA / AVGO / MRVL / COHR / LITE / AMD / TSM / SMCI / ANET / CRDO / POWL / VECO

  1. $AMD · 8-K · filed 2026-09-28
    Advanced Micro Devices Inc
    Items: 3.02
    8-K
  2. $MRVL · 8-K · filed 2026-09-25
    Marvell Technology Inc
    Items: 8.01,9.01
    8-K

03 HUMANOID · COMPANY NEWS

54 items

scanned: figure-ai / 1x / boston-dynamics / unitree / apptronik / sanctuary-ai / neura-robotics / agility-robotics / physical-intelligence / agibot

04 CN PHOTONICS · 公告流

0 items
CN 源 尚未实装 (TIER-1 下一步)