PHYSICAL AI · 2026-09-29

Physical AI Brief

Daily cross-source signals for the Physical AI supply chain — silicon photonics, CPO, VLA models, humanoid hardware, embodied AI. Three streams, one page, zero filler.

1222 items today · 1166 arxiv · 2 SEC 8-K · 54 humanoid · 0 CN photonics

01 ARXIV · PHYSICAL AI PAPERS

1166 items
  1. arxiv:2609.35765 · cs.CL
    Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales
    András Kovács, Alexander Conroy, Daniel Hershcovich, Jens Bjerring-Hansen

    Identifying intertextual references is central to literary scholarship, but computationally difficult when source material is transformed through paraphrase, allusion, historical language, and translation. We investigate this problem through biblical intertextuality in Karen Blixen's Seven Gothic Ta

    benchmark
  2. arxiv:2609.35764 · cs.CV
    Reliability-Gated Fusion of Consumer Head and Foot IMUs for Lower-Body 3D Pose
    Zhilin Guo, Boqiao Zhang, Oszkár Urbán, Josef Bengtson +7

    Sparse inertial pose estimation promises camera-free motion capture from consumer devices, but consumer sensors are unreliable: firmware-fused orientations are biased, mounting varies between sessions, and streams drift or drop out. On a new 35-take single-subject benchmark pairing an earbud head in

    benchmark
  3. arxiv:2609.35761 · cs.RO
    DexRoam: Learning Mobile Bimanual Dexterous Manipulation from Egocentric Whole-Body Human Demonstrations
    Rui Zhou, Yibo Yuan, Junkai Zhao, Fangyuan Zhao +3

    Mobile bimanual dexterous manipulation requires continuous coordination of locomotion, whole-body motion, and finger-level dexterity within a single trajectory, creating a severe robot demonstration bottleneck. Egocentric human demonstrations offer a scalable alternative, but prior approaches ease t

    vlamanipulationdexterouspi0gr00t
  4. arxiv:2609.35760 · cs.LG
    TokenCast: Forecasting Token Consumption During LLM Agent Execution
    Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo +6

    When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call.

    agentllm agent
  5. arxiv:2609.35759 · cs.CL
    Scaling Long-Form Story Generation via Narrative State Tracking
    Zhennan Wan, Jianfei Chen

    LLMs have demonstrated strong capabilities in creative writing. However, scaling them to full-length novels remains challenging, as maintaining narrative consistency becomes increasingly difficult. Existing story-generation methods typically focus on stories of up to about ten thousand words, leavin

    agentagenticbenchmark
  6. arxiv:2609.35752 · cs.LG
    Neural Harmonic Measure Operator
    Jinjin He, Sinan Wang, Yuchen Sun, Bo Zhu

    We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on

    benchmark
  7. arxiv:2609.35750 · cs.LG
    KV-streams for Efficient Compaction in Agentic Reinforcement Learning
    Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer +14

    Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling

    memoryagenticpost-training
  8. arxiv:2609.35748 · cs.LG
    Improving Test-Time Scaling with Adaptive Looped Transformers
    Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv +3

    Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs gro

    post-trainingbenchmark
  9. arxiv:2609.35744 · cs.AI
    FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents
    Hoyoung Lee, Suyeol Yun, Jack Haverty, Yunju Cho +16

    Evaluating finance research agents requires rubrics that reflect expert standards and fix the values correct as of an information cutoff. Expert-reviewed finance benchmarks rely on fixed, per-item rubrics, which are costly to extend and cannot encode each institution's own standard. In FinAutoRubric

    agentbenchmark
  10. arxiv:2609.35743 · cs.CV
    InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
    Kerui Ren, Kaiwen Song, Weiguang Zhao, Yuxi Wang +7

    World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe c

    memorybenchmark
  11. arxiv:2609.35741 · cs.AI
    Shockingly Simple Self-retrospection Improves Agentic Models Without RL
    Jonathan Light, Christopher Zhang Cui, Jeonghye Kim, Roger Creus Castanyer +6

    People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this que

    agentagentic
  12. arxiv:2609.35738 · cs.LG
    Harness Learning Enables Generalizable Test-Time Adaptation
    Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar +5

    A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at h

    agenttool use
  13. arxiv:2609.35734 · cs.CV
    GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
    Kerui Ren, Tao Lu, Linning Xu, Changjian Jiang +4

    Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unse

    memory
  14. arxiv:2609.35732 · cs.AI
    Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models
    Junru Zhu, Shiming Xie, Aime Lu Fan Chen, Xiaoqing Ding +3

    Tool-using agents can fail twice: a required tool can fail, and the agent can then report success without the evidence needed to justify it. Existing benchmarks often entangle this reporting failure with tool selection, recovery, and environment dynamics. We introduce Failure-Transparent Agents (FTA

    agentbenchmark
  15. arxiv:2609.35728 · cs.CV
    FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning
    Ziyao Huang, Zhengkun Rong, Shiyang Qin, Shuang Liang +4

    We present FlowAct-R2, a framework for interactive humanoid video generation that combines continuous multimodal control with proactive agent planning. Our method consists of two coupled components. First, a Streaming Multimodal Reference Diffusion Transformer adapts the pretrained Seedance 2.0 Mini

    humanoidagent
  16. arxiv:2609.35718 · cs.CV
    Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
    Hanoona Rasheed, Mohammed Irfan Kurpath, Bin Ren, Hisham Cholakkal +2

    Frontier general-purpose systems are rapidly expanding beyond visual understanding into capabilities traditionally handled by dedicated computer-vision models. As these capabilities expand, a central question for the computer-vision community is how far this reach extends, and what remains hard. We

    benchmark
  17. arxiv:2609.35717 · cs.RO
    RoboCompiler: Graph-Native Compilation of Closed-Chain Robots for Consistent Modeling, Control, and Simulation
    Mehdi Heydari Shahna, Joongheon Kim, Jouni Mattila

    Robots with kinematic loops, coupled actuators, and changing contacts require consistent models of configuration, motion, force, and dynamics. Yet these interfaces are often reconstructed separately for control and simulation, making closure and actuation consistency difficult to maintain. This pape

    franka
  18. arxiv:2609.35715 · cs.RO
    X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets
    Prithwish Dan, Chenyang Ma, Wei Zhan

    Reinforcement learning (RL) in simulation can train dexterous manipulation policies without robot demonstrations, but training a single generalist policy with task-agnostic rewards faces a severe exploration problem: approaching, grasping, and reorienting diverse objects with many degrees of freedom

    manipulationdexteroussim-to-realgrippergrasp
  19. arxiv:2609.35709 · cs.RO
    Humanoid Loco-Manipulation With Discrete VLA Model
    Wenxin Shao, Siqi Chai, Kun Li, Kerou Zhang +3

    Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, trainin

    vision-language-actionvlavla modelmanipulationhumanoidteleoperation
  20. arxiv:2609.35707 · cs.LG
    ScAn-Bench: Evaluating Scaling Analysis Methodology
    Artin Sermaxhaj, Nastaran Alipour, Donat Sinani, Johannes Hog +3

    Recent progress in machine learning is driven by large-scale foundation models, where scaling laws and finding optimal scaling prescriptions for architecture, data, and hyperparameters are key in advancing the state-of-the-art. Therefore, it is surprising that no systematic study evaluates the metho

    benchmark
  21. arxiv:2609.35706 · cs.AI
    Reinforcing Agentic Creativity in Scientific Ideation with Night Science
    Priyanka Kargupta, Silviu Cucerzan, Shweti Mahajan, Allen Herring +3

    Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to l

    agentic
  22. arxiv:2609.35704 · cs.CV
    DynaTokens: Teaching Dynamics to Camera-Controlled Video Models at Test Time
    Ziqi Ma, Hongqiao Chen, Georgia Gkioxari

    Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle

    world model
  23. arxiv:2609.35703 · cs.LG
    A Unified Uncertainty Representation for Graph Neural Networks via Doubly-Spectral Stochastic Expansion
    Fred Xu, Thomas Markovich, Florence Regol, Yizhou Sun

    Reliable deployment of graph neural networks requires calibration, out-of-distribution (OOD) detection, and robustness to distribution shift, yet existing methods address these needs with separate models and objectives. We model uncertain node embeddings as random graph signals: graph Fourier filter

    benchmark
  24. arxiv:2609.35701 · cs.LG
    MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
    Chang-Wei Shi, Xu Wang, Wu-Jun Li

    The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining

    memory
  25. arxiv:2609.35700 · cs.RO
    LQR-ArUco Fusion: Robust Hierarchical Control for Navigation and Asymmetric Manipulation in Two-Wheeled Robots
    Anupam Chatterjee, Arpita Kumari

    We propose a hierarchical control framework to address severe dynamic instabilities and navigational drift that arise when a two-wheeled inverted pendulum (TWIP) robot attempts asymmetric object manipulation. While two-wheeled platforms are highly manoeuvrable, their constant balancing adjustments m

    manipulation
  26. arxiv:2609.35698 · cs.LG
    Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
    Hanbin Zhou, Shangzhe Li, Alexander Braverman, Weitong Zhang

    We study adversarial imitation learning (AIL), in which an agent learns to imitate expert demonstrations by optimizing a policy against an adversarial reward that distinguishes expert and learner behavior. Historically, reward regularization and entropy-based policy regularization are key components

    agent
  27. arxiv:2609.35694 · cs.AI
    Reasoning with Continuous Latent Diffusion
    Xiang Cheng

    Continuous diffusion generates complete reasoning solutions through iterative refinement in latent space. We introduce Latent Flow Reasoning Models (LFRMs), an ELF-based training and inference recipe. Our experiments show that accurate decoding alone does not ensure strong reasoning performance. We

    iterative refinement
  28. arxiv:2609.35692 · cs.AI
    Report: Progressive Disclosure of Agent Skills
    Guilin Zhang, Kai Zhao, Priyanka Mudgal, Waleed Ammar +3

    Users of Workday's deployed LLM-based agents often request features which can be addressed by defining named procedures, also known as skills, in the LLM context, effectively augmenting agents' capabilities. However, as an agent's skills library grows in size, so does the agent's operational cost. P

    agent
  29. arxiv:2609.35690 · cs.RO
    Agent Priors-guided Policy Learning
    Puming Jiang, Tianrun Hu, Haozhe Du, Yibo Li +3

    Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost b

    grippergraspagent
  30. arxiv:2609.35686 · cs.LG
    Rethinking Circuit Evaluation: Do Circuits Explain Model Errors?
    Li Zhang, Chuqin Geng, Mark Zhang, Chen Yang +3

    Mechanistic interpretability (MI) aims to explain a model's behaviour through analyzing its internal computations; circuit-based explanations aim to isolate these computations with compact subnetworks validated by ablating the rest of the model. We show that circuits validated this way may fail to r

    benchmark
  31. arxiv:2609.35685 · cs.CL
    QuanReview: Offline, Auditable Reconciliation of Human and LLM Span Annotations
    Matteo Musacchio, Juan Cruz Giner Pulero, Isabel Castañeda, Naomi Couriel +3

    Structured span annotations, such as quantities with their units, uncertainty modifiers, and event classes, are expensive to create and hard to keep trustworthy once language models enter the loop. We present QuanReview, an open-source system for auditing and correcting such annotation layers. QuanR

    benchmark
  32. arxiv:2609.35673 · cs.CV
    FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
    Thanh-Long V. Le, Steven Walton, Seunghyun Yoon, Branislav Kveton +3

    Tool-based image editing (image retouching) is commonly formulated with autoregressive multimodal large language models (MLLMs) that sequentially generate reasoning, tool selections, and parameter values. In this work, we present a novel approach to tool-based image editing by framing the task as a

    memorypost-training
  33. arxiv:2609.35671 · cs.AI
    PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents
    Yangqin Jiang, Lingrui Xu, Chao Huang

    Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeate

    agent
  34. arxiv:2609.35664 · cs.AI
    MS-GLA: Multi-Scale Gated Linear Attention for Addressing Representational Bottlenecks via Multi-Temporal Resolution
    Prasoon Dev, Anirudh Sankar, Vasudeva Varma

    Gated Linear Attention (GLA) Transformers advance linear recurrent models through data-dependent gating, but face a core limitation: the fixed-capacity memory matrices across all heads operate at a single temporal resolution, where each token is processed individually, forcing them to simultaneously

    memorylong-contextbenchmark
  35. arxiv:2609.35658 · cs.CV
    Many Eyes, One World: Feed-Forward 3D Reconstruction from Mixed Cameras
    Qiaoge Li, Yifan Zhan, Haijun Yang, Haiyang Liu +2

    Real-world capture is heterogeneous: perspective, fisheye, and $360^\circ$ panoramic images can coexist within a single reconstruction task, yet most feed-forward 3D reconstruction models assume perspective imagery and a uniform input representation. Recent models handling several camera types are e

    benchmark
  36. arxiv:2609.35657 · cs.AI
    CMDO: A Cognitive Memory-Driven Optimization Algorithm for Adaptive Population-Based Search
    Mohammed Yusuf Mujawar, Shahram Rahimi, Noorbakhsh Amiri Golilarz

    Population-based optimization methods often use previous search information through successful solutions, parameter adaptation, or operator performance, but they rarely retain the context in which a search behavior succeeded or failed. We introduce Cognitive Memory-Driven Optimization (CMDO), a deri

    memorybenchmark
  37. arxiv:2609.35652 · cs.RO
    MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
    Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao +5

    Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base act

    manipulationliberoworld model
  38. arxiv:2609.35645 · cs.LG
    CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
    Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin +1

    Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how

    agentbenchmarkevaluation framework
  39. arxiv:2609.35641 · cs.CV
    Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
    Shuyue Stella Li, Xiaochuang Han, Yulia Tsvetkov, Luke Zettlemoyer

    Precise instruction following in image generation, such as satisfying object counts and spatial relations, remains an open challenge at least in part because it is learned using unreliable reward models such as object detectors and vision-language models. We introduce Verifiable Visual Rewards (VVR)

    post-trainingbenchmark
  40. arxiv:2609.35639 · cs.AI
    GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation
    Yuchen Sun, Jinjin He, Sinan Wang, Bo Zhu

    Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. We introduce GPUPhysBench, a benchmark of 50 tasks testing whether coding agents can meet these demands. Tasks co

    benchmark
  41. arxiv:2609.35629 · cs.LG
    SANTA++: Sampling Attention through Representative Keys
    Kyle Lee, Christian Z. Pratt, Ruoyu Fang, Heekyung Lee +3

    Attention often concentrates on a small subset of tokens in the context, but which subset matters changes from one query to the next. To exploit this changing structure, we introduce SANTA++, a training-free stochastic attention method that uses representative keys for memory-efficient selection wit

    memoryretrieval-augmented
  42. arxiv:2609.35627 · cs.CL
    Can LLMs Value the Right Evidence? Evidence-Value Misalignment in Dynamic Medical Diagnosis
    Kehua Feng, Yunsheng Lu, Yitong Qiao, Tiantian He +4

    A correct diagnosis reached from insufficient or misleading evidence can pose a clinical hazard, yet outcome-based accuracy may reward such lucky guesses. We call this mismatch between diagnostic decisions and the value of available evidence Evidence-Value Misalignment (EVM). To disentangle evidenti

    benchmark
  43. arxiv:2609.35622 · cs.LG
    Elicitation and Decision Geometry in Single-Index Bandits
    Sakshi Arya, Cheng Soon Ong

    We study two-arm contextual bandits with arm-specific single indices and a shared unknown monotone link. Monotonicity makes the optimal action depend only on the contrast between the index directions, hence arm-specific reward functions need not be estimated. We introduce Natural Boundary Learning (

    benchmark
  44. arxiv:2609.35621 · cs.LG
    Cartridges++: KV Cache Compression without Off-Context Derailment
    Sonia Laguna, Joao Monteiro, Marco Cuturi, Pierre Ablin +1

    Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-value (KV) cache balloons. Compressed KV (CKV) representations aim to mimic the cache of a document and are typically computed once and for all,

    memorylong-context
  45. arxiv:2609.35620 · cs.LG
    Attention Graphons: A Graph Limit Perspective on Graph Transformers
    Caio F. Deberaldini Netto, Moshe Eliasof, Luana Ruiz

    Graph Transformers produce, for each attention head, a dense $n\times n$ matrix of learned pairwise interactions. We ask a fundamental question: do these attention-induced graphs converge to a stable limit object as $n$ grows, or does the learned interaction pattern remain unstructured and size-depe

    benchmark
  46. arxiv:2609.35619 · cs.RO
    CollisionSplatting: Collision-Aware Motion Planning in 3DGS Scenes with Image-Conditioned Objectives and Adjustable Conservatism
    R. Khorrambakht, Joaquim Ortiz-Haro, Stephan Weiss, Ludovic Righetti

    Incorporating dense visual information into motion planning remains challenging, as geometric planners rely on abstracted scene representations that discard visual richness, while learned visual models often lack geometric interpretability and computational efficiency. This paper introduces Collisio

    manipulation
  47. arxiv:2609.35616 · cs.CV
    EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold
    Junjie Chen, Fei Wang, Kun Li, Yiqi Nie +4

    Interactive 3D head generation requires coordinated speaking and listening motion that responds to an evolving conversation. Existing generators use incoming observations as context but keep their parameters fixed, leaving conversational patterns unused as a learning signal. We introduce EvolvingAva

    benchmark
  48. arxiv:2609.35615 · cs.LG
    Behavioral Foundation Models for Quality Diversity
    Nazim Bendib, Nicolas Perrin-Gilbert, Olivier Sigaud

    Behavioral Foundation Models (BFMs) are an emerging paradigm in reinforcement learning, playing a role analogous to large language models in natural language processing: they have shown remarkable versatility, enabling zero-shot performance, fast imitation, and online adaptation, all by exploiting t

    manipulationbenchmark
  49. arxiv:2609.35614 · cs.LG
    EvE: An Alternate Optimizer to Adam
    Shashank Raj, Kalyanmoy Deb

    Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutiona

    benchmark
  50. arxiv:2609.35611 · cs.LG
    On-Policy Self-Distillation for Multi-Turn Image Editing
    Liangbing Zhao, Le Zhuo, Mohamed Elhoseiny

    Instruction-based image editing has achieved strong performance in single-turn settings, yet practical editing is often iterative, with each instruction applied to the output of the previous turn. We find that existing editing models degrade rapidly under recursive editing and attribute this failure

    benchmark
  51. arxiv:2609.35606 · cs.AI
    TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science
    Chutong Yang, Xiyuan Zhang, Yu Huang, Boran Han +6

    Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether m

    agentagenticbenchmark
  52. arxiv:2609.35603 · cs.LG
    Control-Geometry Straightening for Sampling-Based Latent Planning
    Ziang Fu, Ning Ning

    Joint-embedding predictive architectures enable planning with latent world models, but accurate transition prediction alone does not ensure that the planning objective is easy to optimize. We introduce Control-Geometry Straightening (CGS), a single auxiliary loss that learns planner-friendly represe

    world model
  53. arxiv:2609.35596 · cs.AI
    SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
    Saswat Das, Parvati Viswanathan, Daniel Donnelly, Chang Huang +2

    Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates

    memoryagentllm agentagenticself-evolvingbenchmark
  54. arxiv:2609.35586 · cs.AI
    IMC-CLINIC: Coupled Loss-Informed Newton Iterations for Clipping in Analog In-Memory Computing
    Yung-Chin Chen, Chia-Yu Chen, Naveen Verma

    Analog in-memory computing (IMC) offers a promising path toward energy-efficient large language model (LLM) inference by executing matrix multiplications (MatMul) directly within memory arrays in the analog domain. Its efficiency, however, comes with an additional source of error: limited-precision

    memory
  55. arxiv:2609.35583 · cs.CV
    What Paired Evaluations Reveal under Visual Perturbations
    Yongda Wei, Chen Zhang, Yifei Wang, Xinyu Wang +3

    Robustness evaluation must examine diverse visual perturbations, while benchmarks cover only some real-world conditions and physical testing is costly. Paired evaluations link clean and perturbed predictions for the same image, capturing changes in correctness, confidence, and acceptance beyond aggr

    benchmark
  56. arxiv:2609.35581 · cs.LG
    QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks
    Pranav Gupta

    We introduce QC-Stark, a benchmark for evaluating large language models (LLMs) on 11 quantum computing (QC) tasks, spanning circuit construction, debugging, compilation, error correction, and simulation. Across 2,750 evaluations (10 models $\times$ 11 tasks x 5 difficulty levels x 5 seeds), we find

    benchmark
  57. arxiv:2609.35578 · cs.AI
    FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
    Bowen Yang, Jingbo Zhou, Qinghong Miao, Hua Wu

    Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat

    memory
  58. arxiv:2609.35576 · cs.LG
    Share-Borne AI Virus: Memory-Hopping Attacks Across LLM Agents
    Sidharth Pulipaka, Ansh Sharma, Stanislau Hlebik, Leonidas Raghav +3

    Large language models are increasingly deployed as stateful assistants that retain information across interactions and use tools to read, modify, and create persistent artifacts. As these artifacts are shared between users, they form an indirect communication channel between otherwise independent as

    persistent memoryllm agent
  59. arxiv:2609.35575 · cs.RO
    F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement
    Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren +12

    The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered fai

    vision-language-actionmanipulationsim-to-realagentself-improvement
  60. arxiv:2609.35570 · cs.RO
    EdgeVLN: Runtime-Aware Deployment Ready Quantized Vision Language Navigation Model
    Rithvik Jonna, Man Namgung, Aakash Gurram, Tinoosh Mohsenin

    Vision-language navigation (VLN) models perform well but target compute-rich platforms, limiting deployment on memory- and power-constrained robotic edge devices. Compression alone does not establish whether a VLN model fits the memory, latency, and energy budgets of an edge platform while preservin

    memory
  61. arxiv:2609.35569 · cs.LG
    Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions
    Tianyao Shi, Xipeng Shen, Yi Ding

    Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. Yet these dimensions are largely evaluated in isolation, leaving it unclear when and how they lead to different optimization decisions. We present PRISM,

    embodied
  62. arxiv:2609.35568 · cs.LG
    From Experience to Expertise: Adoption-Aware Memory Learning for Data-Scarce NPU Kernel Synthesis
    Longxiao Fan, Tao Zhang, Han Yan, Jiajun Li +4

    High-performance kernels underpin efficient accelerator execution but require expert tuning and lengthy manual optimization cycles. LLM coding agents promise automation, yet their CUDA knowledge transfers poorly to data-scarce domain-specific architectures (DSAs) such as NPUs, whose execution models

    memoryexternal memoryagentself-improvingpost-training
  63. arxiv:2609.35564 · cs.AI
    Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
    Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir +36

    Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unsee

    benchmark
  64. arxiv:2609.35561 · cs.AI
    RSI-Master: Structuring Experiments to Guide Autonomous Model Improvement
    Yaxin Du, Xiyuan Yang, Zhifan Zhou, Yujie Ge +9

    Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit

    agentself-improvementpost-trainingbenchmark
  65. arxiv:2609.35560 · cs.CV
    WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon
    Haiyu Zhang, Wenqiang Sun, Tengfei Wang, Junta Wu +4

    Interactive world models require responding in real time to versatile controls and maintaining long-horizon consistency. However, modeling heterogeneous controls remains difficult, while explosive contexts and unstable distillation impede achieving both long-horizon consistency and real-time respons

    world modelmemory
  66. arxiv:2609.35559 · cs.AI
    From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining
    Kangcheng Deng, Hui Cai, Jiacheng Lu, Chester Zhongshu Qian +5

    Inference scaling has been shown to improve large language model (LLM) performance, and this principle naturally extends to autonomous LLM agents through increased search budgets, which we refer to as *search scaling*. Although prior work has characterized the mechanisms, scaling behavior, and perfo

    agentllm agent
  67. arxiv:2609.35557 · cs.AI
    The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent
    Shobhan Roy

    The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program

    agent
  68. arxiv:2609.35553 · cs.LG
    Simplex Diffusion Models
    Justin Deschenaux, Alexandre Galashov, Andrew Campbell, Li Kevin Wenliang +3

    Diffusion models have revolutionized generative modeling for continuous data through the gradual refinement of a belief state. This iterative refinement has not yet carried over to discrete diffusion models, which discard uncertainty at intermediate steps through categorical sampling (information co

    iterative refinement
  69. arxiv:2609.35551 · cs.AI
    BaRe-Mem: Bayesian Reliability Memory for Robust and Adaptive Agent Consultation
    Peilin Feng, Zhengyang Huang, Soujanya Poria

    In multi-agent systems, reliable consultation is challenging because advisor capabilities vary across tasks, and misleading information can make consultation worse than autonomous reasoning. We introduce BaRe-Mem, an online Bayesian reliability memory for multi-agent consultation. It estimates advis

    memoryagentmulti-agentagent systembenchmark
  70. arxiv:2609.35549 · cs.AI
    RareDx: Controlled Knowledge Integration and Graph-Grounded Policy Optimization for Rare-Disease Diagnosis
    Bo Zhang, Yuchen Wang, Dongbai Li, Matthew Yu Heng Wong +5

    Rare-disease diagnosis is a long-tail reasoning problem: phenotypes are incomplete, individual disorders are sparsely documented, and relevant evidence is distributed across ontologies, gene annotations, and biomedical text. Language models consequently favor common conditions, miss rare candidates,

    post-trainingbenchmark
  71. arxiv:2609.35545 · cs.LG
    Graph World Models for Constrained Epidemic Policy Planning
    Yiqi Su, Rashed Shelim, Lingyi Wang, Walid Saad +1

    Epidemic policy planning often requires coordination between geographical regions, taking into account mobility-driven spillovers and how to make use of limited resources. Existing methods either lack action-conditioned models of coupled dynamics or cannot guarantee per-period feasibility. We presen

    world modelaction-conditioned
  72. arxiv:2609.35540 · cs.AI
    Continuous Context Management
    William Hoy, Jingxuan Fan, Nurcin Celik, Xu Pan

    Long-horizon large language model (LLM) agents commonly retain their complete interaction history until compaction is triggered at a predefined threshold. We study Continuous Context Management (CCM), which performs compaction at every turn to prevent interaction history from accumulating in the act

    memoryagent
  73. arxiv:2609.35539 · cs.CV
    Learning to Reason with Persistent Object States for Video Instance Segmentation
    Yongxue Xu, Boxue Yang, Ziqian Liu, Shaoqiu Zhang +2

    Video segmentation models maintain object identities by carrying instance information across frames. Under prolonged occlusion, reappearance, or interactions between similar instances, however, an unreliable update can overwrite a valid history and cause persistent identity drift. We introduce POSRe

    persistent statebenchmark
  74. arxiv:2609.35537 · cs.LG
    Optimal Networks for Agentic Information Aggregation
    MohammadHossein Bateni, Zahra Hadizadeh, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz +1

    We study information aggregation in the networked learning model introduced by Kearns, Roth, and Ryu (SODA 2026). There is a fixed distribution over $d$ features and a common label. Agents learn in topological order on a directed acyclic graph. Each observes a subset of the features and its parents'

    agentagentic
  75. arxiv:2609.35536 · cs.CV
    Look Before You Judge: Training-Free Region Mining for Grounded and Explainable Deepfake Detection
    Chia-Ling Chen, Yu-Ting Ta, Jian-Yu Jiang-Lin, Tai-Ming Huang +7

    Multimodal large language models (MLLMs) can explain deepfake verdicts in natural language, but such explanations are not necessarily visually grounded in the visual evidence underlying the prediction. A model may describe plausible artifacts inferred from language priors rather than from image evid

    manipulation
  76. arxiv:2609.35532 · cs.AI
    ARISE: Adapting to Evolving Capability Gaps in Agentic Reinforcement Learning
    Kun Feng, Yuchen Fang, Yiyang Tan, Shuqi Gu +5

    As a long-horizon agent improves through experience, previously observed weaknesses may recede while new limitations emerge, continually changing what it still needs to learn. Yet the learning process often remains tied to a static view of these needs: fixed behavioral criteria and training prioriti

    agentagenticagent benchmarkbenchmark
  77. arxiv:2609.35530 · cs.CV
    AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
    Yuta Oshima, Ku Onoda, Yusuke Iwasawa, Masahiro Suzuki +2

    Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally past

    agentagenticbenchmarkevaluator
  78. arxiv:2609.35525 · cs.LG
    Deep Epistemic Value Functions for Optimistic Exploration
    Leander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas Krause

    Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central chall

    agent
  79. arxiv:2609.35517 · cs.LG
    Reward-Aligned Reweighting for On-Policy Distillation
    Haofeng Xu, Junwei Su, Lansong Diao, Wenchao Zhou +1

    On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's

    benchmark
  80. arxiv:2609.35515 · cs.LG
    MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?
    Zihan Yu, Jiadong Zhang, Jialin Cheng, Jingtao Ding +1

    Scientific discovery requires not only recovering mathematical laws that describe observable behavior, but also identifying the mechanisms that generate them. Existing benchmarks for symbolic regression and scientific agents primarily evaluate phenomenal-law recovery, leaving mechanism discovery lar

    benchmark
  81. arxiv:2609.35507 · cs.CV
    ReVA: A Scene-Centric Dataset Beyond Repetition for Remote Sensing Video Question Answering
    Zhen Yao, Likai Wang, Yuming Yang, Zhihao Zheng +8

    Multimodal Large Language Models (MLLMs) have demonstrated remarkable advances in remote sensing. However, existing remote sensing multimodal reasoning benchmarks exhibit two critical limitations: they rely on (i) template-driven questions, which causes repetitive questions; and (ii) static images t

    benchmark
  82. arxiv:2609.35505 · cs.LG
    An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
    Shangzhe Li, Yuxiao Yang, Tianrun Yu, Kaixiang Zhao +3

    We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that

    benchmark
  83. arxiv:2609.35504 · cs.CV
    SolveEdit: Benchmarking Visual Problem Solving in Generative Models
    Wenjie Shu, Yexin Liu, Harold Haodong Chen, Xuerui Qiu +8

    Machine intelligence is often evaluated through abstract reasoning problems, yet many real-world problems are visual, such as arranging objects, repairing layouts, or tracing routes. Solving these problems requires understanding a scene, inferring what must change to achieve a goal, and realizing th

    benchmark
  84. arxiv:2609.35502 · cs.LG
    Structured Latent Modeling for Supervised Multimodal Information Decomposition
    Wanting Huang, Sanvesh Srivastava, Weiran Wang

    Multimodal prediction relies on diverse forms of evidence: information repeated across modalities, cues specific to a single source, and complex cross-modal dependencies that emerge only when inputs are considered together. While recent methods promote richer interactions, they lack a principled way

    benchmark
  85. arxiv:2609.35501 · cs.LG
    SRHarness: A Harness for Agentic Symbolic Regression
    Zihan Yu, Shixuan Zhou, Hao Huang, Jingtao Ding +1

    Recent agentic symbolic regression approaches increasingly rely on large language models to analyze data, select scientific operations, and refine hypotheses over long search trajectories. In such systems, performance depends not only on the underlying model and search strategy, but also on the runt

    agentic
  86. arxiv:2609.35497 · cs.CV
    Sprout: Building Dynamic Memory While Reasoning for Agentic Video Understanding
    Wei Chen, Xuanyu Zheng, Yancheng Long, Haoyang Xu +5

    Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by seve

    memoryagentagenticbenchmark
  87. arxiv:2609.35493 · cs.RO
    Terrain-Aware Autonomous Planetary Exploration for Exteroceptive-Proprioceptive Mapping with Quadruped Scouts
    Alberto Sanchez-Delgado, João Carlos Virgolino Soares, Victor Barasuol, Claudio Semini

    Autonomous planetary exploration requires robots to navigate unknown, uneven terrain while assessing risk, traversability, and energetic cost. Quadruped scouts are well suited for this task because they can traverse irregular surfaces and gather mobility-relevant information during locomotion. This

    quadruped
  88. arxiv:2609.35491 · cs.CV
    From Scores to Samples: Elastic Forcing for Autoregressive Video Generation
    Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An +4

    Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-traini

    post-training
  89. arxiv:2609.35486 · cs.CV
    Who Is Left of Whom? Tracing Spatial Evidence and Role Binding in Relative-Position Reasoning
    Yingjin Song, Denis Paperno, Albert Gatt

    High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this proce

    benchmark
  90. arxiv:2609.35482 · cs.RO
    ForVis: An In-Field Dataset and Benchmark for VIO Using Under-Canopy UAV Flights in Forests
    Arman Kiani, Masoud Ataei, Elvis Gyaase, Jeffrey Eiyike +3

    Visual-inertial Simultaneous Localization and Mapping (VI-SLAM) for UAVs remains difficult to evaluate in real forest environments, where motion, illumination changes, repetitive vegetation, and vibration can all affect estimation. We present ForVis, an in-field dataset and benchmark for evaluating

    benchmark
  91. arxiv:2609.35479 · cs.RO
    Robot Tool Design from Scratch via Behavior-Aware Hierarchical Optimization
    Yinghan Chen, Xiyao Tian, Yizan Dai, Yuyang Li +1

    The ability to design a tool for a task marks a level of intelligence beyond merely understanding, selecting, or using one. Existing methods for robotic tool design typically optimize a tool's continuous shape and action within a structure that is prescribed or generated beforehand, so the structure

    tool-use
  92. arxiv:2609.35477 · physics.optics
    Plasmon-Enhanced Second-Harmonic Generation in Atomically Thin Crystalline Silver Nanostructures
    Saad Abdullah, Philipp K. Jenke, Andrew P. Weber, Álvaro Rodríguez Echarri +7

    The intrinsically weak nonlinear optical response of existing materials, further constrained by symmetry-forbidden second-order processes in centrosymmetric media, severely limits efficient frequency conversion in deeply subwavelength, ultrathin volumes. Addressing this challenge is crucial for the

    quantum photonic
  93. arxiv:2609.35476 · cs.RO
    CoBrush: A Hierarchical Planning Framework for Human-Robot Co-Painting
    Dantong Qin, Yike Guo, Qinlin Liu, Alessandro Bozzon +1

    Embodied co-painting requires a robot to repeatedly update a shared physical canvas while human intent evolves over interaction. Existing reference-driven painters or reactive assistants are typically optimized for single-shot rendering or sketch completion, limiting their ability to sustain coheren

    embodied
  94. arxiv:2609.35473 · cs.LG
    Handwritten Text Recognition Lives in the High-Pixel Variance Subspace
    Carlos Garrido-Munoz, Jorge Calvo-Zaragoza

    In self-supervised pretraining for Handwritten Text Recognition (HTR), pixel reconstruction methods outperform contrastive methods, unlike in natural-image classification. We argue that this difference follows from where discriminative signal lies in pixel space: for HTR, it is concentrated in high-

    benchmarkevaluation protocol
  95. arxiv:2609.35469 · cs.RO
    Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching
    Chenyu Zhang, Yuhang Cao, Daru Du, Yingxi Lu +7

    Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action token

    vision-language-actionvlamanipulationbenchmark
  96. arxiv:2609.35465 · cs.LG
    Tetra: Serving Leech-Lattice Quantized LLMs at 2.7 Bits per Parameter
    Pier-Jean Malandrino

    Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on

    memory
  97. arxiv:2609.35463 · cs.AI
    A.D.A.M.O. (Agent for language-Driven Actions with Multimodal Observations): A Visual-Symbolic Framework for Virtual Humans
    Alessandro Emmanuel Pecora, Stefano Calzolari, Francesco Strada, Andrea Bottino

    Creating believable vh requires the coherent integration of perception, reasoning, and action mediated by language. A central challenge is to combine these components into a control loop grounded in interactive 3D environments. To this end, we present A.D.A.M.O. (Agent for language-Driven Actions wi

    world modeltool calling
  98. arxiv:2609.35462 · cs.LG
    CLIMB: A Clinical Multimorbidity Benchmark for Diagnosing Co-occurring Conditions through Multiturn Conversations
    Yusuf Kesmen, Aniruddha Mukherjee, Yena Chang, David Sasu +6

    Patients often have several co-occurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multi-turn interaction and multi-label diagnosis. We introduce CLIMB, a benchm

    benchmark
  99. arxiv:2609.35461 · cs.CL
    AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic
    Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian, Muhammad Alqurishi

    As Large Language Models (LLMs) continue to scale both in size and capabilities, their proficiency in the Arabic Language has seen significant advancement. However, a critical gap remains: the extent of their factual knowledge and cultural sensitivity to the diverse Arabic-speaking world remains lar

    benchmarkevaluation framework
  100. arxiv:2609.35456 · cs.AI
    AutoBCI: Forecast-Guided Agentic Neural Architecture Discovery for EEG-Based Brain--Computer Interfaces
    Muyun Jiang, Yi Ding, Wei Zhang, Jinbo Chen +8

    EEG-based brain-computer interfaces support a broad range of applications, yet designing decoding architectures that perform well across diverse tasks remains challenging. We introduce AutoBCI, an agentic framework in which a Designer Agent and a Forecaster Agent support the discovery and selection

    agentagentic
  101. arxiv:2609.35450 · cs.RO
    Uni-VLaT: Whole-Body Tactile Adaptation of VLA Policies for Humanoid Loco-Manipulation
    Zihao Wang, Shutong Liu, Siqi Zheng, Liu Cao +4

    Physical contact often determines how a humanoid should respond during loco-manipulation, yet vision and proprioception alone are often insufficient to characterize physical interaction, especially when the contact region is occluded. Unlike sparse force or torque measurements at predefined regions,

    vision-language-actionvlamanipulationhumanoidtactile
  102. arxiv:2609.35445 · cs.LG
    NeuronSifter: Intervention Planning in CNS Microenvironments
    Haowei Xu, Wanyi Fu, Hongbin Han, Zhaoheng Xie

    Prioritizing central nervous system (CNS) interventions requires predicting how a dose, route, and schedule act on a partially observed microenvironment, then choosing the measurement that would change the decision. Action-conditioned predictors reduce a regimen to an identity token or a scalar expo

    action-conditionedbenchmark
  103. arxiv:2609.35440 · cs.LG
    SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
    Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li

    Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper layer. Conventional local learning removes this update locking by training each module to predict t

    memory
  104. arxiv:2609.35433 · cs.LG
    ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
    Yihang Chen, Yuanhao Ban, Cho-Jui Hsieh

    Reinforcement learning from verifiable rewards (RLVR) frequently reuses rollouts across multiple policy updates, increasing the mismatch between the current policy and the data-generating policy. We identify a sign-dependent gradient starvation problem in clipped policy optimization: clipping suppre

    benchmark
  105. arxiv:2609.35432 · cs.RO
    Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence
    Hongcheng Gao, Jingjing Zhou, Zelin Zheng, Shijia Ge +9

    Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirem

    vision-language-actionagentself-evolving
  106. arxiv:2609.35431 · cs.RO
    Memory in the Sky: Low-Altitude Question Answering with Multi-Agent Memory Aggregation
    Chengyang Li, Yujie Wan, Shuai Wang, Kejiang Ye +5

    This paper studies low-altitude question answering (LAQA), in which distributed unmanned aerial vehicle (UAV) memories are aggregated at a ground server to answer questions about observations over a long horizon. Unlike conventional resource allocation based on sensing, communication, control, or co

    memoryagent memorymulti-agentagent systembenchmark
  107. arxiv:2609.35427 · cs.LG
    LLMs are General Asynchronous Agents
    George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy +3

    Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform a

    embodiedmemoryautonomous agentembodied agenttool calling
  108. arxiv:2609.35426 · cs.LG
    Frontier Learning: Training LLM Reasoners at the Edge of Capability
    Robin Faro, Shyam Sundhar Ramesh, Ilija Bogunovic, Aurelien Lucchi

    Reinforcement Learning-based post-training of Large Language Models (LLM) has been successfully applied to improve their reasoning capabilities. Existing pipelines primarily finetune LLMs on a fixed pool of problems specified prior to training using the GRPO loss. This is fundamentally limiting, as

    post-training
  109. arxiv:2609.35416 · cs.CV
    When Should the Count Change? Learning State Maintenance for Causal Video Counting
    Pengyiang Liu, Dongyue Lyu, Junbo Niu, Zhongyue Shi +2

    Continuous video counting requires distinguishing new observations from new objects or completed events. We introduce StaMina (State Maintenance), which learns to maintain counting state through state-conditioned updates. Recurrent visual context supports recognition; learned transitions maintain vi

    benchmark
  110. arxiv:2609.35412 · cs.AI
    Self-Adapting Group of Experts for Multi-Agent Reasoning
    Mohammad Atif Quamar, Nurbek Tastan, Karthik Nandakumar, Junpei Komiyama

    Multi-agent systems bring together language model agents with different roles to propose, review, and refine solutions. Each agent's response depends on its model's capabilities, the reasoning strategy defined by its system prompt, and the information in its input context. Existing frameworks often

    agentmulti-agentagent systembenchmark
  111. arxiv:2609.35409 · cs.AI
    AwarenessBench: Assessing Cognitive Capabilities of Language Models
    Xiaojian Li, Rongwu Xu, Tianyun Zhang, Yue Wang +8

    As language models (LMs) exhibit increasingly consciousness-like behaviors, evaluating their cognitive abilities becomes essential. We introduce AwarenessBench, the first comprehensive benchmark for assessing the cognitive abilities of LMs in four dimensions: metacognition, self-awareness, social aw

    benchmark
  112. arxiv:2609.35408 · cs.AI
    "Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables
    Yage Zhang, Yukun Jiang, Yang Zhang

    Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the ed

    agentbenchmark
  113. arxiv:2609.35407 · cs.CV
    BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion
    Wanjiang Weng, Yongliang Wu, Xiaofeng Tan, Xingyu Zhu +2

    Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poo

    self-correction
  114. arxiv:2609.35402 · cs.LG
    Persistent Partners Raise Prices Among Learning Agents
    Paul-Peter Arslan, Yubin Kim, Xiao Xiao

    When pricing agents meet repeatedly on a platform, the platform decides who faces whom. We ask whether that choice moves the prices the agents learn, and whether a rise comes with learned punishment. In a pre-registered randomised experiment in the Bertrand duopoly of Calvano et al., each agent's pr

    agent
  115. arxiv:2609.35400 · cs.AI
    Structural Alignment for Reliable Industrial AI: Bridging Physical Reality, Data, Models, and Human Intent
    Lizhi Xiao, Sihong Wu, Victoria Xiao, Yiqiao Song +4

    Artificial intelligence is increasingly deployed in critical industrial domains, including healthcare, energy grids, subsurface exploration, where failures can have severe consequences for human safety, system stability, and economic outcomes. Yet AI is still evaluated primarily through benchmark ac

    benchmark
  116. arxiv:2609.35394 · cs.CV
    Rethinking Visual Token Compression for Video Large Language Models: A Simple Yet Strong Baseline
    Xiao Zhang, Wang Zeng, Sheng Jin, Wentao Liu +2

    Video Large Language Models (Video LLMs) have achieved remarkable progress in video understanding, but their inference efficiency is constrained by the large number of visual tokens produced by long videos. Recent video token compression methods increasingly introduce sophisticated strategies for to

    benchmark
  117. arxiv:2609.35390 · cs.LG
    Inductive Feedback for Mixed-Policy Distillation
    Amir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang

    Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, whi

    agenticpost-trainingbenchmark
  118. arxiv:2609.35381 · cs.AI
    MCP Error Messages Written for Developers Hurt the Most Capable Agents Most
    Xiaonan Xu, Wenjing Wu

    Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 erro

    agentleaderboard
  119. arxiv:2609.35378 · cs.AI
    Multilinguality in Hybrid Attention LLMs
    Lucas Bandarkar, Junlin Hu, Chenyuan Yang, Mohsen Fayyaz +1

    In response to the growing demand for long sequences in agentic and reasoning use cases, many state-of-the-art LLMs combine multiple variants of attention to mitigate the quadratic complexity of traditional softmax attention. These hybrid attention LLMs aim to balance the strengths and limitations o

    agentic
  120. arxiv:2609.35375 · cs.RO
    From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations
    Bangjun Wang, Longyan Wu, Yukun Wei, Shenghe Shao +7

    Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although cu

    manipulationworld modeltool use
  121. arxiv:2609.35367 · cs.CL
    From Input to Output: A Flexible Agent for Dual-End Interpretation of Sparse Autoencoder Features
    Dewen Liu, Zixuan Li, Jonathan Pan, Zhao Wu +3

    Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods characterize input-side activation patterns and output-side intervention effects, yet often leave their functional connection implicit, while

    agentagentic
  122. arxiv:2609.35366 · cs.AI
    Planarian: Managing Agent State with Statepoints
    Jinnan Guo, Hao Mark Chen, Kapil Vaswani, Andrew Paverd +1

    LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or re

    agentllm agent
  123. arxiv:2609.35362 · cs.LG
    d-OPD: Future-Aware On-Policy Distillation for Block Diffusion Language Models
    Ruitao Liu, Qinghao Hu, Song Han

    Large language models (LLMs) typically generate text autoregressively (AR), predicting one token at a time. Block diffusion language models (dLLMs) instead generate blocks sequentially while denoising multiple tokens in parallel within each block, offering a promising way to accelerate generation. R

    benchmark
  124. arxiv:2609.35357 · cs.AI
    Do Coding Agents Reuse Existing Code or Reinvent the Wheel?
    Dongsheng Ma, Sizhe Wang, Xinyi Huang, Zhengren Wang +4

    Coding agents are increasingly deployed for iterative development on real repositories, yet existing evaluation barely answers a basic question: \emph{do coding agents reuse existing code or reinvent the wheel?} The question matters: every duplicated implementation is a fix applied twice and agents

    benchmark
  125. arxiv:2609.35349 · cs.LG
    Quasi Linear Kernel Attention with Infinite Capacity
    Nicolaj Rux, Johannes Hertrich, Sebastian Neumayer

    The evaluation cost of transformers with softmax attention scales quadratically with sequence length. Kernel attention addresses this by replacing softmax with a more general kernel function. In this paper, we aim to identify kernels that retain the expressivity of attention while enabling quasi lin

    benchmark
  126. arxiv:2609.35347 · cs.LG
    Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
    Xin Li, Hao Jiang, Xin Gao, Annan Wang +5

    Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists t

    benchmark
  127. arxiv:2609.35342 · cs.AI
    Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
    Riccardo Porcedda

    The appearance of Jev marked the era of System One Models, foundation models that return structured decisions with probability distributions rather than text. Aside from low cost and great speed, Jev's central promise is that these probabilities are calibrated: such claim is not backed by any public

    benchmark
  128. arxiv:2609.35341 · cs.CV
    Generative Uncertainty as a Self-supervised Signal for Semantic Similarity Learning
    Enrico Pallotta, Sina Raoufi, Lars Doorenbos, Gianni Franchi +1

    Evaluating semantic similarity between videos is a fundamental challenge in computer vision, essential for tasks ranging from out-of-distribution (OOD) detection to video retrieval. However, defining and labeling video similarity is notoriously difficult and expensive due to the complex spatio-tempo

    v-jepa
  129. arxiv:2609.35336 · cs.CV
    TMCS: Tool-Grounded Multi-Agent Reasoning for Compositional Chemical Problem Solving
    Shengqin Wang, Jie Jin, Yu Cheng, Yihang Chen +3

    Despite the promise of Large Language Models (LLMs) in computational chemistry, rigorous combinatorial chemistry problems remain difficult because they require quantitatively constrained molecular modification, candidate validation, and systematic revision after failed attempts. Existing tool-augmen

    multi-agentagent frameworktool use
  130. arxiv:2609.35335 · cs.LG
    Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation
    Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann

    Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level informatio

    benchmark
  131. arxiv:2609.35333 · cs.LG
    Scalable In-Context Reinforcement Learning with Recurrent Algorithm Distillation
    Yuanqing Ma, Zhenrui Zheng, Chenjun Xiao

    Algorithm Distillation (AD) has demonstrated the remarkable ability of Transformers to perform in-context reinforcement learning without explicit weight updates. However, capturing long-term learning progress necessitates expansive context windows, which incur prohibitive memory costs and limit scal

    memory
  132. arxiv:2609.35328 · cs.AI
    Hyper Algorithm Design Agent: Evolving Learnable Optimizer from Zero
    Zipei Yu, Yue-Jiao Gong, Zeyuan Ma, Yuncheng Jiang +1

    Meta-Black-Box Optimization (MetaBBO) is one of the highlights in the recent AI for Optimization trend. This paradigm's bi-level workflow leverages the learnable algorithm design policy at meta level to ensure the performance and generalization improvement on the low-level optimization task. While M

    agentagent frameworkself-improvement
  133. arxiv:2609.35324 · physics.optics
    Contactless Continuous-Variable Quantum Optical State Conditioning With Classical MmWave Phase
    Niloy Ghosh, Sarang Pendharker

    This paper establishes the Heisenberg's picture framework for seamless mmWave-to-photonic data transduction. Based on this framework, contactless modulation of continuous-variable (CV) quantum optical states with digitally modulated classical mmWave beams is shown for the first time. Our analysis re

    quantum photonic
  134. arxiv:2609.35319 · cs.LG
    Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
    Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao +7

    On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for cor

    autonomous agent
  135. arxiv:2609.35318 · cs.RO
    DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library
    Youhui Wang, Yunzhu Li, Li Fei-Fei, Jiajun Wu +1

    Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving arti

    manipulationdexteroussim-to-realagenticself-evolving
  136. arxiv:2609.35316 · cs.AI
    Reliability Engineering for AI Systems: Challenges, Methods, and Directions
    Rong Pan, Yili Hong, Min Xie

    AI reliability concerns whether an AI system performs its intended function dependably over a stated period and under stated operating conditions, with stated evidence. As these systems become more autonomous, that function includes more than a correct output. Retrieval, memory, tool use, permission

    agentictool useself-evolvingbenchmark
  137. arxiv:2609.35312 · cs.CL
    MemoReason: Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs
    Zineddine Tighidet, Andrea Mogini, Jiali Mei, Patrick Gallinari +1

    Large Language Models (LLMs) perform well on reasoning benchmarks, but it remains unclear whether this reflects genuine contextual reasoning or reliance on facts memorized in their parameters. We investigate this by distinguishing two possibilities: a broad \textit{memorization bias}, where familiar

    memorybenchmark
  138. arxiv:2609.35311 · cs.RO
    RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts
    Jin Hyun Kim, Min Young Kim, Soohwan Song, Daekyum Kim

    Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predic

    world modelaction-conditioned
  139. arxiv:2609.35308 · cs.CL
    Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation
    Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah

    Large language models process conversation history as unverified context: false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, isolating distinct fail

    benchmark
  140. arxiv:2609.35303 · cs.CV
    PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents
    Jiazhou Zhou, Hu Zhou, Yucheng Chen, Jinyuan Qu +2

    Reinforcement learning with verifiable rewards (RLVR) via Group-Relative Policy Optimization (GRPO) is widely used for multi-turn VLM agent training, yet it suffers from zero-gradient silence on uniform failures and coarse episode-level credit assignment. While On-Policy Distillation (OPD) and On-Po

    embodiedagentagent benchmarkbenchmark
  141. arxiv:2609.35302 · cs.AI
    Narrowing the Horizon: Quantifying Topic Saliency Shifts in Generative Monoculture
    Oriane Peter, Elena Simperl, Kate Devlin

    As Large Language Models (LLMs) become central to how we access and share information, they play an increasingly powerful role in shaping global knowledge. However, as these models evolve, their outputs risk converging into a \textit{generative monoculture}, where the diversity of perspectives they

    post-training
  142. arxiv:2609.35298 · cs.AI
    Training-Free Clinical Reasoning through Medical Ontologies and Cognitive Mapping: A Symbolic-Probabilistic Knowledge Graph Framework
    Surajit Das

    Most clinical prediction systems learn patient-variable-outcome associations; we investigate a training-free diagnostic paradigm mapping patient observations to explicit medical knowledge. CKG Reasoner integrates candidate-specific Evidence Feature Nodes, patient-reference matching, a bounded Inform

    knowledge graph
  143. arxiv:2609.35291 · cs.LG
    Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
    Shunchang Liu, Lukas Fluri, Xin Chen, Francesco Croce

    Modern AI models are aligned through post-training to adapt them to downstream tasks. Recent work shows that fine-tuning language models on narrow tasks can induce emergent misalignment (EM), causing broadly harmful behaviors beyond the training task. However, EM has been studied almost entirely in

    agenticpost-training
  144. arxiv:2609.35290 · cs.AI
    EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning
    Shihan Dou, Shaofan Liu, Zhonghang Lu, Jiahang Lin +7

    Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, w

    agentbenchmark
  145. arxiv:2609.35288 · cs.LG
    $λ$-JEPA Spectral Anti-Collapse Regularization for Self-Supervised Learning
    Berker Demirel, Clémentine Dominé, Valentino Maiorca, Marco Fumero +2

    Joint-embedding self-supervised learning typically combines an invariance objective across augmented views with additional mechanisms to prevent representational collapse. These objectives are often applied after a projection head, while downstream tasks use the backbone representation before the pr

    v-jepavjepabenchmark
  146. arxiv:2609.35286 · cs.AI
    The Argument and the Letterhead: Source-Position Coherence in AI Evaluation
    Michele Loi

    An argument can be surprising coming from a particular speaker without being a bad argument. Do AI evaluators keep these judgments apart? Two preregistered descriptive studies and a later Jev supplement collected 2,976 usable evaluations of six fixed texts about US AI policy, Germany's debt brake an

    evaluator
  147. arxiv:2609.35285 · cs.AI
    Textual User Taste: Natural-Language User Context for Foundation-Model Recommender System at Scale
    Ghazal Fazelnia, Paul Gigioli, Eliza Klyce, Sharon Zheng +15

    Foundation model recommender systems require user context that can be consumed by large language models, reasoned over, and refined through natural-language interaction. Traditional behavioral embedding vectors remain highly effective for retrieval and ranking, but they are opaque to users and not n

    evaluation framework
  148. arxiv:2609.35279 · cs.CL
    Measuring Collapse and Correction in Homogeneous-Panel LLM Debate
    Xin Li, Mengbing Liu, Chau Yuen

    Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing

    multi-agent
  149. arxiv:2609.35269 · cs.LG
    eval-unlearn: Benchmarking unlearning in Text-to-Image Diffusion Models
    Mansi, Nikhil Raghavan, Zixia Huang, Kai Sheng Ong +3

    The rising number of concept unlearning techniques for text-to-image (T2I) diffusion models has produced a fragmented evaluation landscape. Methods are assessed under heterogeneous experimental conditions making principled cross-method comparison difficult. We present eval-unlearn, an open-source Py

    benchmarkleaderboard
  150. arxiv:2609.35267 · cs.RO
    GuardPIBT: Counterfactually Gated Neural Guidance for Ultra-Large-Scale 3D Multi-Agent Path Finding
    Yuan Zhou, Zhenyu Hou, Guangtong Xu, Xiaoqiang Ji +3

    Large-scale 3D multi-agent path finding becomes increasingly difficult under dense traffic. Priority Inheritance with Backtracking (PIBT) scales well, but its one-step goal-directed ordering may become insufficient under dense interactions and large-scale congestion. We present GuardPIBT, which augm

    multi-agent
  151. arxiv:2609.35262 · cs.CL
    Rubric-Aware On-Policy Self-Distillation for LLM Personalization
    Yilun Qiu, Xiaoyan Zhao, Chengbing Wang, Cilin Yan +5

    LLM personalization aims to generate responses aligned with individual users' preferences and needs. User-specific rubrics make these expectations explicit, providing direct supervision on what a satisfactory answer should cover. Existing rubric-guided approaches, however, exploit such guidance only

    graspbenchmark
  152. arxiv:2609.35255 · cs.AI
    Towards Reliable AI Data Scientists: Data Agents with Workflow Harnesses
    Huachi Zhou, Yujing Zhang, Jiahe Du, Jiacheng Cai +6

    Large language model agents are increasingly deployed for data-intensive work, yet reliable data analysis requires more than general-purpose reasoning and ad hoc tool augmentation. Data Agents, equipped with workflow harnesses, offer a promising paradigm for automating the end-to-end data science li

    benchmark
  153. arxiv:2609.35249 · cs.RO
    Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies
    Dingsheng Liu, Yangzheng Wu, Mahboubeh Asadi, Zhiyuan Li +4

    Pretrained robot manipulation policies such as vision-language-action models (VLAs) or world-action models (WAMs) leave interaction-relevant metric geometry implicit. Recent breakthroughs in spatial reconstruction can supply the necessary geometry reliably, but their features describe local shape wi

    vision-language-actionmanipulationrobotwinbehavior-1kbenchmark
  154. arxiv:2609.35246 · eess.SY
    Hard-Constrained Probabilistic Factor Graph Neural Network for Distribution System State Estimation under Non-Gaussian Uncertainty
    M. Furqan Azam, Marta Vanin, Chris Hermans, Geert Deconinck

    Robust and accurate state estimation is fundamental for the reliable operation and monitoring of active distribution networks. Conventional numerical estimators, such as weighted least squares, are computationally slower and often suffer from convergence issues in the presence of sparse measurements

    benchmark
  155. arxiv:2609.35236 · cs.LG
    Long-Horizon Scaling: How Model Capabilities Shape the Returns to Computation
    Haoyu Zheng, Zhengyu Chen, Huaisheng Zhu, Ruishan Fang +4

    Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab an

    benchmark
  156. arxiv:2609.35233 · cs.AI
    EP-Mem: Elastic Privacy Memory for Social Relationship-Aware LLM Agents
    Fengzhou Sun, Yuan Zhang, Xintong Yu, Jinyao Yan

    Large language model (LLM) agents face critical privacy risks when acting as delegates in human-agent-human communication. To prevent such breaches, agents must understand users' social relationships and adhere to context-dependent social information disclosure boundaries. Current studies on agent m

    memorymemory architectureagent memoryagentllm agentbenchmark
  157. arxiv:2609.35231 · cs.RO
    Zero-Shot Reactive Obstacle Avoidance for Generative Robot Policies
    Weihang Guo, Lydia E. Kavraki

    We propose NUDGE (Nudge Update via Differentiable GEometry), a training-free obstacle-avoidance procedure that can be incorporated in any robot policy based on diffusion or flow matching, including diffusion policies and vision-language-action models. Our work injects gradients from a signed distanc

    vision-language-actionrobot policy
  158. arxiv:2609.35228 · cs.CV
    Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning
    Hao-Xuan Ma, Yihao Liu, Yutao Sun, Yanting Miao +7

    Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sen

    benchmark
  159. arxiv:2609.35226 · cs.CV
    Generative AI-Based Data Augmentation for Oral Lesion Classification: The PhotoMOCI Dataset and Benchmark
    Marco Parola, Mario G. C. A. Cimino, Sabrina Senatore

    Early detection of oral cancer via photographic imaging presents a promising avenue for large-scale oral cavity screening. However, the development of robust deep learning models is frequently hampered by the scarcity of high-quality, annotated datasets. To address this limitation, a novel and well-

    benchmark
  160. arxiv:2609.35225 · cs.CV
    SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
    Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai +1

    Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propo

    benchmark
  161. arxiv:2609.35215 · cs.AI
    ASCT: Attentive Search over Counterfactual Trees for Credit Assignment in Agentic Reinforcement Learning
    Yang Li, Jinhan Yang, hai liu, Di Wan +7

    Terminal utility evaluates a complete agentic workflow, but learning requires credit for the decisions within it. We introduce Attentive Search over Counterfactual Trees (ASCT), a framework that turns training-time multi-step search into local action credit. At actor-visited states, an auxiliary tre

    retrieval-augmentedagentic
  162. arxiv:2609.35212 · cs.LG
    Adversarial Consistency-Guided Representation Learning for Multi-view Clustering
    Yuchen Lin, Kunpeng Xu, Ying Fang, Lifei Chen

    Multi-view clustering aims to capture cross-view consistency while exploiting view-specific information. However, shared representations learned to capture cross-view consistency may still retain view-identifying information, potentially compromising the consistency of cross-view clustering structur

    benchmark
  163. arxiv:2609.35210 · cs.CL
    Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
    Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou

    On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoder

    post-training
  164. arxiv:2609.35208 · physics.optics
    Generating Vector-Vortex $γ$ Photons by Nonlinear Compton Scattering
    Yong-Zheng Ren, Mamutjan Ababekri, Jun-Lin Zhou, Feng Wan +7

    Vector-vortex photons, characterized by a nonseparable coupling between polarization and orbital angular momentum (OAM), offer opportunities for optical manipulation, quantum communication, nuclear photonics, etc. However, their generation in the $γ$-ray regime remains challenging. Here, we put forw

    manipulation
  165. arxiv:2609.35200 · cs.RO
    ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation
    Pankhuri Vanjani, Mostafa Hatab, Can Mizrakli, Vaisakh Shaj +3

    Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors, such as recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time. We present ReCAT, a language-conditioned poli

    manipulationliberomemory
  166. arxiv:2609.35201 · cs.AI
    From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data
    Husrev Taha Sencar, Rezart Beka, Danish Naeem, Seda Ozalkan +5

    Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounde

    post-trainingbenchmark
  167. arxiv:2609.35196 · cs.AI
    GAC-PINN: Geometry-Adaptive and Constraint-Enhanced Physics-Informed Neural Networks
    Yanxin Zhang, Yong Zhang, Houbiao Li

    For systems with steep gradients, sharp interfaces, or severe spatio-temporal coupling, Physics-informed neural networks (PINNs) suffer from spectral bias, geometric inflexibility, and boundary constraint conflicts, which undermine accuracy and convergence. To overcome these issues, we propose a geo

    benchmark
  168. arxiv:2609.35195 · cs.CV
    CarveMix-RC: Addressing Rare-Class Imbalance Through Lesion-Aware Synthetic Augmentation for Brain Metastasis Segmentation
    Md Shibly Sadique, Md Fayaz Bin Hossen, Michael L. Evans, Walia Farzana +3

    Accurate segmentation of post-treatment brain metastases is essential for treatment planning, longitudinal disease monitoring, and quantitative assessment of therapeutic response. The BraTS-MET 2026 Task 1 challenge introduces a clinically relevant segmentation problem involving four anatomically di

    benchmarkevaluation protocol
  169. arxiv:2609.35193 · cs.LG
    ConRAG: Lightweight inference of multi-hop relations
    Kilian Bänziger, Sonia Laguna, Markus Kreft, Robert Jakob +3

    Understanding how two entities are connected often requires tracing multi-hop relations across documents to identify intermediate entities and supporting evidence that explain a connection. This is a task that appears frequently in scientific research and other knowledge-intensive analyses. We forma

    ragknowledge graph
  170. arxiv:2609.35189 · cs.LG
    G$^3$-LoRA: Organizing Reward-Weighted Video Data with Gradient-Guided Grouped LoRA
    Jia Song, Wenhow Li, Lichen Bai, Bada Ye +1

    Post-training foundation video models on heterogeneous reward-weighted data usually assume that all data categories induce compatible updates. This assumption is fragile when categories correspond to different skills, domains, or evaluation dimensions. We study this problem in text-to-video post-tra

    post-trainingevaluator
  171. arxiv:2609.35188 · cs.AI
    Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference
    Suwesh Prasad Sah

    Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. This study evaluates two-token multi-token prediction (MTP) against autoregressive decoding in a controlled

    memorybenchmark
  172. arxiv:2609.35182 · cs.AI
    Research-Native by Construction: Minimal Nodes, Re-verifiable Workflows, and Compounding Memory for Long-Horizon Scientific Agents
    Di Wang, Yu Liu, Bing Cui, Chaoqun Ji +4

    We describe AfS (Agent for Science), a platform built for long-horizon scientific work, where a project runs for tens of hours across dozens of agent runs with a human present only occasionally. Most agents for science are general coding agents with a skills folder attached, and they inherit that li

    memoryagentbenchmark
  173. arxiv:2609.35177 · cs.LG
    Subgroup Rank-1 Lattice for Practical High-dimensional Black-box Integral Approximation
    Yueming Lyu

    Estimating integrals of black-box, high-dimensional functions, from expectations and kernel mean embeddings to the softmax kernel in self-attention, is a basic subroutine in machine learning. Rank-1 lattice rules suit this setting: they query the integrand only at a fixed point set and need no gradi

    memory
  174. arxiv:2609.35160 · cs.AI
    FONDANT: Strong and Best-Effort Planning via Antichains
    Benjamin Aminof, Tuan Khai Nguyen, Sasha Rubin

    A classical solution concept in fully observable nondeterministic (FOND) planning, is the strong policy (aka winning strategy in the closely related area of reactive synthesis), i.e., such a policy ensures that the goal is reached in an adversarial environment. When strong policies are not available

    agentbenchmark
  175. arxiv:2609.35158 · cs.AI
    PEARL: Adaptive Prefill-Decode Execution with Elasticity for Agentic Reinforcement Learning
    Jiaan Zhu, Wei Gao, Youhui Bai, Zewen Jin +5

    Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but adding rollout replicas yields diminishing returns while training GPUs remain idle between updates. We observe that effective resource use als

    agentic
  176. arxiv:2609.35152 · cs.LG
    Coordinated Lane-Level Variable Speed Limits and Ramp Metering for Successive Weaving Segments Considering Merging/Diverging Risks: A Hybrid Model Predictive Control and Multi-Agent Reinforcement Learning Approach
    Guodong Ma, Baofeng Sun, Wenyu Yang, Zhihong Yao

    Successive weaving segments (SWSs) on urban expressways are bottlenecks prone to recurrent congestion and collisions, requiring fine-grained active traffic management (ATM). Existing approaches struggle to balance the adaptive performance of data-driven optimization with the resilience and transfera

    multi-agent
  177. arxiv:2609.35149 · cs.AI
    From Migration to Calibration: Preserving Agent Capabilities across Models, Jurisdictions, and Scale
    Yaxiao Liu, Pengbo Liu, Yiwen Liu, Yihua Guan +1

    Agents need calibration when deployment conditions change: replacing a driving model, including a foundation-to-post-trained transition; crossing jurisdictions; or scaling across heterogeneous markets and sources. Interface compatibility alone does not establish capability retention or target-contra

    agent
  178. arxiv:2609.35143 · cs.CV
    Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
    Gunin Gupta, Nirmit Arora, Pavan Kalyan Tankala

    AI agents increasingly carry out long-horizon professional work, but their evaluations rarely require a finished creative deliverable. To this end, we introduce Timeline-Bench, a benchmark of 56 real video-editing tasks, each asking an agent to turn raw production material into a finished video. Tas

    agentai agentbenchmark
  179. arxiv:2609.35142 · cs.MA
    Strategically Robust Game-Theoretic Multi-Agent Trajectory Optimization
    Victor L. Qin, Nicolas Lanzetti, Saverio Bolognani, Hamsa Balakrishnan

    Aviation authorities worldwide expect Advanced Air Mobility (AAM) traffic management to be decentralized among service providers, requiring AAM flights to autonomously plan trajectories by predicting other flights' control inputs rather than relying on centralized coordination. Game-theoretic approa

    agentmulti-agent
  180. arxiv:2609.35139 · cs.LG
    CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion
    Genglin Wang, Wangsong Yin, Yeerzhati Abudunuer, Haoxuan Xu +2

    Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the asse

    retrieval-augmentedrag
  181. arxiv:2609.35138 · cs.LG
    FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
    Shidu Ren, Qilin Gu, Zhenghao Ni, Junhan Sun +3

    Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a

    world modelbenchmark
  182. arxiv:2609.35134 · cs.CV
    VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction
    Conghan Yue, Yuanjie Chen, Yue Han, Ya Gao +3

    Video editing has advanced substantially in recent years, with methods increasingly accounting for the visual consequences of edits, such as changes to shadows and occlusions. However, the physical consequences of edits, including changes to subsequent motion and interactions, remain less explored.

    benchmark
  183. arxiv:2609.35117 · cs.AI
    Tool Mediation Alters Refusal Mechanisms in Large Language Models
    Abel Rodríguez, Giuseppe Garofalo, Lieven Desmet, Vera Rimmer

    Large language models (LLMs) are increasingly deployed with access to external tools, yet harmful tool-mediated interactions are less likely to be refused when compared to regular conversational ones. As this change in refusal behavior remains underexplored, we investigate its underlying mechanisms

    llm agent
  184. arxiv:2609.35115 · cs.AI
    DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines
    Haixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie

    Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives.

    memory
  185. arxiv:2609.35113 · cs.LG
    SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression
    Ziwen Zhang, Xiju Wu, Yuheng Jing, Runxiang Wang +8

    Symbolic regression (SR) seeks concise and interpretable mathematical expressions from data for scientific equation discovery. Existing SR benchmarks face a tradeoff between evaluation cost and benchmark validity. Repeated evaluation of large task pools is expensive, and compact benchmarks lack syst

    benchmark
  186. arxiv:2609.35110 · cs.LG
    Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
    Yitong Li, Jincheng Yu, Junsong Chen, Haopeng Li +5

    Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substant

    memoryself-improvement
  187. arxiv:2609.35107 · cs.AI
    DoAtlas-2: A Foundation for Self-Evolving Causal Biomedical Discovery
    Yulong Li, Rong Xia, Yuxuan Zhang, Jianxu Chen +11

    We introduce DoAtlas-2, a foundation for self-evolving causal biomedical discovery that organizes knowledge around causal mechanisms and advances through external evidence from human populations. DoAtlas-2 integrates 771 research resources covering more than 720,000 participants in 48 countries, fro

    self-evolving
  188. arxiv:2609.35106 · cs.LG
    DRIFT: Disentangled Responsive-Invariant Flow Transport for Single-Cell Perturbation Prediction
    Mustapha Bounoua, Giulio Franzese, Pietro Michiardi

    Predicting cellular responses to perturbations is a central problem in cellular biology, with broad applications in systems biology and drug discovery. This task is challenging because cellular responses can be complex and cell-state dependent, intrinsic cell-to-cell variability can be confounded wi

    benchmark
  189. arxiv:2609.35099 · cs.LG
    E3J: An Efficient and Open-Source Backend for Euclidean Equivariant Operations on GPU and TPU
    Olivier Peltre, Armand Picard, Adrien Pichard, Miguel Bragança +5

    We present e3j, a fast Euclid-equivariance backend for geometric deep learning applications with JAX bindings for GPU and TPU. Leveraging both optimized CUDA and Pallas kernels and algorithmic improvements, the library achieves state-of-the-art throughput and runtime on both forward and backward pat

    memorybenchmark
  190. arxiv:2609.35097 · cs.LG
    SpikeLite: Lightweight Spiking Neural Networks for Time-Series Forecasting
    Bang Hu, Changze Lv, Mingjie Li, Xiaoqing Zheng +2

    Spiking neural networks (SNNs) offer an energy-efficient paradigm for time-series forecasting through spike-driven computation. However, recent SNN forecasters often pursue higher accuracy through increasingly complex attention mechanisms, or specialized neuronal dynamics, weakening the lightweight

    benchmark
  191. arxiv:2609.35096 · cs.LG
    DF-CBM: Region-Aware Concept Bottleneck Models for Deepfake Detection
    Georgios Tsoumplekas, Vazgken Vanian, Alexandros Doumanoglou, Panos K. Papadopoulos +3

    Deepfake detection methods have become increasingly effective yet most provide limited insight into the evidence behind their predictions. However, in forensic settings users also need to know which manipulation cues support the decision and where they appear. Existing explainability methods only pa

    manipulation
  192. arxiv:2609.35094 · cs.RO
    QuadHand: A Compact Quadrotor Aerial Manipulator with MRC-SDF-Based Whole-Body Motion Planning
    Rui Jin, Ruiyang Liu, Xinhang Xu, Haotian Jin +3

    Uncrewed aerial manipulators (UAMs) integrate robotic arms with aerial platforms for three-dimensional physical interaction. However, enlarging the workspace increases arm-induced disturbances, while existing geometric representations face a trade-off between geometric fidelity and computational eff

    manipulationmanipulatorgripper
  193. arxiv:2609.35090 · cs.CV
    Advancing Video-Text Pretraining with Multi-View Captions
    Fida M. Thoker, Renaud Vandeghen, Karen Sanchez, Marc Van Droogenbroeck +1

    Video-text pretraining has achieved remarkable progress through the scaling of models and datasets, yet the quality of language supervision remains underexplored. Existing web-scale datasets often provide only a single sparse caption per video that fails to capture rich spatiotemporal semantics, whi

    benchmark
  194. arxiv:2609.35089 · cs.AI
    Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research
    Zehao Lu, Xingguo Xiong, Wopke van der Werf, Thijs L. van der Plas +1

    Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-inte

    multi-agentagent system
  195. arxiv:2609.35088 · cs.AI
    When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents
    Geonwoo Kim, Brent ByungHoon Kang

    Tool-enabled agents form calls from model-visible interfaces, while hosts later select their implementation. Standard dispatch omits the descriptor-handler relation. An unchanged and schema-valid call can therefore acquire a different security effect during rollout, reconnect, or delayed approval. W

    llm agent
  196. arxiv:2609.35086 · cs.LG
    Retrieval-Augmented Diffusion Modeling for Stochastic Discount Factor Portfolios
    Kelvin J. L. Koa, Xinyang Li, Ke-Wei Huang

    In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial markets exhibit non-stationary dynamics wit

    retrieval-augmented
  197. arxiv:2609.35084 · cs.LG
    GraphHCA: Closed-Form Hindsight Credit Assignment for Long-Horizon LLM Agents
    Haodong Zhu, Yangyang Ren, Changbai Li, Sheng Xu +3

    Group-based reinforcement learning (RL) has advanced large language models (LLMs) and is increasingly extending to agentic tasks, where sparse terminal rewards make step-level credit assignment essential. Existing methods assign credit from what follows an action in sampled rollouts, but do not expl

    llm agentagentic
  198. arxiv:2609.35083 · cs.LG
    Continuous Variational Synthesis
    Alan N. Amin, Mattia G. Gollub, Andrei Slabodkin, Elizabeth B. Wood +1

    Biological machine learning was long bottlenecked by the ability to synthesize designed DNA. Variational synthesis models control chemical reactions to physically manufacture quadrillions of designed sequences in DNA. However, training these generative models is challenging: constraints on chemical

    post-training
  199. arxiv:2609.35082 · cs.LG
    Cross-Rollout Bellman Closure for Long-Horizon Agentic Reinforcement Learning
    Yangyang Ren, Haodong Zhu, Linlin Yang, Sheng Xu +2

    Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon settings, these rollouts revisit shared anchor states, offering cross-rollout evidence for step-level credit. Ideally, step-level credit should inc

    llm agentagenticbenchmark
  200. arxiv:2609.35078 · cs.CV
    RefineDrive: Reliable Failure-Guided Learning for Vision-Language-Action Driving
    Zhe Sun, Ziyi Luo, Yehao Lu, Lei Zhou +1

    Vision-Language-Action (VLA) models for autonomous driving rely heavily on successful expert demonstrations, leaving model-specific failures underexploited. Learning from these failures is hindered by unreliable diagnoses, poorly matched correction targets, and coarse rewards. We propose RefineDrive

    vision-language-actionpost-training
  201. arxiv:2609.35072 · cs.LG
    Not All Rollouts Are Worth Learning: On Trajectory Valuation for Post-Training Reinforcement Learning
    Xuesong Jia, Ziao Yang, Zhanhe Huang, Hongfu Liu

    We consider the problem of trajectory valuation in reinforcement learning: how to identify and mitigate detrimental trajectories during online training. Unlike classification, where data valuation relies on fixed training and validation sets, reinforcement learning involves dynamically generated tra

    post-training
  202. arxiv:2609.35065 · cs.LG
    TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash
    Jay H. Park, Hyungjun Kim, Dong Kim

    Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical KV hit is not necessarily ready for GPU retrieval. Demand staging exposes SSD latency, whereas imme

    memory
  203. arxiv:2609.35060 · cs.LG
    Beyond Gradient Flow: Identifiability and Recovery from Distribution Snapshots
    Nam D. Nguyen, Valeriya Malysheva

    Inferring dynamics from snapshots of evolving distributions is fundamentally underdetermined: the Fokker-Planck equation constrains the drift $F$ only through its score-weighted divergence $\nabla\cdot F+F\cdot\nabla\logρ$, leaving a $ρ$-solenoidal gauge invisible to any single-time constraint. Time

    benchmark
  204. arxiv:2609.35058 · cs.LG
    TIDE: Teacher-Student Transition via Informative Distillation and Exploration for Agentic RL
    Yibin Huang, Xinming Xu, Conghui Zhu

    Effective multi-turn agents require interaction strategies that coordinate information gathering, actions, and feedback over long horizons. GRPO is a reinforcement learning algorithm used to train these agents, but sparse trajectory-level rewards limit early exploration in small models. Recent metho

    agenticbenchmark
  205. arxiv:2609.35055 · cs.LG
    Universality and Generalization of Causal Transformers Across Context Lengths
    Takashi Furuya, Maarten V. de Hoop, Gabriel Peyré

    Long contexts are central to modern transformer systems, but most expressivity results choose a different network for each fixed sequence length. We study whether one masked transformer can approximate causal token-to-token maps uniformly over sequences of arbitrary length sampling a fixed normalize

    long context
  206. arxiv:2609.35052 · cs.CV
    OPIS: An Input-Grounded Benchmark for Multi-Object Memory in Video World Models
    Hao Wang, Tao Yu, Liuzhou Zhang, HeXin Wang +12

    Video world models must preserve the visual state of the world over time, but existing evaluation protocols often rely on generated histories, video reference, or selected revisit viewpoints that can confound the assessment of a model's true memory capability. To address this, we introduce OPIS, an

    embodiedworld modelmemorybenchmarkevaluatorevaluation protocol
  207. arxiv:2609.35047 · cs.RO
    EMPIRIC: Experiment-Driven Learning of Residual World Models for Robot Planning
    Yichao Liang, Amber Li, Dat Nguyen, Emily Bunnapradist +12

    A robot should be able to learn through experiments how unfamiliar objects behave and interact, then plan with that knowledge. It need not start from scratch: physics engines supply knowledge of motion and contact, but can omit entire mechanisms, such as glue curing, water heating, or wind. We prese

    manipulationworld modelagent
  208. arxiv:2609.35046 · cs.CV
    LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation
    Hongli Xu, Zhaowei Lu, Junwen Huang, Jiaqi Hu +4

    Category-level 6D pose estimation from a single RGB-D observation is inherently under-constrained, since partial visible geometry must be interpreted together with a canonical object structure before a stable pose can be determined. We present LEGAU, a unified framework that jointly predicts NOCS co

    benchmark
  209. arxiv:2609.35043 · cs.CV
    Mixed-Prior Decision Risk for Open-Set Recognition
    L. A. Erlygin, P. D. Proskura, A. A. Zaytsev

    In open-set recognition (OSR), a probe must either be identified as one of the known gallery classes or rejected as unknown, so three error types coexist: false acceptance, false rejection, and misidentification. An uncertainty score for selective recognition should rank probes by the risk of the de

    benchmark
  210. arxiv:2609.35039 · cs.RO
    Do Not Cut When Uncertain: Rejectable and Calibrated Decision Heads for VLA Policies in Robotic Harvesting
    Heng Zhang

    Vision-Language-Action (VLA) policies trained with behavior cloning or flow matching are optimized to output an action trajectory, but they cannot express "I don't know" or "I should not act." In robotic harvesting, occlusion makes single-frame decisions fundamentally ambiguous: identical pixels can

    vision-language-actionvlamanipulation
  211. arxiv:2609.35038 · cs.LG
    GUIDE-FBO: Guidance via Uncertainty Intervention and Distributional Exchange for Federated Bayesian Optimization
    Jintao Wei, Chenxi Li, Songhao Wang

    Federated Bayesian Optimization (FBO) enables distributed agents to collaboratively optimize expensive black-box objectives without sharing raw local observations. However, effective knowledge transfer remains challenging under communication constraints and task heterogeneity. We propose GUIDE-FBO,

    agentbenchmark
  212. arxiv:2609.35035 · cs.LG
    THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout
    Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni

    The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDS

    benchmark
  213. arxiv:2609.35034 · cs.LG
    Recommendation Ranking Off-Policy Evaluation under Ranking-Dependent Examination via Examination-Relevance Decomposition
    Riki Okamura, Toshiharu Sugawara

    Off-policy evaluation, which estimates evaluation policy performance from logged data, is key for recommender ranking policies. However, logged clicks cannot distinguish unexamined items from examined non-clicks, causing bias in existing estimators when the assumed examination structures fail. We pr

    policy evaluation
  214. arxiv:2609.35032 · cs.RO
    JRDB-AVR: An Active Visual Reasoning Benchmark for Embodied Agents in Real-World Environments
    Zhixi Cai, Fucai Ke, Sukai Huang, Maria Garcia de la Banda +3

    In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object,

    embodiedworld modelagentagenticembodied agentbenchmark
  215. arxiv:2609.35028 · cs.LG
    VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation
    Hanxun Huang, Yutao Wu, Qizhou Wang, Silvia Montaña-Niño +8

    Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content

    benchmarkscalable evaluationscalable evalllm-as-judge
  216. arxiv:2609.35026 · cs.AI
    WebPageBench: Event-Level Verification and Controlled UI-Variant Generation for Web Agents
    Anton Emelyanov, Maria Tikhonova, Zaven Martirosian, Sergei Averkiev +1

    We present WebPageBench, an open framework for evaluating web agents in which every task is verified from the interface's own event log. Six instrumented mock sites with brand identifiers removed (a marketplace, a bookstore, a grocery service, rail ticketing, hotel search and a document cabinet) emi

    agentjudge modelleaderboard
  217. arxiv:2609.35025 · cs.AI
    AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
    Haotian Luo, Haoyu Wang, Zeyu Qin, Huanjin Yao +4

    Recent gains in language model capability have come more from data than from architecture. Frontier labs and data companies produce verifiable agentic tasks, which supervised finetuning and reinforcement learning then turn into capability.This production line still rests on human labour and on human

    agentagenticself-improvementhuman-in-the-loopbenchmark
  218. arxiv:2609.35023 · cs.CV
    Proxy2World: Learning to Generate Worlds From Lightweight Proxies without Seeing Them
    Hongli Xu, Weilong Yan, Anbang Wang, Chunyu Zou +2

    Lightweight scene proxies let creators control scene layout and motion while leaving room for imagination in appearance, lighting, and visual effects. However, a suitable proxy is not uniquely defined, making paired proxy-video data difficult to construct automatically at scale. We present Proxy2Wor

    world model
  219. arxiv:2609.35022 · cs.CV
    Detection of Adversarial Attacks on Super-Resolvers Using Spectral Features
    Emma J. Reid, Haley Duba-Sullivan, Tony G. Allen

    The integration of deep learning models into image preprocessing pipelines such as super-resolution introduces a largely unexplored attack vector for adversaries targeting downstream tasks. To ensure trustworthiness of critical imaging pipelines, we must be able to detect adversarial behavior within

    benchmark
  220. arxiv:2609.35021 · cs.LG
    Addressing Spatial Indistinguishability in Spatiotemporal Prediction via Optimal Transport-Guided Masking
    Guangyu Wang, Jiawei Tong

    Spatiotemporal prediction aims to learn discriminative representations from correlated temporal signals over spatial structures for accurate future inference. A central challenge is \emph{spatial indistinguishability}: different nodes may share similar historical patterns yet evolve toward divergent

    benchmark
  221. arxiv:2609.35017 · cs.AI
    TermJudge: A Document-Level Metric Judging, Not Counting, Terminology in Machine Translation Evaluation
    Nicolas Dahan, Fran{\cc}ois Yvon, Rachel Bawden

    Existing automatic metrics for evaluating terminological use in machine translation (MT) penalise any divergence from a fixed reference, conflating translation errors with the valid terminological variation that human translators routinely produce. We introduce TermJudge, a document-level terminolog

    llm-as-judge
  222. arxiv:2609.35005 · cs.LG
    Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device
    Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański +1

    Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must m

    memory
  223. arxiv:2609.35003 · cs.RO
    Learning to Act under Visual Interruptions with Vision-Language-Action Models
    Mingle Jiang, Rui Xu, Yunke Wang, Chang Xu

    Vision-language-action (VLA) models have demonstrated strong capabilities in robotic manipulation, but they are typically developed and evaluated with all camera streams available throughout task execution. When a camera stops delivering frames during task execution, the policy must continue acting

    vision-language-actionvlavla modelmanipulationgr00tworld model
  224. arxiv:2609.35000 · cs.RO
    Graph-Based Simultaneous Path and Foothold Planning for Multi-Limbed Intra-Vehicular Robots in Space Stations
    Masazumi Imai, Kentaro Uno, Toshinori Kuwahara, Kazuya Yoshida

    Robot-aided operations in space stations are essential for reducing the workload of astronauts and improving the efficiency of on-orbit activities. Multi-limbed intra-vehicular robots (MLIVRs) equipped with grappling end-effectors have emerged as a promising solution, as they can securely grasp pre-

    manipulationgrasp
  225. arxiv:2609.34988 · cs.CL
    The Right Lesson at the Right Step: Deriving Control Updates for Self-Evolving Agents
    Yunhe Su, ZiYi Dong, Tong Yu, Weijian Deng +3

    Self-evolving agents improve future behavior by reusing past experience, typically as global prompts, memories, or reflections. Yet these mechanisms rarely control where experience takes effect. In long tool-use workflows, the same lesson may correct one decision but distract another, making experie

    memoryagenttool-useself-evolvingbenchmark
  226. arxiv:2609.34982 · cs.RO
    ActionUNet: Improving Robustness of VLA Models with Efficient Multi-scale Fine-tuning
    Di Zhu, Ziheng Yan, Fang Wan

    Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited gener

    vision-language-actionvlavla modelmanipulationopenvlalibero
  227. arxiv:2609.34979 · physics.app-ph
    Far-field excitation of symmetry-protected bound states in the continuum by nonlinear virtual sources
    Marc Martí-Sabaté, Yijie Zhang, Shengming Sun, Ruxin Li +3

    Bound states in the continuum (BICs) enable complete wave confinement within open systems through destructive interference or symmetry protection, giving rise to ideally infinite quality factors and extreme field enhancement. Their practical exploitation, however, is hindered by a fundamental parado

    manipulation
  228. arxiv:2609.34978 · cs.CV
    One Sensor, Whole Body - 3D Body Pose from a Single Consumer Earbud IMU
    Zhilin Guo, Boqiao Zhang, Oszkár Urbán, Josef Bengtson +7

    Consumer earbuds already stream inertial motion data from the head, one of the most widely worn sensor locations on the body. We ask how much of the 3D body pose a single such head IMU can recover, and whether adding more consumer sensors actually helps. We build a multimodal capture pipeline that r

    benchmark
  229. arxiv:2609.34975 · cs.LG
    Teach to Learn: Hint Annealing for Self-improving LLM Reasoning
    Zile Wang, Zijian Li, Haodong Wang, Jian Liu +4

    Group Relative Policy Optimization (GRPO) improves language-model reasoning by comparing verified rewards among multiple solution rollouts for each query. However, difficult training queries can yield only incorrect rollouts, leaving GRPO with no reward contrast or learning signal. Prior hint-based

    self-improvingself-improvementbenchmark
  230. arxiv:2609.34974 · cs.AI
    Before Acting, Change the State: Prospective State Intervention for Web Agents under Deceptive Interfaces
    Ruozhao Yang, Mingfei Cheng, Xiaofei Xie

    LLM-based Web agents can autonomously complete user tasks, yet deceptive interfaces can steer them toward outcomes that conflict with users' interests. Existing defenses primarily intervene on agent behavior through blocking, guidance, or replanning. We identify a distinct failure mode: a task-valid

    agentbenchmark
  231. arxiv:2609.34973 · cs.AI
    APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
    Puneet Mathur, Dinesh Manocha

    Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes su

    benchmark
  232. arxiv:2609.34972 · cs.CV
    Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models
    Jingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang +2

    Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers

    benchmark
  233. arxiv:2609.34971 · cs.AI
    Action-Space Shaping for LLM Agents: Measuring and Mitigating Tool-Schema Bias
    Yinhong Liu, Zhili Tan, Zilin Wang, Zhijiang Guo

    Large Language Models (LLMs) have shown strong performance on tool-use agentic tasks when given a fixed tool schema. Yet a tool schema is not the action space of an agent; it is merely one interface representation of it. The same executable action can be exposed through many different, functionally

    agentllm agentagentictool-use
  234. arxiv:2609.34969 · cs.RO
    NavJev: Efficient Vision-Language Navigation via Action-Centric Visual Compression and Discriminative Action-Semantic Memory
    Kai Sheng, Liuyi Wang, Jinlong Li, Haojie Dai +2

    Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigat

    embodiedmemorysemantic memoryembodied agent
  235. arxiv:2609.34968 · cs.RO
    RoboFL: Federated Expert Assembly for World Action Models
    Rongyu Zhang, Ruizhi Fan, Yunfan Lou, Hengyu Fang +7

    Vision-language-action and world-action models are increasingly popular, yet remain bottlenecked by physical interaction data that is scarce, institutionally siloed, and task-heterogeneous. A natural federated solution is to let each client adapt a shared foundation model through parameter-efficient

    vision-language-actionvlarobotwinfranka
  236. arxiv:2609.34963 · cs.AI
    JevVibe: Efficient Classification-Guided Secure Code Generation
    Arshak Rezvani, Sasha Behrouzi, Ahmad-Reza Sadeghi

    Large language models can generate functionally correct code that still contains security weaknesses, motivating repair pipelines that first diagnose a weakness type before deciding how to fix it. The Common Weakness Enumeration (CWE) provides a standardized vocabulary for such diagnoses, but asking

    agentbenchmark
  237. arxiv:2609.34962 · cs.LG
    ALICE: In-context, Zero-shot, Mutual Information Estimation
    Giulio Franzese, Simone Rossi, Pietro Michiardi

    Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are more

    benchmark
  238. arxiv:2609.34960 · cs.AI
    ProofLoom: Proof-Obligation-Driven Theory Construction for Autoformalizing Research-Level Stochastic Optimization
    Feiming Wang, Daibo Li, Kun Yuan

    Formalizing research-level stochastic optimization in Lean requires both an algorithm model and domain theory connecting foundational libraries to convergence proofs. Revising a model to restore provability can change the mathematical claim. We introduce ProofLoom, a fully automated LLM-agent system

    agent system
  239. arxiv:2609.34951 · cs.AI
    AX is the New AEO
    Ido Finder, Assaf Elovic, Gad Shalev

    In 2023, AI models answered from training data and hallucinated when it ran out, and businesses were told to seed that knowledge. Models' training knowledge has since given way to live web search, and the advice followed it there: answer-engine optimization, or AEO, now tells businesses to scatter b

    agentagentic
  240. arxiv:2609.34949 · cs.AI
    VD-DeepStack: Bridging Visual Comparison and Language Reasoning for Few-Shot Anomaly Detection
    Mengyang Zhao, Zhuolin He, Haiyang Yu, Yuxuan Liang +8

    Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abst

    benchmark
  241. arxiv:2609.34944 · cs.RO
    Adjoint Guidance Flow: Amortized Critic Guidance for VLA Policies
    Jeongsol Kim, Youngjun Jun, Kyumin Choi, Youngmin Kim +5

    Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic through a one-step surrogate of the sampl

    vision-language-actionvlavla policyliberomemory
  242. arxiv:2609.34934 · cs.CV
    TaoTex: Boosting Texture Detail Fidelity for Native 3D Material Generation
    Xiuchao Wu, Shuichang Lai, Jiangjing Lyu, Chengfei Lyu

    Recent 3D generation models can produce accurate geometries while still struggling to reconstruct detailed textures. We propose a diffusion-based native 3D material generation model TaoTex, which faithfully recovers intricate textures through tailored strategies and improvements. First, we develop a

    agent
  243. arxiv:2609.34930 · cs.AI
    PDEU-Bench: Benchmarking the Personalized Planning Lifecycle of Tool-Calling LLM Agents
    Huayi Lai, Shichao Song, Qingchen Yu, Simin Niu +3

    Large language model (LLM) agents are evolving from tool-calling systems that execute isolated instructions into task-oriented agents that pursue user goals through sustained, multi-step interactions. However, existing benchmarks for personalized tool use largely assess isolated calls or reactive ex

    memoryllm agenttool usebenchmark
  244. arxiv:2609.34924 · cs.LG
    Audit the Scaffold, Not the Checkpoint: A Stationarity Dichotomy for Recursive Self-Improvement in Agentic Coding
    Sebastian Bobadilla-Suarez, Bob Suh, Ryan Fortin

    An auditor who checks whether a system's weights are frozen is checking the wrong thing. Our stationarity dichotomy says that iterative self-modification hits strict diminishing returns whenever the agent's reachable set of edits stays fixed, and can escape only if that set expands. Rewriting scaffo

    agentagenticself-improvement
  245. arxiv:2609.34913 · cs.AI
    DGF-Bench: A Benchmark for Simulating and Auditing Deception Against Multi-Agent Governance Boards
    Jeremy Canale

    Tool-using language-model agents can review enterprise projects as governance boards do: they read the evidence, apply written rules and decide whether the project may proceed. Part of that evidence comes from suppliers and project members with a stake in the decision. DGF-Bench is a benchmark in wh

    agentmulti-agentbenchmark
  246. arxiv:2609.34911 · cs.RO
    Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
    Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim +4

    Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive

    vision-language-actionmanipulation
  247. arxiv:2609.34905 · cs.CV
    ReSight-SMC: Two-Stage Power Sampling via Island SMC with Visual Scouts
    Yaowen Zhang, Xiangyu Qiu, Junyi Hu, Zhi Lu +4

    Power sampling has emerged as a training-free approach to LLM reasoning, eliciting capabilities comparable to reinforcement learning by sharpening the model distribution over complete responses. Despite this success, power sampling remains underexplored in large vision-language models (LVLMs). We tr

    post-trainingbenchmark
  248. arxiv:2609.34900 · cs.CV
    FILIGREE3D: Scaling Sparse Latent Flow Matching for Ultra-High-Resolution Image-to-3D Generation
    Hongjie Li, Xinran Yang, Xiuchao Wu, Jiangjing Lyu +1

    Scaling image-to-3D generation to ultra-high resolutions requires controlling rapidly growing computational costs without sacrificing fine geometric detail. We present \textbf{Filigree3D}, a sparse latent flow-matching framework that generates 3D geometry from a single image at voxel resolutions up

    memory
  249. arxiv:2609.34895 · cs.CV
    LVMT: Video Mask Transformer for Long-term Video Segmentation
    Narges Norouzi, Niccol`o Cavagnero, Idil Esen Zulfikar, Bastian Leibe +2

    Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across tim

    memorybenchmark
  250. arxiv:2609.34893 · cs.RO
    ECHO: Event-Augmented Context with Hindsight and Outlook for Wrist-Only Manipulation
    Xinyue Wang, Yicheng Jiang, Zesen Gan, Junhao He +7

    Learning-based manipulation policies relying on RGB cameras often suffer from degraded observations under extreme exposure. Event cameras mitigate this degradation by asynchronously detecting pixel-level intensity changes to offer a high dynamic range. However, their observations heavily depend on c

    manipulationgripperevent camera
  251. arxiv:2609.34886 · cs.AI
    Fewer Assumptions by Design: A Reusable Skill for LLM-Assisted Verus Verification
    Andrada-Livia Antoneac, Dorel Lucanu, Dragoş Teodor Gavriluţ

    LLM-assisted Verus verification is a less tedious method to verify Rust implementations, but paired with self-referential structures, e.g., Doubly Linked Lists (DLLs)—notoriously difficult to formalise for verification—it becomes a substantially more demanding verification task. Moreover, a spec

    agentllm agent
  252. arxiv:2609.34884 · cs.CV
    SubRot: Signed Gradient Subspace Calibration for VLM Rotation Quantization
    Zhenhao Shang, Haizhao Jing, Haokui Zhang, Guoting Wei +3

    Post-training quantization reduces the deployment cost of vision-language models (VLMs), but preserving multimodal capabilities at low bit widths remains challenging. Existing methods rely on modality- or token-level gradient statistics, which are susceptible to cross-sample variations in visual-to-

    post-trainingbenchmark
  253. arxiv:2609.34879 · cs.AI
    One Readout, Many Repairs: Diffusion-Guided Hierarchical Search for Tool-Agent Repair
    Xiang Xia, Cheng Yan, Fan Xu, Zhijun Fan +2

    Tool agents use large language models to act through external tools, yet successfully executed calls can still leave user requests unfulfilled. Tool-agent repair seeks alternative call sequences that execute successfully and fulfill the original requests. However, repair requires exploring both oper

    benchmark
  254. arxiv:2609.34867 · cs.CV
    P4Q: Co-designing Token Pruning and Quantization for Vision-Language Model Acceleration
    Haizhao Jing, Zhenhao Shang, Haokui Zhang, Rong Xiao +1

    Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions,

    memorypost-training
  255. arxiv:2609.34866 · cs.LG
    From Attention Sensitivity to Layer Role: Revisiting Mixed-Precision Quantization of Transformers
    Nafiseh HosseinpourFardi, Negar Alihadi, Mahmoudreza Babaei, Milad Hosseini +1

    Most post-training quantization pipelines fit each weight matrix to its pretrained counterpart, one matrix at a time. Whether that proxy tracks what an attention block actually computes, or how errors in the Q, K and V projections compound inside the softmax, is rarely checked. We write the objectiv

    post-training
  256. arxiv:2609.34864 · cs.AI
    On the Limits of Metacognitive Monitoring in LLMs
    Dongqi Han, Yifan Yang, Dongsheng Li

    Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier m

    benchmarkevaluator
  257. arxiv:2609.34863 · cs.CV
    Revisit to Segment: Working Memory Distillation for Reasoning Segmentation
    Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu +3

    Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Their generated responses contain reasoning traces and localization proposals that can serve as working memory when revisiting the same image and query. Our

    memorybenchmark
  258. arxiv:2609.34861 · cs.LG
    When Text Matters: Design Principles for Visual Token Pruning in Vision-Language Model
    Minchan Kang, Kyeonghye Park, Seoyoung Cho, Daeshik Kim +1

    Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-based token selection c

    benchmark
  259. arxiv:2609.34860 · cs.LG
    Attention-based Hierarchical Variational Information Bottleneck for Robust Multi-Agent Communication under Variable Bandwidth
    Lukas Koch Vindbjerg, Qi Zhang, Yury Brodskiy, Lukas Esterle

    Learning-based multi-agent communication under limited bandwidth does not only require deciding what to communicate, but also structuring messages so that partial transmissions remain useful. We study this problem under prefix truncation, where only the first part of each message is received. To add

    multi-agent
  260. arxiv:2609.34853 · cs.CV
    EviSplat: Preserving Multi-View Evidence in 3D Gaussian Splatting for Open-Vocabulary Segmentation
    Sungho Moon, Kota Shimomura, Junwoo Park, Wonhyeok Choi +3

    Open-vocabulary 3D scene understanding enables object localization and segmentation from free-form text queries without a fixed category vocabulary. Many recent methods build on 3D Gaussian Splatting and consolidate multi-view observations, such as masked crops from individual views, into language f

    evaluation protocol
  261. arxiv:2609.34851 · cs.RO
    Learning High-Risk High-Precision Motion Control
    Nam Hee Kim, Markus Kirjonen, Perttu Hämäläinen

    Deep reinforcement learning (DRL) algorithms for movement control are typically evaluated and benchmarked on sequential decision tasks where imprecise actions may be corrected with later actions, thus allowing high returns with noisy actions. In contrast, we focus on an under-researched class of hig

    benchmark
  262. arxiv:2609.34849 · cs.LG
    When Sparse Reward Meets Dense Distillation: Training Dynamics of On-Policy Distillation
    Xinke Jiang, Tao Feng, Zhibang Yang, Zhixin Zhang +3

    Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token receives the same sequence-level advantage regardless of its individual contribution. To complement this sparse supervision, a growing family o

    post-training
  263. arxiv:2609.34848 · cs.AI
    Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
    Yugu Li, Zehong Cao, Peizhen Li, Yang Zhang +2

    RLVR provides reliable trajectory-level credit, while OPSD offers dense supervision for token-level credit. This exposes a fundamental coupling when updating step-level credit direction and magnitude with teacher supervision, preventing steps from receiving reliable credit directions and contributio

    benchmark
  264. arxiv:2609.34843 · cs.CV
    ORAV: Benchmarking Audio-Video Generation from Multimodal Contexts
    Jiacheng Hua, Xiaokun Feng, Jiaqi Hua, Chang Liu +2

    Audio-video generation using heterogeneous multimodal references has emerged as a new challenge, requiring both compositional control over generation and grounded understanding of multimodal context. In this paper, we introduce ORAV Bench for Omni Reference Audio-Video Generation, comprising 380 tas

    benchmark
  265. arxiv:2609.34842 · cs.LG
    QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities
    Hanyin Cheng, Linfeng Wang, Zhengbo Qu, Yang Shu +6

    Existing multimodal time series foundation models (TSFMs) typically model heterogeneous modalities through largely shared mechanisms, overlooking the distinct forecasting roles of endogenous and exogenous modalities. In this work, we propose QiYao-M, a role-aware multimodal TSFM that models the two

    benchmark
  266. arxiv:2609.34840 · cs.AI
    Nociception as a Control Primitive: Afferent Channels and Nociceptive Memory for Agents Deployed in One Body
    Wolfgang Maass

    An agent deployed in a single body cannot learn how fast that body wears, because every trial that would reveal its wear resistance wears the body it would protect. We study this \emph{epoch-one} setting, in which the parameters of a fixed-weight policy are set before the body is drawn and never upd

    memoryagent
  267. arxiv:2609.34839 · cs.CL
    OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations
    Faadil Mustun, Chiara Semenzin, Roberto Dessi, Pablo Robin Guerrero +6

    Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute

    benchmarkevaluation protocol
  268. arxiv:2609.34838 · cs.LG
    DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents
    Hanyang Wang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan +6

    On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stal

    benchmark
  269. arxiv:2609.34834 · cs.CV
    Transform-Aligned Learned Features for Lossy Point Cloud Attribute Compression
    Yueru Chen, Pengpeng Yu, Dingquan Li, Wei Gao +2

    Transform-based methods provide an effective framework for point cloud attribute compression by representing attributes as transform coefficients. Introducing learned spatial context into this framework requires mapping spatial representations to the transform domain, but this known basis change is

    benchmark
  270. arxiv:2609.34833 · cs.CV
    Multi-Scale Semantic Mapping in Urban Environments via Observation Calibration and Policy Dependence Regularization
    Runling Long, Junhao Feng, Jia Wan

    Semantic mapping is fundamental to embodied navigation, yet existing methods are developed for indoor environments, where objects exhibit relatively limited scale variation and are observed from a restricted range of viewpoints. Urban environments pose substantially greater challenges: agents must m

    embodied
  271. arxiv:2609.34832 · cs.AI
    BV Loss: Block Verification-Aware Loss for Block Diffusion Speculative Decoding
    Suyoung Kim, Jahyun Koo, Hyeonjin Kim, Inhyeok Bang +4

    Diffusion drafters accelerate speculative decoding by proposing multiple tokens in parallel. Despite recent advances in speculative decoding through sequence-level drafting and verification, existing training objectives remain largely designed around token-level verification. To address this mismatc

    benchmark
  272. arxiv:2609.34831 · cs.LG
    Structured Neural SDEs for Functional Calibration
    Francesco Piatti, Andrea Iannucci, Thomas Cass

    Neural Stochastic Differential Equations (Neural SDEs) provide flexible continuous-time generative models, but generic neural drift and diffusion networks are costly to simulate on long horizons and can give unstable gradients when the training signal is a path functional rather than a pointwise obs

    benchmark
  273. arxiv:2609.34829 · cs.CL
    From Weak Task Specifications to Scientific Extraction Agents: Optimizing Task Construction
    Zixiao Dong, Wei Yang, Zihao Liu, Chenshu Li +3

    Most methods that optimize LLM prompts and agent workflows assume that task-specific output schemas, extraction instructions, and evaluation criteria are predefined. For scientific extraction agents, however, a short task goal may not fully determine these components, while specifying them manually

    agent
  274. arxiv:2609.34826 · cs.CV
    WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
    Yuheng Zha, Yilei Wang, Qiyue Gao, Junrong Chen +4

    Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, w

    world model
  275. arxiv:2609.34823 · cs.RO
    AGRO-SUVIDE: Agentic Robotics for Surgical Viscoelastic Debridement
    Shutong Jin, Ziyang Chen, Preethi Satish, Medow Shen +3

    Augmented dexterity has the potential to reduce the fatigue experienced by surgeons during repetitive surgical tasks. In this paper, we propose the first AGentic RObotics framework for SUrgical VIscoelastic DEbridement (AGRO-SUVIDE), the repeated removal of small fragments attached to a viscoelastic

    agentagenticself-improving
  276. arxiv:2609.34821 · cs.MA
    CEO Arena: Evaluating Long-Horizon Multi-Agent Decision-Making in Competitive Markets
    An Yan, Yu Huo, Zhiwei Shang, Yiran Peng +1

    Long-horizon competition tests agents' ability to coordinate business decisions under uncertainty and adapt to changing rival strategies. We introduce CEO Arena, a benchmark that uses matched replacement evaluation to assess operating returns alongside an agent's effects on rivals and the market. Ea

    agentmulti-agentbenchmarkarena
  277. arxiv:2609.34817 · cs.CV
    ESTHER: Egocentric Stereo Hand Estimation and Reconstruction in the Wild
    Hongyu Ma, Hairong Qu, Shiqi Zhao, Yongsong Yang +1

    Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still ha

    manipulationbenchmark
  278. arxiv:2609.34810 · cs.AI
    UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning
    Zenghuang Fu, Zhaoyang Li, Qiuyuan Ai, Xiaofeng Han +7

    Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by ev

    agentic
  279. arxiv:2609.34807 · cs.CV
    ControlTrace: Recovering Control Fields for Hidden-Content Recognition
    Zijian Liu, Yaoguang Chen, Liwei Liu, Weixi Wu +3

    Spatially conditioned diffusion models can embed words and contours in natural-looking images, but vision-language models (VLMs) may fail to recognize the hidden content. Transformation-based recovery depends on parameter and view selection. To evaluate hidden-content recovery and recognition, we co

    benchmark
  280. arxiv:2609.34805 · cs.AI
    SIPO: Selective-Inference Policy Optimization for Tree-Structured Agentic RL
    Zenghuang Fu, Ningqi Chen, Mingda Jia, Xiaofeng Han +9

    Tree-structured reinforcement learning trains search agents by comparing alternative continuations and propagating terminal rewards to intermediate decisions. Adaptive expansion, however, creates a statistical asymmetry: an incumbent is selected using its own generation statistic, whereas fresh sibl

    agenticbenchmark
  281. arxiv:2609.34804 · cs.LG
    Physics-Attested Federated Learning: Securing Collaborative Anomaly Detection in Critical Water Infrastructure
    Jeff Nijsse, Shu Su, Benjamin Oholeguy, Sreenivas Sremath Tirumala

    Federated learning enables industrial operators to train shared intrusion detection models without disclosing proprietary operational telemetry. However, existing defenses operate strictly in update space, leaving aggregators blind to data poisoning; model updates derived from fabricated telemetry r

    benchmark
  282. arxiv:2609.34800 · cs.CL
    Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks
    Panagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis

    The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in th

    benchmark
  283. arxiv:2609.34799 · cs.AI
    STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts
    Wanchun Ni, Tao Qi, Leonel Aguilar, Jiugeng Sun +3

    Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every s

    evaluation frameworkscalable evaluationscalable evalevaluation protocol
  284. arxiv:2609.34798 · cs.CV
    InfiMed2: A Generalist Medical Multimodal Foundation Model from Contextual Evidence and Stability-Aware Supervision
    Guanghao Zhu, Zeyu Liu, Zhitian Hou, Pengkai Wang +8

    Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity,

    post-trainingbenchmark
  285. arxiv:2609.34792 · cs.CV
    D$^2$-VLA: Dual-Memory Dual-Frequency Vision-Language-Action Model For Long Dynamic Manipulation
    Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng +8

    Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-action (VLA) policies often rely on the latest observation, and refreshing their visual context typically requires another costly vision-language model (VLM)

    vision-language-actionmanipulationliberorobotwinmemorybenchmark
  286. arxiv:2609.34790 · cs.AI
    CoSec: Benchmarking Agent Security in Communities
    Hao Chen, Wenhui Dong, Ye Chen, Jiezhi Yao +11

    LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks and prevent unauthor

    persistent memoryagentllm agentagent systembenchmark
  287. arxiv:2609.34785 · cs.LG
    BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification
    Yuheng Wu, Berk Gokmen, Sujeeth Jinesh, Lauren McLane +4

    Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design and reference outputs cycle by cycle can reject valid designs. To address this, we introduce BEHAVE,

    agentagenticself-improvingself-improvementevaluator
  288. arxiv:2609.34783 · cs.CL
    TQTS-Bench: A Multi-Syntax Benchmark for Text-to-Query over Time-Series Databases
    Fei Lyu, Zhiyi Peng, Jiaming Liu, Yixuan Yang +4

    Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domain

    benchmark
  289. arxiv:2609.34782 · cs.RO
    CoHuB: A Simulation Benchmark for Multi-Humanoid Collaboration
    Hyunjin Park, Jebeom Chae, Minwoo Park, Sunghyun Park +10

    Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduc

    humanoidteleoperationbenchmark
  290. arxiv:2609.34781 · cs.CV
    When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
    Yuxing Cheng, Yuan Wu, Yi Chang

    Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of

    benchmark
  291. arxiv:2609.34776 · cs.AI
    Page-Aware Retrieval-Augmented Generation for EvalLLM 2026: A Five-Variant Study on French PDFs
    Abdelhak kelious

    We study retrieval-augmented generation (RAG) for questions about French PDF documents when both the answer and its supporting document pages are evaluated. Five system variants add dense retrieval, rank fusion, reranking, and query decomposition to a BM25 baseline. On 595 challenge questions, the c

    retrieval-augmented
  292. arxiv:2609.34772 · cs.AI
    Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs
    Yadong Wang, Siping Yue, Yu Tian, Chuanxing Geng +1

    Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot deter

    benchmark
  293. arxiv:2609.34769 · cs.AI
    LongPuzzleBench: Evaluating GUI Agents on Long-Horizon Visual Puzzles
    Bingo Zhang, Haochuan Lu, Zongjie Li, Genjian Li +2

    GUI agents need long-horizon visual reasoning: they must interpret a changing interface while keeping a multi-step plan viable as earlier actions constrain later ones. Existing benchmarks evaluate grounding, computer use, and game play, but rarely test whether agents stay coherent across long chains

    memorybenchmark
  294. arxiv:2609.34768 · cs.CV
    Privacy-Preserving Full-Body Meshing from mmWave Radar via Mesh Foundation Model Supervision
    Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin

    Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip sensors (mean ~6.5 points/frame; ~28% empty frames) has confined prior art to body-part keypoints or discrete action classification. We present a cross-mod

    benchmark
  295. arxiv:2609.34765 · cs.LG
    Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models
    Minchan Kang, Kyeonghye Park, Seungyeon Sa, Seoyoung Cho +2

    Post-training quantization (PTQ) enables efficient deployment of large vision-language models (LVLMs), but is typically calibrated on a small set while expected to generalize across diverse downstream tasks. Although recent PTQ methods for LVLMs incorporate sensitivity signals, they still minimize r

    post-training
  296. arxiv:2609.34764 · cs.AI
    WeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM Plans
    Min Jia, Shihao Zhou, Jun-Peng Zhu, Peng Cai +8

    Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. How

    knowledge graphself-evolving
  297. arxiv:2609.34759 · cs.RO
    PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
    Zhen Wang, Changpeng Wang, Zhe Liu, Zhangyang Qi +4

    Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward

    quadruped
  298. arxiv:2609.34754 · cs.CL
    Draft-KV: Learning Useful Latent Communication Between Language Models
    Linquan Wu, Shichang Meng, Tianxiang Jiang, Haoyu Yang +6

    Latent communication passes internal states between language models instead of decoded text, but higher receiver accuracy does not show that the receiver used the message content. Across five method-dataset pairs, replacing each message with one from an unrelated question changes accuracy by at most

    memory
  299. arxiv:2609.34749 · cs.CV
    CoDrive: Cross-Vehicle World-Consistent Video Generation with Precise Trajectory Control for Cooperative Driving
    Yu Meng, Baining Zhao, Junta Wu, Tengfei Wang +9

    Real-world driving is inherently multi-agent, yet most existing driving world models generate observations from a single ego vehicle. Independently extending them to multiple vehicles does not ensure that different agents observe a consistent shared world. We present CoDrive, a cross-vehicle, multi-

    world modelmulti-agentbenchmark
  300. arxiv:2609.34745 · cs.LG
    No Pain, More Gain: Iterative Merging for Effective Multi-Teacher On-Policy Distillation
    Seonghyeon Kim, Chaeyun Jang, Noah Lee, Boseop Kim +1

    Multi-teacher on-policy distillation (MOPD) combines independently developed domain teachers into a single student by distilling their predictions on student-generated samples. We study a setting where teachers share a reference model but undergo different post-training procedures, and find that MOP

    post-trainingbenchmark
  301. arxiv:2609.34743 · cs.RO
    Simulation for Planetary Robotic Perception and Autonomy: A Concise Survey of Recent Capabilities and Gaps
    Hoyun Kim, Giseop Kim

    Planetary robotics is an important enabler of scientific exploration in environments where direct human-in-the-loop operation is costly, hazardous, or infeasible. However, developing and validating planetary robotic systems remains difficult because representative field testing is expensive, limited

    human-in-the-loop
  302. arxiv:2609.34730 · cs.RO
    Action Sequence Transfer via LLMs for Heterogeneous Environments
    Choongho Chung, DongHwan Shin, Sung-Hee Lee

    We present an action sequence transfer system that adaptively transfers user action sequences across different target spaces. Given an input action sequence from a source space and scene graph representations of both the source and target environments, our system predicts a corresponding action sequ

    scene graph
  303. arxiv:2609.34727 · cs.AI
    Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs
    Zhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun +5

    On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and cost. Although KV caching is widely used to reduce latency in long-context inference, existing design

    memorylong-context
  304. arxiv:2609.34724 · cs.RO
    DexWeave: Learning Dexterous Humanoid Loco-Manipulation from Human Demonstrations
    Naichuan Sun, Haotian Shen, Yizhang Zhang, Luying Feng +4

    Learning dexterous humanoid loco-manipulation from human demonstrations requires transferring not only human motion, but also the coordinated interaction structure underlying the demonstrated behavior. This is challenging because embodiment differences distort the coupling among body motion, wrist p

    manipulationdexteroushumanoid
  305. arxiv:2609.34722 · cs.CV
    Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
    Zesong Yang, Weikai Chen, Liyuan Cui, Lutao Jiang +9

    Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric

    memory
  306. arxiv:2609.34720 · cs.CV
    DBCF: Dual-Branch Complementary Fusion of Foundation Models for Generalized Deepfake Detection
    Fengming Gu, Mingjie He, Zonghui Guo, Jie Zhangb +1

    As image generation and editing technologies have progressed substantially, facial forgeries pose significant challenges to privacy and public safety. Due to limited ability to capture forgery cues, existing small-scale forgery detection models often struggle to generalize across various domains and

    manipulationbenchmark
  307. arxiv:2609.34718 · cs.LG
    Quality Determines Direction, Length Shapes Magnitude: Length Control for Open-Ended Reinforcement Learning
    Zijun Weng, Zhongan Bi, Xuanang Gao, Xiaohui Hu +4

    Reinforcement learning (RL) changes not only what language models say, but also how much they say, often increasing response length at the cost of token efficiency. Controlling this length growth is particularly challenging in open-ended RL because (i) response length is entangled with quality, (ii)

    benchmark
  308. arxiv:2609.34717 · cs.CL
    ReMCTS: Reflection-Enhanced Monte Carlo Tree Search for Code Generation
    Huifei Wang, Xinying Huang, Yiheng Sun, Yifan Yuan

    Open-weight large language models (LLMs) can generate function-level programs from natural-language prompts, but plausible candidates still fail on hidden semantics and repeat mistakes across repair attempts. We present ReMCTS, an execution-grounded, memory-augmented, LLM-guided MCTS-style search fr

    memory
  309. arxiv:2609.34715 · cs.AI
    PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs
    Zhentao Tan, Jianrong Zhang, Ruijie Quan, Yi Yang

    Physical trajectories contain more than snapshots of a system: they also reveal how its states evolve under governing conditions. However, representation learning for parametric partial differential equations (PDEs) has largely relied on reconstruction-based objectives that emphasize recovering obse

    latent dynamicsbenchmark
  310. arxiv:2609.34712 · cs.AI
    RSI-Router: Evolving Subtask-Level LLM Routing and Skills for Cost-Efficient Agents
    Hao Li, Hangfan Zhang, Zhiyao Cui, Chunjiang Mu +4

    Practical deployment of large language model (LLM) agents requires strong task performance at affordable inference cost. For long-horizon agentic tasks, this performance-cost trade-off can be improved through within-task large-small model collaboration, as smaller models can handle some stages even

    agenticself-improvementbenchmark
  311. arxiv:2609.34711 · cs.LG
    Learning Propagation Geometry from Message-Passing Feedback
    Yingxu Wang, Kunyu Zhang, Xinwang Liu, Mengzhu Wang +3

    Learning local geometry enables graph neural networks (GNNs) to adapt how they compare and integrate neighborhood information. However, estimating geometry from aggregated representations can overlook variation among individual messages and dependencies across feature dimensions. We propose GeoF, a

    benchmark
  312. arxiv:2609.34710 · cs.AI
    FromPitch2Board: Benchmarking LLM Agents in Long-Horizon Football Management
    Peiyu Zang

    Long-horizon agent benchmarks typically report how far an agent progresses, but do not identify whether its performance comes from the foundation model, scaffold, responsibility scope, match-control granularity, or horizon. We introduce FromPitch2Board, a deterministic football-management benchmark

    agentllm agentagent benchmarkbenchmark
  313. arxiv:2609.34707 · cs.RO
    SOR-Nav: Search or Relocate? Context-Gated Exploration and Cross-Region Relocation for Object Navigation
    Yuan Ji, Zirui Li, Yuxin Cai, Shuge Wu +2

    Object navigation requires an embodied agent to find an object in an unseen environment under partial observability and a limited motion budget. Existing methods primarily optimize where the robot should go next by ranking candidate destinations. In contrast to these methods, we present SOR-Nav, a h

    embodiedagentembodied agentbenchmark
  314. arxiv:2609.34702 · cs.RO
    MarsLab: A Martian Rover Simulator for Planetary Rover Autonomous Navigation
    Hoyun Kim, Beomsu Kim, Giseop Kim

    Future Mars missions will require rover autonomy that can operate across unstructured terrain, changing illumination, atmospheric dust, and limited communication. Simulation is a practical way to study these conditions before deployment, but existing Mars-relevant resources differ in scope, includin

    benchmark
  315. arxiv:2609.34701 · cs.LG
    ResonAct: Streaming Metrics for Runtime Diagnosis and Self-Healing in Multi-Agent Systems
    Tarun Chintada, Neelamadhav Gantayat, Ishaan Romil, Renuka Sindhgatta +2

    Multi-agent systems (MAS) are increasingly used to automate enterprise workflows involving multiple specialized agents, external tools, and long-running task execution. Failures may arise from tool degradation, context propagation errors, coordination breakdowns, or repeated agent interactions that

    agentmulti-agentagent systembenchmark
  316. arxiv:2609.34697 · cs.CV
    Triangular Resampling for Long-Horizon Motion Generation
    Kunhang Li, Yiyi Cai, Xiangyue Zhang, Fangyuan Tu +4

    We introduce Triangular Resampling (TR), a post-training method for mitigating long-horizon error accumulation in motion diffusion models. Built on FloodDiffusion's triangular denoising schedule, TR addresses the mismatch between ground-truth-derived training windows and model-generated inference st

    post-training
  317. arxiv:2609.34691 · cs.AI
    Using LLMs to Detect LLM-Generated Texts: A Cross-Generation Analysis
    Haiyue Yuan, Jie Guo, Weidong Qiu, Zheng Huang +2

    Automated detection of LLM-generated texts (LGTs) is critical, yet dedicated detectors often struggle to generalize across domains and models. While general-purpose LLMs offer flexible zero-shot authorship classification with explanatory rationale, their detection behavior, especially regarding self

    benchmark
  318. arxiv:2609.34688 · cs.CV
    Unified Trajectory Matching Policy Optimization: Diverse T2I Generation and VLA Generalization
    Zhiyuan Ma, Jiaming Li, Lingzhen Li, Yu Liu +6

    Reward-maximizing reinforcement learning (RL) is widely used to post-train stochastic diffusion and flow policies for text-to-image (T2I) generation. However, reward-maximizing RL causes policy mode collapse even under reference KL or entropy regularization, reducing the policy to a single high-rewa

    vision-language-actionvlapost-training
  319. arxiv:2609.34687 · cs.AI
    VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience
    Siqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu +5

    Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning

    embodiedagentembodied agentbenchmark
  320. arxiv:2609.34686 · cs.AI
    Jailbreak Context Lingers: Divergent Safety Routing and Its Cross-Task Predictability in Tool Agents
    Xi Wang, Songlei Jian, Yiming Zhang, Bin Ji +4

    As large language models increasingly operate as tool-using agents, post-jailbreak safety feedback is often assumed to serve as a reliable safeguard; however, how lingering jailbreak context shapes subsequent agent behavior remains largely unexplored. To systematically examine this dynamic, we intro

    agentautonomous agent
  321. arxiv:2609.34684 · cs.RO
    Natural State-Prediction Accuracy can Hide Weak Controlled Responsiveness in VLA Readouts
    Hyungjoon Kim, Wonbin Son, Mi Young Lee, Jun Young Lee +1

    Accurately decoding object states from the internal representations of vision-language-action (VLA) models does not establish that the predictions respond faithfully to changes in the target physical state. In natural observations, object state, robot configuration, occlusion, and task progress vary

    vision-language-actionvlaevaluation framework
  322. arxiv:2609.34683 · cs.LG
    AgentPerfBench: A Benchmarking and Evaluation Suite for Inference Performance of Agentic LLMs
    Cheuk Hang Lau, Zeyu Cao, Kevin Wong Cheuk Yin, Yao Lai +5

    The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system designs are all selected based on representative workloads. However, a significant mismatch has emerged in the agentic era. Existing benchmarks prima

    memoryagentictool-usetool callingbenchmark
  323. arxiv:2609.34682 · cs.CV
    V-Gym: Enhancing Agentic Visual Reasoning via Skill-Data Co-Evolution
    Bei Yan, Yuecong Min, Jie Zhang, Junqi Yang +2

    Advances in multimodal understanding, reasoning, and tool use enable agents to tackle increasingly complex visual reasoning tasks. By distilling past execution experience into reusable skills, agents can transfer lessons from both successes and failures into future reasoning, reducing repeated error

    agentictool useself-improvementbenchmark
  324. arxiv:2609.34681 · cs.LG
    SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
    Qiulin Shang, Binyu Wang, Yongqi Qiao, Songde Rao +2

    Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimiza

    online learning
  325. arxiv:2609.34680 · cs.LG
    QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization
    Qiulin Shang, Zhoutong Wu, Jie Hu, Kun Yuan

    Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the netw

    memorypost-trainingevaluator
  326. arxiv:2609.34679 · cs.LG
    From Preference to Reciprocity: Decentralized Matching with Empirically Grounded LLM-agent Based Modeling
    Wangxuan Fan, Xiaoyu Nie, Zhoutian Shi, Xiangcheng Meng +4

    Bipartite matching is a fundamental problem in game theory and market design. Classical approaches such as Gale--Shapley assume complete preferences and centralized computation, whereas many real-world matching processes are decentralized, asynchronous, and shaped by sequential interaction under lim

    llm agentonline learning
  327. arxiv:2609.34678 · cs.AI
    Fair Fact-Checking: Closing the Cross-Lingual Gap in LLM Factual Judgement with RoSh
    Muhammad Ahmad, Fatemeh Seyedin, Adrian Weller, Dongwon Lee +1

    Misinformation on social media remains a critical problem, and more and more people settle it by asking a language model instead of a fact checker. Whether models judge such claims reliably is debated; whether they judge them equally well in every language people ask in has gone almost unasked. We t

    benchmark
  328. arxiv:2609.34677 · cs.LG
    Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
    Beomsu Kim, Chieh-Hsin Lai, Bac Nguyen, Amir Bar +2

    World models predict future observations from current experience and actions, yet prediction can depend on observations seen far in the past. Episodic memory preserves past observations for later recall; however, as memory accumulates, it raises a fundamental question: which memories are useful for

    world modelmemoryepisodic memory
  329. arxiv:2609.34674 · cs.RO
    HOI-Retarget: Contact-Centric Retargeting for Human-Object Interaction
    Jihwan Shin, Adrià López Escoriza, Junzhe He, Matthias Heyrman +1

    Learning from demonstration (LfD) has enabled humanoid robots to acquire diverse whole-body skills, but extending this paradigm to human-object interaction (HOI) is limited by the availability of robot-compatible interaction references. We present HOI-Retarget, a contact-centric retargeting method t

    humanoid
  330. arxiv:2609.34666 · cs.RO
    On the Numerical Reliability of Differentiable Physics-Based Optimization for Robotic Material Manipulation
    Xintong Yang, Minglun Wei, Yu-Kun Lai, Ze Ji

    Differentiable physics is increasingly used in robotic material manipulation for system identification, trajectory or skill optimization, demonstration generation, and robot or end-effector design. These applications depend on gradients propagated through long, contact-rich simulation rollouts. We s

    manipulationbenchmark
  331. arxiv:2609.34661 · cs.AI
    Codoku: Renewable Program-Reasoning Challenges for Frontier Coding Agents
    Cong Li, Hao Sun, Zenan Li, Zhendong Su

    Existing program-reasoning benchmarks ask large language models to predict a program's behavior on a given input. Coding agents break two assumptions on which these benchmarks rest: an agent can recover the answer by executing the program instead of reasoning about it, and fixed task sets drawn from

    agenttool usebenchmark
  332. arxiv:2609.34660 · cs.CL
    Rewarding Novel Deductions: Solver-guided Process Rewards for Logical Reasoning
    Muhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza

    Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing i

    benchmark
  333. arxiv:2609.34658 · cs.CV
    CapField-OPD: Learning Continuous Capability Fields via Joint-Anchored Multi-Teacher On-Policy Distillation for Flow Models
    Pengyang Ling, Yujie Zhou, Jiazi Bu, Yibin Wang +4

    Reward-specialized post-training produces strong experts for flow-based generative models, while multi-teacher on-policy distillation (OPD) consolidates their capabilities into a single student. Existing methods, however, route each prompt to a single teacher according to its semantic category, impl

    post-training
  334. arxiv:2609.34656 · cs.LG
    Minimax Last-Iterate Convergence in Matrix Games with Observed Actions
    Yuheng Zhang

    We study last-iterate convergence in unknown two-player zero-sum matrix games with bandit payoff feedback and observed opponent actions. For games with $d$ actions per player, we develop an algorithm achieving a duality gap of $\widetilde{\mathcal{O}}(\sqrt{d/t})$ with high probability, simultaneous

    memory
  335. arxiv:2609.34654 · cs.AI
    A General Harness for Protein Foundation Model Fitness Prediction
    Yang Tan, Qijia Tian, Gangyu Sun, Bozitao Zhong +4

    Accurate fitness prediction is central to protein engineering and understanding sequence-function relationships. With advances in deep learning, protein foundation models (PFMs) have become widely used for this task. Recent analyses, however, show that these models share preferences reflecting their

    benchmark
  336. arxiv:2609.34653 · cs.LG
    OmniTide: Co-Designing Algorithms and Systems for Efficient On-Device Omni-LLM Streaming
    Zongshang Shen, Wangsong Yin, Daliang Xu, Mengwei Xu +1

    On-device streaming omni-modal inference safeguards user privacy and eliminates prohibitive per-token API costs, but faces a critical bottleneck: the continuous influx of multimodal data rapidly exhausts constrained memory and compute budgets via monotonic KV cache growth. Existing sparse attention

    memorybenchmark
  337. arxiv:2609.34652 · cs.CV
    Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective
    Guoqi Yu, Juncheng Wang, Shujun Wang

    Medical time series (MedTS) underpin many clinical classification tasks, yet existing methods usually represent them only as numerical sequences and underuse the morphology that is explicit in waveform inspection. To bridge this gap, we introduce Vision-Informed Retrieval (ViRe), which uses a frozen

    benchmark
  338. arxiv:2609.34649 · cs.AI
    Beyond Skill Evolution: Self-Evolving Context Management Policies for Long-Horizon Agent Harnesses
    Weiyuan Li, Jinghan Xu, Aili Chen, Xintao Wang +3

    Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottl

    agentllm agentself-evolvingbenchmark
  339. arxiv:2609.34648 · cs.AI
    SEmoEdit: Probing and Harnessing the Editability of Pre-trained Speech Flows
    Tianxin Xie, Pengfei Zhang, Kai Jiang, Zelin Zhao +1

    Existing training-based speech emotion editing methods often require substantial task-specific training and can be unstable. This motivates us to investigate whether the pretrained generative dynamics of large-scale text-to-speech (TTS) models can be directly manipulated for training-free emotion ed

    benchmark
  340. arxiv:2609.34645 · cs.AI
    Nereus: Adaptive Parallelism for LLM Post-Training
    Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang +2

    Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. Several factors may change during a run, including resource availability, sequence length, memory pressure, and stage bottlenecks. As a c

    memorypost-training
  341. arxiv:2609.34642 · cs.LG
    Tilted Schrödinger Bridge Matching
    Sergei Kholkin, Evgeny Burnaev, Alexander Korotin

    Schrödinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion model

    post-training
  342. arxiv:2609.34641 · cs.CV
    Backdoor as Probe: Test-Time Adversarial Defense for CLIP
    Zhongqi Wang, Jie Zhang, Nie Sen, Zhiyu Chen +2

    Test-time adversarial defense improves the robustness of vision-language foundation models such as CLIP without retraining. However, adversarial activation shifts are typically treated as distortions to suppress, rather than signals to exploit. We turn these shifts into defense signals by repurposin

    benchmark
  343. arxiv:2609.34636 · cs.AI
    MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics
    Danilo Gusicuma, André Freitas

    This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying stru

    benchmark
  344. arxiv:2609.34634 · cs.AI
    A Persistent State for Auditable Mixture-of-Experts Routing
    Abdurrahman Javat, Allan Kazakov

    Mixture-of-Experts (MoE) models repeatedly route tokens to sparse subsets of experts, but conventional routers expose no routing-specific record of how cross-layer influences accumulate. We introduce Scratchpad-Augmented Mixture-of-Experts (SA-MoE), which gives each router access to a low-dimensiona

    persistent state
  345. arxiv:2609.34633 · cs.LG
    GenMem: Generative Symbolic Memory for Self-Evolving Harness
    Xinke Jiang, Tao Feng, Weixuan Xu, Zhixin Zhang +6

    Long-term memory supports the self-evolution of LLM agents by retaining experience and skills across tasks and enabling their retrieval, reuse, and revision in subsequent long-horizon decision-making. Yet existing memory management approaches remain limited to discriminative retrieval and to address

    memoryagentllm agentmulti-agentself-evolving
  346. arxiv:2609.34630 · cs.CV
    Long Time No See: Benchmarking VLMs for Out-of-Sight Spatiotemporal Reasoning in Egocentric Videos
    Fangzhou Ma, Ivo Alexander Ban, Eren Homburg, Gabriele Goletto +4

    Real-world AI systems must reason about objects that are no longer visible: an AR assistant guiding a user back to an object used earlier, a household robot retrieving an item someone put away. This requires not just recalling where an object was last seen, but updating its state when it is moved an

    benchmark
  347. arxiv:2609.34629 · cs.LG
    DisKO: Deep Koopman Learning in Distribution Space from Unpaired Snapshots
    He Ma, Xiaochen Liu, Wanfeng Lu, Ying Wang +2

    Many complex systems are observed only through temporally unpaired distribution snapshots, making trajectory-based dynamical learning difficult without additional assumptions. We therefore formulate the problem directly in distribution space, treating the distribution itself as the dynamical state.

    benchmark
  348. arxiv:2609.34619 · cs.AI
    LLMs for Executable Multi-Agent System Specification Generation
    Andreas Kouvaras, Periklis Mantenoglou, Alexander Artikis

    MAS specifications express the effects of the actions of the agents and their environment, as well as other temporal phenomena, such as the intervals during which an agent may perform an action. The specification of a MAS should also be executable in order to allow for run-time monitoring. Construct

    agentmulti-agentagent system
  349. arxiv:2609.34613 · cs.LG
    Probabilistic Geodesic Flow Matching on Location-Scale Families
    Zeyuan Yu, Zhi Chang, Shiwei Lan

    Flow matching (FM) has recently emerged as a promising framework for generative modeling due to its conceptual simplicity and strong empirical performance. In FM, samples are transported along a vector field parameterized by a neural network, inducing a probability path that evolves from a simple no

    benchmark
  350. arxiv:2609.34608 · cs.RO
    Efficient World Action Model Inference with Adaptive Intermediate States
    Zhinnan Liu, Haozhi Han, Ruge Zhang, Teng Ma +7

    World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, obser

    liberorobotwin
  351. arxiv:2609.34606 · cs.CV
    WorldAttention: An Efficient Attention Architecture for Interactive Video World Models
    Zeyu Zhang, Jinyuan Mao, Dakai An, Wangbo Zhao +5

    Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current

    embodiedworld modelmemory
  352. arxiv:2609.34605 · cs.LG
    PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
    Youzhi Liu, Ruobing Zheng, Boyuan Tong, Tianqi Li +4

    Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teach

    post-training
  353. arxiv:2609.34604 · cs.LG
    Shaping Persistent Representations from Independent Interactions
    Ji Dai, Quan Fang, Junyu Gao, Rongfeng Guo +3

    World models learn environment dynamics from interaction experience. These dynamics depend on the current state and actions, as well as on properties that persist across interactions. Yet standard predictive training can reduce error using local evidence alone, without organizing persistent informat

    tactileworld model
  354. arxiv:2609.34603 · cs.AI
    After the Fix: How Corrected Agent Histories Transfer to Related Tasks
    Yanfei Zhang, Xu Lin

    Does repairing an episode make its experience a better memory for the next task? We transfer the same failed source before and after accepted repair to a fixed target, alongside independent execution. Our 3,300 runs cover 100 ThinkingBox pairs and the same 100 APEX pairs with and without source-stat

    memoryagent
  355. arxiv:2609.34599 · cs.LG
    The Low-Rank Structure of VLA Reinforcement Learning
    Minjae Oh, Yoonah Park, Jongwon Lim, Yohan Jo

    Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including $π_{0.5}$ and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CA

    vision-language-actionvlavla modelgr00tliberopost-training
  356. arxiv:2609.34598 · cs.CV
    Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding
    Nanxing Hu, Xiaoyue Duan, Qiwei Yan, Kailin Lyu +2

    Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs

    memory
  357. arxiv:2609.34596 · cs.CV
    Temporal Modelling for Burn Scars on Sentinel-3
    Luca Barco, Edoardo Arnaudo, Andrea Bragagnolo, Claudio Rossi +1

    Rapid and accurate burn scar delineation from satellite imagery is essential for post-fire damage assessment. Sentinel-3 OLCI, with daily revisit and 21 spectral bands, suits rapid mapping, yet most pipelines treat acquisitions independently, leaving the pre/post-fire change signal unexploited. We p

    benchmark
  358. arxiv:2609.34594 · cs.RO
    Aerial GRIPPER: A Gradient-based Real-time Inverse-game Predictor and Planner
    Zeshuai Chen, Meng Wang, Jindou Jia, Xiang Yu +1

    Accurate capture of non-cooperative targets is critical. In an attempt to tackle this intractable challenge, an aerial gripper system integrated with a Gradient-based Real-time Inverse-game Predictor and PlannER (GRIPPER) framework is proposed. The interaction is formulated as a general-sum pursuit-

    grippergrasp
  359. arxiv:2609.34593 · cs.LG
    Deep Weighted Bellman Residual Minimization for $Q^*$ Estimation
    Lican Kang, Jerry Zhijian Yang, Cheng Yuan, Chen Zhong

    Off-policy evaluation is a foundational component of offline reinforcement learning, aiming to assess and optimize policy performance using pre-collected datasets. However, such datasets often suffer from pronounced challenges, including distribution shift, $Q$-value overestimation, and low sample u

    policy evaluation
  360. arxiv:2609.34590 · cs.LG
    The Model Knows When to Stop: Training-Free Early Stopping for Long-Context Reading
    Muath Alyobi, Mohamed Eltahir, Almoayyad Abuljdail, Riyadh Almutawa +2

    Language models often process long inputs sequentially in chunks, but continuing to read after sufficient evidence has been acquired wastes computation. Existing stopping mechanisms either learn sufficiency from internal activations or train an exit gate, while a simpler alternative asks the model w

    long-contextbenchmark
  361. arxiv:2609.34587 · cs.CV
    Reinforcement Learning from Intermediate Renders for Image-to-Code Generation
    Omri Kaduri, Kate Feingold, Phillip Isola, Tali Dekel

    Reinforcement learning is increasingly used to post-train vision-language models for image-to-code generation, such as generating SVG code from a reference image, by optimizing rewards computed from the final rendered output. However, relying on a single terminal reward provides sparse feedback that

    post-training
  362. arxiv:2609.34586 · cs.LG
    AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum
    Michalis Kasioulis, Moysis Symeonides, George Pallis, Marios D. Dikaiakos

    Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent development, observability, or benchmarking s

    agentagenticbenchmark
  363. arxiv:2609.34585 · cs.LG
    Compute Time Scaling with Recursive Models for Combinatorial Optimization
    Zhengxi Zhang, Paul Swoboda

    We propose Tiny Recursive Models for Combinatorial Optimization (\ours{}), a general neural method for combinatorial optimization that scales both depth (how often we recursively invoke our network) and width (how much we sample in parallel). Both are fundamental for combinatorial optimization: hard

    benchmark
  364. arxiv:2609.34581 · cs.CV
    Counterfactual Attention Policy Distillation for Temporal Video Grounding
    Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu +4

    Temporal video grounding is a key capability of advanced \emph{Multimodal Large Language Models} (MLLMs) for the thorough understanding of video events, which is however often limited by repeated actions and visually similar contexts in long videos. In this paper, we study this issue from the perspe

    benchmark
  365. arxiv:2609.34575 · cs.AI
    Diffusion Subgoal Planning for Long-Horizon Offline Goal-Conditioned Reinforcement Learning
    Hengrui Zhang, Yuhu Cheng, C. L. Philip Chen, Xuesong Wang

    Offline goal-conditioned reinforcement learning (GCRL) learns goal-directed policies from reward-free data, but in long-horizon tasks, goal-conditioned value functions often provide unstable guidance due to sparse rewards and discounting. Hierarchical methods partially mitigate this issue via subgoa

    manipulationbenchmark
  366. arxiv:2609.34572 · cs.LG
    Nudgeability: Reasoning Models Follow Confidence Signals Without Tracking Their Own Competence
    Rohit Saxena, Utkarsh Upadhyay

    Reasoning language models that can call tools must decide during inference whether to answer unaided or delegate. Any self-reflection mechanism for this must answer three questions: where the reflective signal comes from (verbal reports, output distributions, hidden states, a separate predictor), ho

    post-training
  367. arxiv:2609.34571 · cs.AI
    PersonaManifold: Revealing and Exploiting Curved Geometry in LLM Persona Representations
    Rui Xu, Yinghui Xu, Libo Wu

    Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations---addition, scaling, and linear interpolation---u

    benchmark
  368. arxiv:2609.34567 · cs.CV
    Recent Advances in Agentic Agri-Robotic Phenotyping: A Perspective Review from Fragmented Multimodal Sensing to Unified PhenoAgent Intelligence
    Muhammad Owais, Ehtesham Iqbal, Samee Ullah Khan, Muhammad Umraiz +2

    This review examines the evolution of plant phenotyping from conventional manual trait measurement to high-throughput, robotic, and artificial intelligence-driven crop monitoring. Despite significant advances in imaging, autonomous platforms, multimodal sensing, and deep learning, current phenotypin

    agenticagent frameworkbenchmark
  369. arxiv:2609.34565 · cs.AI
    FlowState: Execution State as Memory for Long-Horizon LLM Agents
    Minghao Li, Bangyan Li, Zifan Wang, Yulong Li +4

    Long-horizon tasks require LLM agents to continually draw on information from earlier interactions. However, retaining the full history increases context costs, while compressing it risks losing details needed later, and the relevance of historical information often becomes apparent as the task prog

    memoryllm agent
  370. arxiv:2609.34564 · cs.CV
    GLF-Q: Global-Local Feature-based Quantization for Vision Transformers
    Peilin Sun, Guang Liang, Jin Tong, Jianxin Wu

    Post-training quantization (PTQ) efficiently compresses Vision Transformers (ViTs) without retraining, yet suffers severe accuracy degradation at low bit-widths. Existing optimization-based PTQ methods guide block reconstruction via either soft logits or second-order Hessian proxies. Logit supervisi

    post-training
  371. arxiv:2609.34561 · cs.LG
    Brain-Conditioned Action Policies for Neural Motor Decoding
    Luyao Jin, Running Zhao, Huan Zhao, Vincent C. K. Cheung +1

    Motor brain-computer interfaces (BCIs) aim to decode motor intention, enabling people with paralysis to control external devices. Neural motor decoding typically learns task-specific mappings from neural activity to kinematics, yet remains constrained by scarce paired neural-action data. We propose

    vision-language-actionvlaopenvla
  372. arxiv:2609.34556 · cs.LG
    RoPE is Dead, Long Live RoPE: Towards Scalable Data-aware Positional Encodings
    Jarod Lévy, Mathurin Videau, Jad Yehya, Jean-Rémi King +2

    Transformers process tokens without any inherent notion of order, making positional encoding a fundamental requirement rather than an architectural refinement. Rotary Position Embedding (RoPE) has become the default positional encoding in modern language models, yet it is heavily biased toward nearb

    benchmark
  373. arxiv:2609.34557 · cs.AI
    SkillRubric: Co-Evolving Actor Guidance and Evaluator Rubrics for Multimodal Agents
    Bingqing Jiang, Guoxi Zhang, Jasper Wang, Auric Wang +6

    Recent work incorporates reusable skills distilled from past interactions into multimodal agent training, providing procedural guidance for long-horizon planning and tool use. However, policy optimization in these methods remains driven primarily by sparse outcome rewards, providing little supervisi

    agentagent benchmarktool usebenchmarkevaluator
  374. arxiv:2609.34555 · cs.LG
    PulseInfer: I/O-Centric Sparse KV Cache Offloading for Efficient Long-Context LLM Decoding
    Qiuyang Zhang, Kai Zhou, Kai Lu, Haocheng Lu +5

    Long-context LLM serving is increasingly bottlenecked by decode, where large KV caches limit batch size and underutilize GPUs. Sparse KV cache offloading expands effective capacity by storing most historical KV blocks in CPU DRAM and recalling only selected blocks on demand. However, we find that ex

    long-context
  375. arxiv:2609.34554 · cs.RO
    Where Memory Belongs: Ledger, an Object Ledger for Memory-Augmented VLAs
    Tanguy Dieudonné, Jack B. Jedlicki, Heng Yang

    Memory is essential for long-horizon, partially observed robotic manipulation: a robot must remember which object was placed in a drawer, whose cup it moved, or how many action cycles have elapsed. Recent vision-language-action (VLA) models embed memory directly inside the policy, but benchmarks sho

    vision-language-actionmanipulationmemorybenchmark
  376. arxiv:2609.34550 · cs.RO
    Gaze Prompts: Temporally Dense Human Attention for Vision-Language-Action Fine-Tuning
    Yihan Zhou, Rui Yan, Mingcong Li, Zheyuan Huang +3

    Vision-Language-Action (VLA) fine-tuning pairs images with actions at every step, yet typically provides only a task-level language instruction, leaving moment-to-moment visual relevance implicit. We introduce \emph{eye-tracker-supervised gaze prompting}, which uses gaze recorded during VR teleopera

    vision-language-actionvlamanipulationteleoperation
  377. arxiv:2609.34548 · cs.AI
    SGG-ReflAct: Sub-Goal Guided ReflAct with Structured Planning for Reliable Long-Horizon Reasoning
    Jaeho Jung, Sung Hoon Jung

    Recent advances in reasoning backbones have empowered large language model (LLM)agentstotackle complex, multi-step tasks. However, as reasoning horizons grow, inconsistent internal beliefs induce intermediate errors that cause agents to drift from their goals. This limitation also persists in REFLAC

    llm agent
  378. arxiv:2609.34547 · cs.CV
    ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models
    Gueter Josmy Faure, Min-Hung Chen, Hao Ping Wang, Timothée Lardy +2

    Video-capable vision-language models score above 80\% on popular benchmarks yet struggle with spatial-temporal binding: associating the right action with the right person at the right moment. We introduce ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targe

    benchmark
  379. arxiv:2609.34545 · cs.AI
    Remember Before You're Asked: MemDream for Self-Probing Memory Evolution
    Mingfei Lu, Mengjia Wu, Runsong Jia, Zhe Luo +1

    Memory is essential for enabling LLM-based agents to maintain coherent, personalized behavior over long-horizon interactions. However, existing memory systems share a fundamental limitation: they never proactively test their own memory, repairing it only after real queries expose weaknesses. This re

    memoryllm agent
  380. arxiv:2609.34540 · cs.AI
    APOLO: Automatic Prompt Optimization for Ontology Learning
    Huu Tan Mai, Roman Kochnev, Cuong Xuan Chu, Lukas Lange +2

    Ontology Learning (OL) from text has advanced with the emergence of Large Language Models (LLMs), but it remains challenging due to the limited availability of annotated training data and the difficulty of adapting LLMs to perform OL effectively. We address this via APOLO - Automatic Prompt Optimiza

    multi-agentagent system
  381. arxiv:2609.34539 · cs.CV
    TSGate: Timestep-Aware Gated Attention for Diffusion Transformers
    Boyu Zhang, Yifan Liu, Shuxia Lin, Qingjian Ni +2

    Diffusion Transformers (DiTs) have emerged as the dominant architecture for high-fidelity image and video generation. Recent DiT systems increasingly use structured prompts for training, improving caption quality and prompt adherence. However, their generation quality can degrade severely under out-

    benchmark
  382. arxiv:2609.34537 · cs.AI
    The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions
    Xiaoting Lyu, Xinbo Ma, Yufei Han, Hangwei Qian +4

    Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \te

    benchmark
  383. arxiv:2609.34528 · cs.CV
    Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
    Hyun-Kurl Jang, Jihun Kim, Kuk-Jin Yoon

    Open-vocabulary semantic segmentation (OVSS) enables pixel-level prediction over arbitrary text-specified vocabularies and has shown strong generalization on common benchmarks. However, OVSS performance often degrades in specialized domains such as medical imaging, remote sensing, and industrial ins

    benchmark
  384. arxiv:2609.34526 · cs.AI
    PairPref: When Should Memory Guide the Answer? A Benchmark for Contextual Preference Use
    Mingfei Lu, Mengjia Wu, Yi Zhang

    Memory-augmented assistants use retrieved preferences to guide their responses. A small change in the situation can change whether a preference is appropriate while barely affecting its retrieval similarity. Memory benchmarks typically test whether systems store and retrieve preferences, with less a

    memorybenchmark
  385. arxiv:2609.34510 · cs.AI
    Can AI Make Money in Crypto? Measuring the Gap from Backtests to Real Markets
    Xingtong Yu, Jiarun Zhou, Guanlin Ding, Wenkang Wei +11

    AI-based trading methods have rapidly evolved from machine learning and reinforcement learning to large language models (LLMs) and trading agents, yet their performance is still predominantly assessed through historical backtesting. Such evaluations provide limited evidence of whether a method can g

    benchmark
  386. arxiv:2609.34506 · cs.AI
    Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
    Manya Singh, Arjun Pakrashi

    Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree on labels, a signal of genuine ambiguity. We investigate whether models struggle on

    benchmark
  387. arxiv:2609.34503 · cs.LG
    Distribution-Conditioned Task Routing for Class-Incremental Learning
    Longhuan Xu, Zhipeng Zhou, Wei Ji, Chunyan Miao +2

    Parameter-efficient adaptation enables continual learners to acquire task-specific knowledge through compact model updates while maintaining strong within-task performance. However, class-incremental inference requires each input to be classified among all classes seen so far without access to its t

    benchmark
  388. arxiv:2609.34502 · cs.CV
    SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling
    Xinyu Wang, Huafeng Shi, Zian Li, Yan Zhou +3

    We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retai

    memory
  389. arxiv:2609.34499 · cs.LG
    Scalable GNN-based Knowledge Graph Representation Learning with Efficient Message Passing
    Huu Tan Mai, Cuong Xuan Chu, Heiko Paulheim, Daria Stepanova

    Graph neural networks (GNNs) excel at representation learning on Knowledge Graphs (KGs), achieving stateof-the-art performance on tasks like link prediction or entity classification. However, their high computational complexity, inherent to their user-defined message passing (MP) algorithm, still pr

    knowledge graph
  390. arxiv:2609.34497 · cs.LG
    QAMM: Adjoint MeanFlow Matching for Few-Step Offline Reinforcement Learning
    Yuehu Gong, Shutong Ding, Mokai Pan, Yimiao Zhou +3

    Flow policies can model rich action distributions, but their iterative sampling limits decision speed. Adjoint matching uses the critic's action gradient to improve a flow policy without backpropagating through its sampling trajectory, yet its supervision is defined for instantaneous velocities. We

    humanoid
  391. arxiv:2609.34496 · cs.MA
    MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems
    Yapeng Li, Songze Li, Shuang Yu, Jing Yu +3

    LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost

    agentmulti-agentagent systembenchmark
  392. arxiv:2609.34492 · cs.AI
    PowerBench: A Benchmark for Agentic Retrieval and Reasoning in Power Systems
    Xijing Wang, Yinsheng Yao, Jinru Ding, Yidong Jiang +4

    Large language model (LLM) agents offer new opportunities for automated analysis in industry. However, rigorous evaluation of such agents-for example, within power system scenarios-remains hindered: real operational data are confidential, and existing public resources fail to fully capture the chain

    llm agentagentictool usebenchmark
  393. arxiv:2609.34491 · cs.LG
    M3OS: A Monte Carlo Graph Search-Orchestrated Multi-Agent LLM System for Evidence-Traced Molecular Optimization
    Junjie Wang, Yaowei Jin, Ruohui Tang, Guonan Cui +6

    Small-molecule optimization integrates medicinal-chemistry reasoning and computational evidence through iterative, multi-objective decisions. When large language models (LLMs) reason over optimization histories stored primarily in conversational context, they must recover candidate identities, prior

    multi-agentbenchmark
  394. arxiv:2609.34488 · cs.LG
    FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL
    Xun Wang, Ruishuo Chen, Yu Chen, Zhuoran Li +1

    Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to natur

    post-training
  395. arxiv:2609.34486 · cs.RO
    Model-Informed Safe Reinforcement Learning for Bipedal Locomotion via Step-to-Step Prediction
    Victor Paredes, Ayonga Hereid

    Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches.

    humanoidwhole-body control
  396. arxiv:2609.34484 · cs.RO
    ARS: Agentic Reward System for Robot Learning
    Sheng Hu, Weiyi Lu, Lingbing Zeng, Gan Weng +4

    Progress reward modeling is the problem of estimating how a robot's behavior changes task progress over time. Reliable estimation requires distinguishing meaningful state changes from failed attempts and task-irrelevant actions. We introduce the Agentic Reward System (ARS), an inference framework fo

    manipulationagentagenticbenchmark
  397. arxiv:2609.34481 · cs.CL
    CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR
    Bashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh +35

    We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed

    benchmark
  398. arxiv:2609.34480 · cs.CV
    When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes
    Sungguk Cha, Mintae Kim, Youngsub Han, Byoung-Ki Jeon +1

    Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each ques

    benchmark
  399. arxiv:2609.34470 · cs.CV
    Precise Editing and Flexible Referencing for Interactable Worlds
    Xinyao Liao, Xianfang Zeng, Zhu Liang, Zhoujie Fu +4

    We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends

    world model
  400. arxiv:2609.34469 · cs.CL
    The Last Mile Is the File: OfficeEditBench for Preservation-Aware Office Editing
    Zhiwen Wu, Chengxu Wu

    A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, a 170-task benchmark for change-scoped m

    benchmark
  401. arxiv:2609.34467 · cs.LG
    Alignment-Guided Flow Transformer for Efficient Vision-Language-Action Policy Learning
    Shengchao Hu, Peng Wang, Qiyang Zhou, Guodong Zheng +4

    Recent advances in Vision-Language-Action (VLA) models point toward general-purpose robotic intelligence by unifying perception, instruction, and control. Despite impressive progress, existing VLA models often adapt poorly due to \emph{tri-modal misalignment} among vision, language, and action, whic

    vision-language-actionvlavla modelmanipulationbenchmark
  402. arxiv:2609.34463 · cs.AI
    CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents
    Xiao Yang, Yangchen Ou, Yuhan Gao, Le Wang +2

    Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However, existing training-based def

    agentbenchmark
  403. arxiv:2609.34460 · cs.LG
    When Does Structured Knowledge Help Neural Theorem Proving?
    Sareh Nabi, Roland Vogl, Marzieh Nabi

    Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analog

    knowledge graph
  404. arxiv:2609.34459 · cs.AI
    Escaping Local Views: Discovering Latent Concepts for Interpretable Multi-Agent Reinforcement Learning
    Yijie Sun, Sanquan Sun, Yanda Zhu, Yuanyang Zhu +2

    Efficient cooperation is challenging due to the usual partial observability of each agent in multi-agent reinforcement learning. Recurrent networks encode local interaction histories, but their hidden representations provide limited insight into the information underlying individual decisions. To ad

    agentmulti-agent
  405. arxiv:2609.34457 · cs.LG
    ZonoGPT: Towards An Abstract Domain for Verifying Large GPT Models
    Hai Duong, Thanh Le, ThanhVu Nguyen

    Transformer-based models are widely used for reasoning, coding, and multimodal agentic tasks. To provide formal assurance of desirable behaviors, such as robustness, safety, and fairness, neural network verification techniques prove required properties and provide auditable guarantees before deploym

    agentic
  406. arxiv:2609.34455 · cs.CL
    RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications
    Jianpeng Zhao, Haihua Xu, Haoyang Zhang, Shuang Qian +6

    We study reasoning in Rule-Governed Decision Tasks (RGDTs), where models apply external rules to case facts and justify decisions, as required in policy, contract, and compliance settings. Beyond the deductive capability emphasized by standard mathematical and logical reasoning tasks, RGDTs require

    benchmarkevaluator
  407. arxiv:2609.34453 · cs.LG
    SPACE-LoRA: Allocating Activation-Subspace Protection for Continual Learning
    Seunghyun Yoo, Kiseok Kim, Hyeontae Joo, Junyeop Bang +1

    This study addresses the catastrophic forgetting problem that occurs when sequentially learning successive tasks using Low-Rank Adaptation (LoRA) from a lifelong learning perspective. While existing approaches have primarily constrained parameter updates or learning subspaces to reduce interference

    lifelong learning
  408. arxiv:2609.34447 · cs.LG
    Unbiased Top-$k$ Estimation for On-Policy Distillation
    Linjian Meng, Siyuan Gan, YuHan Li, Xiran Wang +4

    On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student vi

    post-training
  409. arxiv:2609.34446 · cs.LG
    Livin' on a Prior: Likelihood Score Approximation for Inverse Problems
    Rostislav Makarov, Tal Peer, Danilo de Oliveira, Timo Gerkmann

    Generative models have found great success as data-driven methods of solving inverse problems. Two popular approaches work either by combining a pretrained generative prior with a known degradation model, or by training a conditional generative model directly from paired data. We target a setting th

    post-trainingbenchmark
  410. arxiv:2609.34444 · cs.AI
    Social Circuits behind Multi-agent Echo Chambers
    Chuiyang Meng, Wenlu Yu, Ming Tang, Cheng Li

    Language-model agents exchange messages to combine evidence, but their communication can also create echo chambers that reinforce shared errors. However, overall task performance does not explain how a message changes the receiving agent's internal activations and affects its decision. In this work,

    multi-agent
  411. arxiv:2609.34442 · cs.LG
    Making LLMs Truly Forget: Deep Unlearning by Searching, Selecting, and Severing Knowledge Paths
    Jialu Wang, Peizhi Niu, Haoteng Yin, Hans Hao-Hsun Hsu +2

    While an unlearned language model may no longer recall a fact directly, the fact often remains recoverable through multi-hop reasoning over related knowledge. Most existing unlearning techniques overlook this vulnerability, targeting facts in isolation while leaving their supporting knowledge intact

    knowledge graph
  412. arxiv:2609.34440 · cs.RO
    When the Score Becomes the Target: Rethinking Metric Validity in Autonomous Driving
    Morui Zhu, Deyuan Qu, Qi Chen, Kentaro Oguchi +1

    Driving benchmark scores are increasingly used not only for evaluation but also as optimization targets. This raises a fundamental question: do score gains remain reliable evidence of driving improvement once the score itself is optimized? We address this question by examining how the scoring proces

    benchmark
  413. arxiv:2609.34438 · cs.AI
    Remember by Asking: Retrieval-Induced Memory Evolution for LLM Agents
    Wanqi Zhou, Jiawei Lu, Yang Wang, Zhaolong Xing +3

    Long-term memory is essential for language agents to maintain coherent and effective behavior over extended, multi-session interactions. Existing memory systems mainly use retrieval at read time, while write-time memory formation still relies on direct extraction or compression. However, when future

    memoryllm agent
  414. arxiv:2609.34428 · cs.AI
    AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
    Chanhee Park, Jeongho Yoon, Sungbin Han, Hyeonseok Moon +1

    Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagn

    agentagenticagent benchmarkbenchmarkleaderboard
  415. arxiv:2609.34427 · cs.LG
    LLMs as Adaptive Meta-Solvers: Strategy-Diverse RL for Industrial-Scale Optimization
    Shihao Zhang, Weiting Liu, Siyu Shao, Yitian Chen +3

    Scaling LLM-based optimization from textbook-scale instances to real-world, industrial tasks remains a critical open challenge. Existing approaches are predominantly evaluated on small, self-contained textual problems and often commit to a solver-integrated paradigm, limiting their ability to handle

    benchmark
  416. arxiv:2609.34426 · cs.LG
    Q-learning Penalized Transformer for Safe Offline Reinforcement Learning
    Shengchao Hu, Peng Wang, Jifeng Hu, Qiyang Zhou +4

    This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety co

    benchmark
  417. arxiv:2609.34425 · cs.AI
    Zero-Shot Cue-Grounded Topic Segmentation of Spoken Documents
    Suhwan Choi, Myeongho Jeon, Myungjoo Kang

    Topic segmentation structures spoken documents into coherent sections, facilitating navigation and downstream understanding. The appropriate granularity can vary substantially, ranging from broad thematic shifts to fine-grained subtopics. Existing LLM-based segmenters, however, often struggle to ada

    benchmark
  418. arxiv:2609.34423 · cs.LG
    On the Relation Between Interval Regret and Dynamic Regret
    Yi-Han Wang, Peng Zhao, Zhi-Hua Zhou

    Non-stationary online learning has attracted much attention in recent years, as static regret is insufficient to guide algorithm design in changing environments. To address this limitation, interval regret and dynamic regret have been introduced as two representative performance metrics that strengt

    online learning
  419. arxiv:2609.34422 · cs.LG
    Coding Agent Memory Post-training: Unlocking the Memory Potential of Pre-trained File Operations for Long-Horizon Tasks via Reinforcement Learning
    Lirui Luo, Kelong Mao, Heming Xia, Rongqing Li +9

    Language-model agents increasingly tackle long-horizon tasks whose interaction histories exceed the model's active context. Recent work has begun to use reinforcement learning to make memory control part of the policy, often relying on predefined memory tools within domain-specific training environm

    memoryagent memoryagentagenticpost-training
  420. arxiv:2609.34420 · cs.LG
    MegaGraph: Towards Efficient Training of Large-Scale Graph Transformers with Automated Hybrid Parallelism
    Tong Qiao, Ao Zhou, Yingjie Qi, Chunming Hu +1

    Graph Transformers (GTs) offer superior representation capabilities by overcoming the depth limitations and over-smoothing issues of traditional Graph Neural Networks (GNNs). However, scaling GTs to large graphs poses critical bottlenecks. Specifically, the attention score matrix and its associated

    memory
  421. arxiv:2609.34419 · cs.AI
    Beyond End-to-End Black Box Mapping: An Intentional Agent Framework for Cognitive-driven Facial Reaction Generation
    Hanzhong Zhang, Jindong Wang, Siyang Song

    Automatic human-like facial reaction generation (FRG) is essential for building intelligent systems that can engage in human-computer interaction (HCI). While diverse and context-appropriate facial reactions can reflect latent appraisal and affective processes in human interaction, most existing FRG

    agentagent framework
  422. arxiv:2609.34415 · cs.LG
    PROACT-Agent: Progressive Runtime Oversight and Active Circuit-breaking for Real-Time Safety
    Ding Jia, Wei Liu, Xianglong Du, Yingjie Li +4

    The transition from Large Language Models (LLMs) to agents shifts safety stakes from toxic text to irreversible environmental harm. While current defenses remain largely retrospective, proactive runtime intervention is bottlenecked by the lack of large-scale, causally-consistent data. We propose PRO

    long-contextbenchmark
  423. arxiv:2609.34414 · cs.RO
    From World Models to World Action Models: Rethinking Next-State Prediction
    Tingyu Yuan, Ziming Ji, Biaoliang Guan, Wen Ye +8

    Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of

    liberoworld modelaction-conditioned
  424. arxiv:2609.34412 · cs.RO
    From Language to Task Maps: Compiling Semantic Relations While Preserving Task-Relevant Freedom
    Jaegyun Park, Jingwang Lee, Jungsoo Lee, Soonwoong Hwang +1

    Natural-language manipulation instructions specify qualitative relations, whereas continuous controllers require state-evaluable task quantities, differentials, and completion conditions. Because a qualitative relation generally leaves part of the relative configuration unspecified, expanding it int

    manipulationfranka
  425. arxiv:2609.34408 · cs.CV
    Distilling Visual Reasoning into Text Space
    Wenhan Yang, Nilay Naharas, Ali Payani, Baharan Mirzasoleiman

    Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these represe

    benchmark
  426. arxiv:2609.34406 · cs.LG
    Unlocking Few-Step Diffusion for Faithful Previews
    Jing Jia, Sifan Liu, Guanyang Wang

    Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can cl

    benchmark
  427. arxiv:2609.34404 · cs.AI
    Eval4DiRec: A Unified and Systematic Evaluation Framework for Diffusion-based Recommender Systems
    Cong Wang, Shoujin Wang, Yishuo Li, Qi Zhang +2

    Leveraging the strong generative capabilities and stable training dynamics of diffusion models, diffusion-based recommender systems (RSs) have recently emerged as a novel recommendation paradigm, attracting increasing attention from both academia and industry. However, despite the rapid growth of di

    benchmarkevaluation frameworkevaluation protocol
  428. arxiv:2609.34398 · cs.LG
    GeoCFM: Positive-Only Conditional Flow Matching for Mineral Occurrence Sampling
    Moshe Eliasof, Eldad Haber

    Critical mineral discovery is a positive-only problem: deposits are observed as sparse locations, while unlabeled regions are not reliable negatives, and similar geophysical signatures can arise from different subsurface states. We therefore model mineral targeting as learning a conditional spatial

    benchmark
  429. arxiv:2609.34397 · cs.AI
    SkillFocus: Evolving Agent Skills via Capability Decomposition
    Ning Wang, Zhiren Gong, Bingdong Li, Peng Yang +1

    Agent skill evolution seeks to improve reusable procedural guidance for large language model (LLM) agents through iterative revision. Existing methods base each revision mainly on execution trajectories or feedback, leaving recurring behavioral requirements across tasks implicit and tying revision t

    agentbenchmark
  430. arxiv:2609.34396 · cs.CV
    HyperDAM: Hyperspectral Distractor-Aware Memory with Amodal Expansion for SAM 3 Tracking
    Ryoga Yuzawa, Tasuku Takagi

    Hyperspectral video provides material cues that can disambiguate targets with similar false-color appearance, yet foundation-model trackers update memory primarily from spatial and appearance evidence. We present HyperDAM, a DAM4SAM3-based hyperspectral tracker with three principal contributions. Fi

    memoryleaderboard
  431. arxiv:2609.34392 · cs.AI
    Org-Agent: Beyond Personal Assistants Towards Organizational Agents
    Luyao Zhuang, Yujing Zhang, Zijin Hong, Yilin Xiao +1

    Language model agents serving organizations must coordinate requests from multiple users while using knowledge distributed across their interactions. We identify two complementary capabilities for this setting, namely cross-user interaction and decision-making, as well as cross-user memory and knowl

    memorytool use
  432. arxiv:2609.34391 · cs.LG
    P2P: Cross-View Population Denoising for Unpaired Single-Cell Perturbation Response Prediction
    Haojie Yang, Ran Su

    AIVC (AI Virtual Cell) is a learned simulator of cellular behavior across conditions. Predicting how a cell population responds transcriptionally to a genetic perturbation is a core task. Perturb-seq records that response by destructive sequencing, so a control cell and a perturbed cell are never ob

    memory
  433. arxiv:2609.34390 · cs.CV
    VastMAT: A Large-Scale Multi-Category Benchmark for Multi-Animal Tracking
    Zhizhen Li, Zan Wang, Huidong Peng, Bohan Tan +3

    Multi-animal tracking (MAT) supports the study of animal movement, behavior, and group interactions. However, general multi-object tracking (MOT) benchmarks primarily focus on pedestrians and vehicles, whereas dedicated MAT benchmarks remain limited in jointly supporting broad animal coverage, large

    benchmark
  434. arxiv:2609.34387 · cs.CV
    CAR-VLA: Complexity-Aware and Risk-Adaptive Reasoning for Autonomous Driving
    Xiaolei Chen, Zhuolin He, Yuxuan Liang, Xu Li +12

    Existing adaptive reasoning methods for driving Vision-Language-Action (VLA) models primarily focus on whether to reason, overlooking how reasoning should differ across driving situations. Our key insight is that while scene complexity informs reasoning depth, dynamic risk is equally critical for de

    vision-language-actionvlavla model
  435. arxiv:2609.34385 · cs.LG
    Just-In-Time Agent Memory with Runtime Agentic Research
    Bingyu Yan, Chaofan Li, Hongjin Qian, Shuqi Lu +2

    Memory is critical for AI agents. Many existing agent-memory systems follow an Ahead-of-Time (AOT) design, constructing memory before a specific request arrives. While this reduces online serving cost, such request-agnostic memory construction can discard fine-grained information that later becomes

    memorylong-contextagent memoryagentai agentagentic
  436. arxiv:2609.34384 · cs.RO
    RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models
    Aernaer Akelijiang, Jiannan Li, Zhineng Chen, Jingjing Chen +1

    Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicit

    vision-language-actionmanipulationfrankamemorybenchmark
  437. arxiv:2609.34380 · cs.AI
    DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory
    Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung +1

    Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this design by allowing a single stored model to support both full-accuracy and lower-precision execution

    memory
  438. arxiv:2609.34379 · eess.SY
    Derivative-Free Generalized Multivariable Super-Twisting Control for Constrained Euler-Lagrange Systems
    Chidre Shravista Kashyap, Jishnu Keshavan

    Robotic manipulators performing payload lift-and-transfer must track prescribed trajectories under abrupt load changes, with limited actuation and inaccurate plant knowledge. Such plants are inherently described by multivariable Euler--Lagrange~(EL) dynamics with cross-coupling established through i

    manipulator
  439. arxiv:2609.34378 · cs.CV
    Marathoner: Ultra-Long-Horizon Autonomous Intelligence
    Zhang Ruiyang, Ou Jinpeng, Xie Yifan, Zhou Jingang +3

    Humans naturally possess the ability to work persistently toward long-term goals. Given a challenging task, humans can continuously work for months or even years to accomplish a specific objective. In this paper, we propose Marathoner, an autonomous agentic model possessing the ability of ultra-long

    autonomous agentagenticpost-trainingbenchmark
  440. arxiv:2609.34375 · cs.RO
    LRC-JEPA: Disentangling Dynamics and Residual Context for Efficient World Models
    Luzhe Huang, Lei Chu, Jingyi Liang, Yuhuan Zhao

    Compact JEPA world models enable efficient latent-space planning, but low-dimensional representation trained under reward-free self-supervision must encode both action-conditioned dynamics and predictable visual context. This competition can entangle controllable state with high-rank nuisance appear

    world modelv-jepaaction-conditioned
  441. arxiv:2609.34374 · cs.LG
    ABC-Align: Prediction-Powered Alignment with Adaptive Bias Control
    Eric Frankel, Banghua Zhu, Sewoong Oh, Lillian J. Ratliff

    Language model post-training is often bottlenecked by the need for human-collected preference data, which is expensive and difficult to scale. Reinforcement learning from AI feedback (RLAIF) style approaches that leverage pseudo labels offer an abundant alternative but introduce systematic biases th

    post-training
  442. arxiv:2609.34373 · cs.RO
    Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication
    Mihir Chauhan, Aniket Bera

    Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited

    multi-agentself-playarena
  443. arxiv:2609.34372 · cs.AI
    PersMem: Internalizing Personality into Dual-Pathway Memory for LLM Agents
    Hanzhong Zhang, Ziwei Xiang, Weicheng Xie, Shizhe Liu +1

    The profile of a role-playing agent usually depends on the pre-defined personality in a system prompt, whereas its memory processing pipeline, including prioritisation of stored memories and subsequent retrieval, remains independent of this personality. This separation causes the agent's memory proc

    memoryagentllm agent
  444. arxiv:2609.34370 · cs.LG
    Spexis: Speculative Lookahead Scheduling for LLM Inference
    Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim +2

    Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new para

    memory
  445. arxiv:2609.34367 · cs.CV
    Rate-Distortion Adaptive Primitive Selection for Omnidirectional Gaussian Splatting
    Yulong Cheng, Youneng Bao, Junfeng Zhou, Mu Li +1

    Learned image codecs (LICs) achieve high reconstruction quality, but their decoding speed is often insufficient for immersive virtual reality (VR). Gaussian splatting (GS) codecs render much faster, yet still lag in reconstruction quality and typically decide primitive allocation without considering

    benchmark
  446. arxiv:2609.34366 · cs.AI
    When Harness Beats Scale, and When Reading Beats Both
    Ivan Bondarenko, Nikolay O. Nikitin

    We describe our system for DocSem, the document-grounded quantitative reasoning shared task at DocInsights 2026, and analyze why it succeeded on labeled data and failed on the test set. The pipeline pairs hybrid block retrieval with Program-of-Thoughts (PoT) generation executed in a sandboxed interp

    knowledge graphleaderboard
  447. arxiv:2609.34363 · cs.CV
    SyncRA: Learning Temporal Correspondence in Omni-Modal Models
    Zelong Xu, Yan Li, Wenhe Hu, Xiyang Hu

    Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible

    benchmark
  448. arxiv:2609.34362 · cs.RO
    FutureDuet: Decoupling Observation Access from Future Supervision in World Action Models
    Jie Wu, Yuzhi Huang, Junqi Liu, Weichen Zhang +4

    World Action Models (WAMs) augment robot action generation with future visual supervision. Existing WAMs commonly fuse main and wrist observations into one visual stream and train both with the same future-video objective, despite their different visual dynamics. A stable main camera reveals scene-l

    liberorobotwin
  449. arxiv:2609.34359 · cs.AI
    Improving Large Language Models for Code through Runtime Program-State Reasoning
    Hongwei Li, Spandan Garg, Yufan Huang

    Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-

    agentpost-training
  450. arxiv:2609.34358 · cs.LG
    FORGE: Form-Optimal Routing of Grounded Evidence for Frozen LLM Agents
    Xi Xiao, Yunbei Zhang, Chen Liu, Lin Zhao +6

    In agentic AI systems, frozen foundation models are increasingly deployed as closed-weight API endpoints, making downstream adaptation possible only through the inputs and inference procedures surrounding the model. As a result, for each input query, two coupled decisions largely determine both answ

    llm agentagenticbenchmark
  451. arxiv:2609.34353 · cs.AI
    SemRD-V2X: Closure-Guided Communication with Bounded Inference for Cooperative Perception
    Hu Xu, Chun Li, Siyuan Qiu, Zeyan Li +1

    Vehicle-to-Everything (V2X) cooperative perception improves 3-D detection by sharing intermediate features, but dense remote features may repeat context that the ego agent can infer locally. Most communication-efficient designs optimize masks or codes empirically, leaving a more basic question open:

    agent
  452. arxiv:2609.34349 · cs.AI
    Fuzzy Distribution Modeling for Synthetic Tabular Data Generation with Causality Preservation
    Michael Vasilakakis, Dimitris K. Iakovidis

    Synthetic tabular data generation provides an effective alternative for the training of machine learning models when real-world data is limited or inaccessible. However, the heterogeneous, non-smooth, and incomplete nature of tabular data poses fundamental challenges to conventional probabilistic an

    benchmark
  453. arxiv:2609.34348 · cs.LG
    Commutator Memory: Sparse, Path-Local Reading and Steering in Language Models
    John Sweeney

    Gradient updates on different data generally do not commute: training a language model on two data sources in opposite orders gives different weights, even with the same data and total exposure. Loss or benchmark deltas show that the models differ, not where. We ask whether this path dependence leav

    memorybenchmark
  454. arxiv:2609.34347 · cs.AI
    SAIL: Spatial Audio Intelligence with Large Language Models via Disentangled Acoustic-Spatial Encoding and Dual-Stream Q-Former
    Zhengding Luo, Jinyang Wu, Haozhe Ma, Yanghao Zhou +2

    Spatial audio large language models (LLMs) enable embodied agents, wearable assistants, and immersive systems to recognize sound events, localize sources, and reason about their spatial relationships. However, existing spatial audio LLMs often rely on early fusion of acoustic and spatial features an

    embodiedembodied agent
  455. arxiv:2609.34346 · cs.CV
    E-WAVE: Event-based Continuous Optical Flow via Warping-Aligned Visual Encoding
    Jiale Wu, Xiaoyang Bai, Haoming Yu, Yiwei Chen +2

    Temporally dense optical flow is essential for dynamic perception in immersive VR/AR systems, where rapid head, hand, and object motion must be continuously captured and tracked. Existing frame-based optical flow estimation methods are constrained by the tradeoff between temporal resolution and comp

    memoryevent camera
  456. arxiv:2609.34345 · cs.CL
    CRISP: Cultural Reward Modeling for Implicit Situated Propriety
    Zekun Yuan, Yangfan Ye, Baohang Li, Shuaibo Zhao +5

    As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined respon

    multi-agentagent framework
  457. arxiv:2609.34344 · cs.LG
    Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
    Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu +9

    Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify

    post-training
  458. arxiv:2609.34342 · cs.AI
    SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing
    Zhiwei Chen, Tianchun Wang, Zhongtao Rao, Haiming Zhu +2

    A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents' behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupporte

    agentllm agent
  459. arxiv:2609.34335 · cs.CV
    SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering
    Yanwei Huang, Mingxuan Zhu, Shujie Li, Shiyuan Liu +2

    Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. S

    benchmark
  460. arxiv:2609.34330 · cs.CV
    MiCo: Mutual Information Coverage Optimization through Semantic Erasure Modeling for Efficient MLLM Inference
    Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang +10

    Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visual tokens results in high computational costs. While many methods have been proposed to reduce the number of visual tokens, most of them rely on heuristi

    benchmark
  461. arxiv:2609.34327 · cs.AI
    Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
    Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo +2

    Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model f

    self-refinementbenchmark
  462. arxiv:2609.34326 · cs.LG
    Routing Without Embeddings: Fast And Interpretable Routing With Regular Expressions
    Yifan Lu, Qiyue Zhang, Haotian Shan, Hanjie Chen +1

    Large Language Model (LLM) routers commonly rely on neural query embeddings, with larger encoders expected to better capture query intent and difficulty. Yet scaling Qwen2.5 encoders from 0.5B to 72B parameters brings little improvement in routing accuracy (Figure 1b), suggesting that small encoders

    benchmark
  463. arxiv:2609.34325 · cs.CV
    DORA: Dynamic Online Reinforcement Agent for Token Pruning in Vision Transformers
    Kaixuan He, Song Chen, Yi Kang

    Vision Transformers (ViTs) incur quadratic self-attention cost in the number of tokens. Most token-reduction methods adapt token identities within a prescribed layer-wise compression schedule, or search a static mask offline, and thus limit online adaptation of when and how much to prune. We propose

    agent
  464. arxiv:2609.34322 · cs.AI
    Test-Time Scaling via Budgeted Multi-Attribute Verification
    Bo Xue, Ji Cheng, Shen-Huan Lyu, Yuanyu Wan +1

    Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated alon

    benchmark
  465. arxiv:2609.34321 · cs.LG
    One Rollout Is All You Get: Fully Test-Time Adaptation for GUI Agents
    Ziqiang Wang, Li Gu, Zhixiang Chi, Linlian Jiang +5

    GUI agents are deployed with frozen weights and discard everything they experience on the job. Existing ways to update an agent's weights assume something deployment withholds: ground truth, rollouts beyond the single attempt (retries, samples, practice runs), or a learning phase other than deployme

    memoryagent
  466. arxiv:2609.34320 · cs.CL
    Certified Selective Automation of LLM Agent Evaluation
    Chengguang Gan, Yunhao Liang, Qinghao Zhang, Shiwen Ni

    Evaluating LLM agents still ends with a human reading trajectories, because automatic judges carry no guarantee on how often they are wrong. We ask the operational question: what fraction of agent evaluation can a judge take over, with a certificate that the error rate among auto-decided trajectorie

    agentllm agenttool-use
  467. arxiv:2609.34319 · cs.RO
    Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference
    Qianer Li, Chengjie Zhang, Jingwen Chen, Zanjia Tong +2

    Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy,

    vision-language-actionvlavla modelopenvlabenchmark
  468. arxiv:2609.34314 · cs.CV
    PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
    Shayekh Bin Islam, Hwanjun Song

    Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Th

    agenticbenchmarkjudge model
  469. arxiv:2609.34313 · cs.AI
    ControlScope: Workflow Revision and Reliability in LLM Agents
    Jingjie Ning, Xueqi Li, Yibo Kong, Dongting Li

    How much of a running workflow should a language model agent revise? ControlScope compares continuing generated code, editing the next tool call's data arguments, and replacing the unfinished workflow from the same public execution state. The nested permissions separate available repairs from the ac

    agentllm agent
  470. arxiv:2609.34309 · cs.CV
    MaLiang-Harness: A Programmable Path to Image and Video Generation
    Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu +3

    Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to

    benchmark
  471. arxiv:2609.34300 · cs.RO
    When World Models Lie: Adaptive Safety Analysis Under Wrong Imaginations
    John Cao, Somil Bansal

    World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safe

    world model
  472. arxiv:2609.34299 · cs.CV
    PSM: Dataset Distillation Based on Precise Statistical Matching by Difficulty
    Hongxu Ma, Guang Li, Shijie Wang, Dongzhan Zhou +5

    Dataset distillation (DD) condenses a large original dataset into a small distilled dataset with high training utility. Decoupled statistical matching methods substantially reduce distillation time and memory overhead while achieving strong performance. However, they typically supervise all distille

    memory
  473. arxiv:2609.34297 · cs.RO
    TLC-DiT: Task-Aligned Local Visual Conditioning for Robust Multitask Robot Manipulation
    Xianbo Cai, Hideyuki Ichiwara, Zihang Wang, Yijun Lu +1

    Language-conditioned robot policies have made clear progress in multitask manipulation, but task-relevant local visual evidence usually stays hidden inside a visual backbone or attention layers. This leaves the policy difficult to inspect and fragile under visual change, two symptoms of a missing ex

    manipulationlibero
  474. arxiv:2609.34296 · cs.AI
    Dr.Credit: Rubric-Grounded Process Credit Assignment for Deep Research Agents
    Yingjian Zhu, Zhenyi Wang, Jiaxin Guo, Kun Ding +5

    Rubric-based tasks are increasingly addressed through reinforcement learning (RL), with rubric scores used as training rewards. However, these rewards typically supervise final answers without distinguishing the contributions of intermediate decisions. Many existing credit assignment methods rely on

    agentbenchmark
  475. arxiv:2609.34294 · cs.CV
    Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings
    Duanning Chen, Ke He, Bin Yang, Yongxiang Yao

    Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without

    memory
  476. arxiv:2609.34287 · cs.LG
    ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining
    Zichun Yu, Jiarui Yan, Shlok Sanghvi, Nihar Atri +1

    LLM pretraining corpora are normally cleaned by a stack of hand-written heuristics. A heuristic scraper extracts the main content from HTML, and dozens of rule-based filters then clean it, so corpus quality is capped by the coarseness and accuracy of the rules. In this work, we propose ReScraper, a

    multi-agent
  477. arxiv:2609.34286 · cs.RO
    Dexterous Tactile World Model
    Ziyao Zeng, Xiatao Sun, Hao Wang, Yueyang Pan +5

    World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video wo

    manipulationdexteroustactileworld model
  478. arxiv:2609.34284 · cs.CL
    Over-Personalization Is a Decision Failure: Generation-Induced Apply Bias in LLMs
    Haeun Jang, Yonghyun Jun, Hwanhee Lee

    Personalized LLMs must decide, for each stored preference, whether the current context calls for applying or suppressing it, which we call its applicability. They frequently over-personalize, applying preferences the context rules out, yet existing benchmarks score only the final response and cannot

    benchmark
  479. arxiv:2609.34281 · cs.LG
    Agentic High-Dimensional Bayesian Optimization with Hypothesis- and Evidence-Guided Search
    Zhixuan Gao, Ke Xue, Rongxi Tan, Ming Chen +1

    High-dimensional Bayesian optimization (HDBO) seeks sample-efficient optimization when the number of variables is large relative to the evaluation budget. Recent LLM-based and agentic BO methods incorporate task knowledge and adapt search decisions during a run, but have primarily been evaluated on

    agentagenticbenchmark
  480. arxiv:2609.34279 · cs.LG
    Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training
    Yuyang Deng, Yu Wang, Jiayun Wang

    Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}ire

    self-evolving
  481. arxiv:2609.34277 · cs.CV
    See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology
    Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li +7

    Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that

    benchmark
  482. arxiv:2609.34276 · cs.RO
    NavHarness: Towards Lifelong Embodied Navigation
    Xunyi Zhao, Jian Zhou, Sihao Lin, Gengze Zhou +5

    Frontier models can now perform well on individual embodied navigation tasks through multi-round multimodal reasoning with simple tools. Across successive tasks, however, an agent must also rely on an evolving map and earlier search records, both of which may be incomplete or conflict with new obser

    embodiedmemoryagentagentic
  483. arxiv:2609.34274 · cs.AI
    BIABench: Evaluating AI agents on real-world bioimage analysis tasks
    Zixuan Pan, Davide Panzeri, Lukas Johanns, Marilin Moor +4

    Artificial-intelligence (AI) agents hold promise for automating bioimage analysis, yet no benchmark evaluates whether they can carry out real-world analyses end to end. Such analyses are hard for agents because 2D images, 3D volumes and time-lapse sequences are often too large to read as context, so

    agentai agentbenchmark
  484. arxiv:2609.34271 · cs.CV
    Scaling Versatile 3D Assets Editing with a Million-Scale Dataset
    Badi Li, Tianxin Huang, Yu Zhou, Wei-Shi Zheng +2

    Although recent 3D generative models produce increasingly realistic assets, controllable 3D asset editing remains challenging. Existing methods are limited by scarce training data, insufficient source-aware modeling, and a lack of practical evaluation protocols. To address these limitations, we pres

    benchmarkevaluation protocol
  485. arxiv:2609.34268 · cs.RO
    SAGE: Symbolic Action-Gating and Editing for LLM Task Planners
    Trung Minh Bui, JongSul Moon, YoungOuk Kim, Quang-Ngoc Phung +2

    Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be s

    embodiedmemorybenchmark
  486. arxiv:2609.34262 · cs.LG
    Maintaining Benchmarks Against Increasingly Capable Agents: Detection and Remediation of Unearned Passes
    Weijun Luo, Kelvin Luu, Xinyi Liu, Guangze Luo +5

    Agentic benchmarks guide model selection and training. Yet an agent can pass a task without demonstrating the intended capability. Such outcomes constitute unearned passes; their proportion among all passes defines the integrity gap. As agents improve, benchmark surfaces that once seemed harmless ca

    agentagenticbenchmark
  487. arxiv:2609.34261 · cs.RO
    RoboICL: Embodied In-Context Learning with GPT-6 Astra
    Fangcheng Liu, Yeqing Shen, Anda Cheng, Weishi Mi +6

    General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-contro

    embodiedmanipulationmemoryleaderboard
  488. arxiv:2609.34258 · cs.AI
    Investigating Human--AI Discrepancies via Multiple-Solution Problems
    Zihao Wang, Francesco Insulla, Andrea Montanari

    Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and m

    benchmark
  489. arxiv:2609.34256 · cs.RO
    UMR: Universal Manipulation Representation
    Song Liu, Linyi Li, Yanshun Zhao, Rxuan Li +10

    General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new

    vlaembodiedmanipulationliberobenchmark
  490. arxiv:2609.34250 · cs.RO
    WAM-OPD: Sharpening World Action Models via On-Policy Distillation
    Panjun Liu, Xiaohan Lei, Shiqi Zhang, Yikun Wang +7

    Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. W

    manipulation
  491. arxiv:2609.34246 · cs.LG
    CasEm: A Cascade Architecture for Long-Horizon Neural Emulation
    Zhaoyi Li, Jingtao Ding, Shihua Li

    Autoregressive neural emulators can drift or diverge over long rollouts despite accurate short-term predictions. We introduce Cascaded Emulation (CasEm), a one-way rollout architecture that augments an existing full-state backbone with an independently evolving model of physically specified aggregat

    benchmark
  492. arxiv:2609.34247 · cs.CL
    SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents
    Wenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen +4

    Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conve

    agentllm agent
  493. arxiv:2609.34245 · cs.LG
    Certified Multi-Source Integrity for Structured Agent Actions
    Anmol Pandey, Aditya Jain, Liang Chen, Carsten Maple +1

    LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself to extract attacker-c

    agentllm agent
  494. arxiv:2609.34242 · cs.LG
    Stashbird: Efficient Speaker-Indexed Memory for Conversational Agents
    Chidera Biringa, Lucas Yannul, Xiaowen Wang, Marco Ayala +4

    AI agents require memory that preserves information across user-agent exchanges, user-to-user conversations, and group conversations with or without agent participation, while supporting updates as evidence changes or is removed. We present Stashbird, an agent memory system that links source episode

    memoryagent memoryagentai agentbenchmark
  495. arxiv:2609.34241 · cs.AI
    AdaGuard: An Adaptive Guard Model with User-defined Policies
    Yunhao Feng, Yifan Ding, Yuxiang Xie, Zheng Li +3

    Guard models support the safe deployment of language model agents, but fixed risk taxonomies limit their ability to accommodate requirements that vary across applications and tasks. Under user-defined policies, detecting violations requires interpreting both the applicable rules and the agent's beha

    agent
  496. arxiv:2609.34237 · cs.CV
    DecFlowEdit: Self-Localized Flow-based Image Editing via Guidance Decoupling
    Zheyuan Zhan, Can Wang, Jiawei Chen, Chun Chen +3

    Flow-based image editing (FlowEdit) enables inversion-free semantic changes through the difference between source and target velocities. In this paper, we observe that FlowEdit's default classifier-free guidance (CFG) configuration, with asymmetric source and target scales, causes substantial backgr

    manipulation
  497. arxiv:2609.34235 · cs.CV
    SegBanana: Steering Unified Multimodal Models into Medical Segmenters
    Xiaoye Liang, Ye Yan, Mingze Yin, Shikun Feng +4

    Medical image segmentation remains challenging in practical deployment, as models often struggle to generalize beyond the distributions covered by their training data and high-quality pixel-level annotations are typically unavailable for adaptation. Inspired by the cross-task transferability of larg

    agenticpost-training
  498. arxiv:2609.34234 · cs.CL
    MAS-OPD: On-Policy Distillation for Multi-agent Systems
    Qiyong Zhong, Mao Zheng, Mingyang Song, Houcheng Jiang +4

    Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competenc

    agentmulti-agentagent systempost-trainingbenchmark
  499. arxiv:2609.34233 · cs.RO
    GAE: General Action Expert for Real-Time Humanoid Teleoperation
    Yuefan Wang, Huaicheng Zhou, Xiao He, Zhijie He +5

    Humanoid avatars extend human physical presence beyond the body, enabling people to participate in social, service, and labor activities through remotely operated robots. This requires teleoperation systems capable of realizing diverse and dynamic whole-body behaviors while maintaining responsive hu

    humanoidteleoperation
  500. arxiv:2609.34232 · cs.CV
    Trustworthy synthetic visual media: Evidence across the media lifecycle
    Zexi Jia, Zhiqiang Yuan, Jie Zhou, Jinchao Zhang

    Images and videos have long helped people understand what happened and how a work came into being. Generative systems complicate that role. Realistic media can now be produced and revised without leaving a stable history, so appearance no longer reveals whether a scene was captured, synthesized, or

    manipulation
  501. arxiv:2609.34231 · cs.CV
    ReGDiff: Guided Diffusion in Regulated Latent Space for Exploring Metamaterial Voxel Geometry
    Wangzhi Zhan, Jianpeng Chen, Dongqi Fu, Dawei Zhou

    Metamaterials are artificially engineered structures whose mechanical and physical behaviors are strongly shaped by geometry rather than composition. Voxel representation provides a unified format for metamaterial geometry generation, as it can express diverse classes such as truss, shell, and porou

    benchmark
  502. arxiv:2609.34228 · cs.LG
    SleuthBench: Benchmarking Statistical LLM Evaluation Using Tabular Hidden Signals
    Jingyun Jia, Antoine Remond-Tiedrez, Aaron Alvarez, Joshua Shunk +2

    Evaluating statistical discovery by large language model (LLM) agents requires verifiable analytical ground truth. Establishing such ground truth for real-world datasets is costly, and prior knowledge of public datasets can influence agent responses. We introduce SLEUTHBENCH, a benchmark that addres

    agentbenchmark
  503. arxiv:2609.34227 · cs.AI
    When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
    Rishabh Sharma, Rishika Lall

    Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran

    memoryagent memoryagent
  504. arxiv:2609.34225 · cs.CL
    USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents
    Qiyong Zhong, Mao Zheng, Mingyang Song, Huwei Ji +5

    On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them,

    agentic
  505. arxiv:2609.34223 · cs.CV
    Uncovering Ordinal-Matching Bias in Audio-Visual LLMs
    Jihoo Jung, Youngjoon Jang, Hyebin Cho, Suho Yoo +1

    This work aims to improve how audio-visual large language models (AVLLMs) associate speech with the correct visible speaker in multi-speaker scenes. We find that current AVLLMs frequently fail at this task, and analyze the nature of these failures. To this end, we construct a synthetic diagnostic da

    benchmark
  506. arxiv:2609.34222 · cs.RO
    Proprioceptive Force Estimation for Quadruped Locomotion and Human-Robot Interaction
    Run Wang, Xu Yang, Alapati Tuerxun, Yilin Mo

    Payload forces must be accommodated during locomotion, while leash forces can specify desired motion. We investigate whether a shared three-dimensional force estimate in newtons, inferred from proprioceptive history under sustained loading, can support both tasks. An estimator and locomotion policy

    quadruped
  507. arxiv:2609.34221 · cs.CV
    WorldWeave: Growing Persistent Geometric Worlds for Video Generation
    Yifan Huang, Lifan Jiang, Qingyue Hao, Cheng Chen +4

    Despite rapid progress, world models still lack explicit, persistent structural memory, making it difficult to preserve consistent world structure during continual scene expansion and cross-view revisits. To address this limitation, we present WorldWeave, a world generation framework that decouples

    world modelmemoryagent
  508. arxiv:2609.34220 · cs.RO
    mmHRI: Towards Privacy-Preserving Human-Robot Interaction with Millimeter-Wave Radar
    Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu +6

    Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on RGB cameras that continuously observe humans to respond to non-verbal commands, such as hand gestures

    vision-language-actionmanipulation
  509. arxiv:2609.34215 · cs.LG
    Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures
    Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang +4

    Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purel

    llm agentagenticbenchmark
  510. arxiv:2609.34214 · cs.AI
    GlyphBench: A Playground for Language-Model Reinforcement Learning
    Roger Creus Castanyer, Marc-Alexandre Côté, Matthew James Sargent, Augustine N. Mavor-Parker +2

    We introduce GlyphBench, an environment suite for reinforcement learning (RL) post-training of language-model agents, with over 360 tasks spanning diverse games. GlyphBench renders spatial observations as two-dimensional Unicode grids and connects training, evaluation, and trajectory replay through

    agentpost-training
  511. arxiv:2609.34211 · cs.AI
    Behavior-Grounded Semantic Enrichment for Financial Fraud Modeling and Reasoning
    Linbo Shao, Huilin He, Yating Lou, Dawei Cheng

    In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However, public real-world financial datasets often lack rich semantics due to privacy constraints. Consequently, synthetic datasets incorporate generated semanti

    multi-agent
  512. arxiv:2609.34210 · cs.RO
    RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
    Haitong Ma, Chenxiao Gao, Rushi Qiang, Na Li +1

    Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agent

    agentbenchmark
  513. arxiv:2609.34207 · cs.LG
    Functional Autoencoders for Amplitude-Phase Representation Learning
    Peida Wu, Xinyang Xiong, Pengcheng Zeng

    Functional data are intrinsically infinite-dimensional, and often exhibit phase variation, where corresponding events occur at different times across observations. Existing linear dimension reduction methods struggle with nonlinear amplitude variation, while functional autoencoders without an explic

    benchmark
  514. arxiv:2609.34206 · cs.CV
    WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies
    Lin Liu, Lu Zhang, Ziying Song, Wu Yang +5

    Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed i

    vision-language-actionvlaliberoworld model
  515. arxiv:2609.34205 · cs.LG
    Learning to Optimize through Solver-Grounded Self-Play
    Xia Jiang, Yaoxin Wu, Chenyu Zhou, Mengzhu Xu +2

    Optimization modeling is central to many decision-making scenarios, but traditionally requires extensive domain expertise. While Large Language Models (LLMs) have shown promise in automating this process, current training paradigms mainly rely on human-annotated or teacher-generated datasets. This d

    self-play
  516. arxiv:2609.34199 · cs.RO
    WB-WAM: Heterogeneous Body-Hand Pre-training for Humanoid Loco-Manipulation
    Chuan Qin, Shaoting Zhu, Siyuan Luo, Siqiao Huang +2

    Humanoid loco-manipulation demands coordinated body and hand behavior, while conventional robot pre-training data provide limited coverage of such whole-body motion. We present WB-WAM, a World Action Model that incorporates explicit whole-body action supervision into generative video pre-training. A

    manipulationdexteroushumanoid
  517. arxiv:2609.34198 · cs.LG
    Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
    Jiapeng Li

    Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 ex

    agent
  518. arxiv:2609.34196 · cs.CV
    ConvCue: Complementary Visual Inductive Biases for Vision-Language Models
    Zixuan Lan, Shichu Sun

    Modern vision-language models (VLMs) achieve strong performance across a broad range of multimodal tasks, yet still struggle with visual questions that require fine-grained discrimination and spatial understanding. These limitations motivate investigating whether supplementary visual representations

    benchmark
  519. arxiv:2609.34195 · cs.AI
    PainterBench: A Figural Divergent-Thinking Benchmark for Tool-Using Language Models
    Shane K. A. Dalumura Hettige, Jonas Oppenlaender

    Figural divergent thinking is the ability to develop a given shape fragment into an original drawing. In humans, this ability is assessed with incomplete-drawing tasks. We introduce PainterBench, a benchmark that ports the incomplete-drawing task to the agentic setting. The agent draws on a canvas t

    agentagentictool usebenchmark
  520. arxiv:2609.34190 · cs.CV
    MotionSpaceFlow: Representation-Aware Flow Matching in Direct Motion Space
    Qing Yu, Kent Fujiwara

    Recent advances in diffusion and flow models have substantially improved text-driven human motion generation. Yet most methods generate in low-dimensional, temporally downsampled latent spaces learned primarily for reconstruction, a bottleneck that can limit generation quality and preclude direct ma

    manipulation
  521. arxiv:2609.34188 · cs.LG
    AlphaPareto: Formulaic Alpha Discovery with LLM-Guided Multi-Objective Reinforcement Learning
    Yingbo Zhao, Zeyu Yang, Zhoufan Zhu

    Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha p

    agent
  522. arxiv:2609.34185 · cs.LG
    EntroPack: Fast and Accurate Entropy-Coded Weight Compression at Arbitrary Bitrates
    Hong Zhang, Zhongjie Duan, Yingda Chen

    Weight compression helps large neural networks fit deployment memory budgets, but common fixed-width formats offer only coarse storage choices. Entropy coding supports finer rates, yet the achieved size depends on the quantized weight distribution and coding overhead. Exploiting this flexibility req

    memory
  523. arxiv:2609.34184 · cs.AI
    CASS: Contribution-Aware Structured Sparsity for Model Merging
    Yan Li, Guiping Cao, Meng Xu, Tao Jiang +4

    Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, trea

    benchmark
  524. arxiv:2609.34182 · cs.RO
    Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation
    Wenqiao Li, Qianyou Zhao, Jiawen Hao, Xuezhou Zhu +4

    Dexterous manipulation requires tactile feedback.However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile intera

    manipulationdexterousteleoperationtactile
  525. arxiv:2609.34181 · cs.AI
    Efficient Reasoning via Constrained Optimization in Latent Space
    Zhinan Hou, XingChen Li, Keyou You

    Large Reasoning Models (LRMs) have shown remarkable reasoning capabilities, yet they still suffer from overthinking, generating redundant reasoning steps which incur substantial token consumption. Existing methods, such as suppressing reflective keywords or forcing shorter reasoning lengths, attempt

    benchmark
  526. arxiv:2609.34179 · cs.LG
    RAGWarrant: Evidence-Preserving Governance for RAG Policy Promotion Under Quality, Cost, Latency, and Risk Constraints
    Richard Krueger, Lucas Krause, Zach Pocquette

    Retrieval-augmented generation systems are extensively instrumented with metrics, benchmarks, traces, and automated judges, but these tools do not decide whether a proposed policy change is safe to release. We present RAGWarrant, an open-source promotion-control framework that treats deployment as a

    retrieval-augmentedragbenchmarkevaluatorleaderboard
  527. arxiv:2609.34178 · cs.CV
    Enhanced Video Text Editing with Trajectory-Aligned Glyph Rendering
    Shulian Zhang, Xiangyu Shu, Wenbo Li, Jian Chen +1

    Video text editing aims to replace or add text in a video while keeping the rest of the video unchanged, which requires the edited text to be correct in every frame and to move coherently with the scene. Despite the remarkable progress of video diffusion models, they struggle to reproduce exact stro

    benchmark
  528. arxiv:2609.34177 · cs.LG
    ReplayLens: Auditing Agents' Use of Outcomes
    Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang +4

    When an agent reuses logged experience, a changed decision may reflect the recorded score, the action's name, or the record's position in storage. Standard memory evaluations do not reveal which relationship drives that change. We introduce ReplayLens, a black-box audit that changes one relationship

    memoryagent
  529. arxiv:2609.34176 · cs.RO
    AGILE-GS: Anchor-Guided Fast Next-Best-View Selection for Active 3D Gaussian Splatting
    Amirhossein Mollaei Khass, Nader Motee

    Radiance fields need hundreds of views, and their placement matters as much as their number. Next-best-view (NBV) selection for 3D Gaussian Splatting (3DGS) usually scores every candidate in the pool and keeps one. Searching for information and choosing a camera, however, are separable problems. We

    embodiedbenchmark
  530. arxiv:2609.34175 · cs.RO
    FailPatch: Failure Residual Patching for Vision-Language-Action Models
    Peng Yu, Jiacheng Wang, Ziheng Zhang, Xuchong Zhang +5

    Vision-Language-Action (VLA) policies are typically adapted using successful demonstrations, which provide direct action supervision but rarely cover failure-prone states. Deployment failures expose these states, yet lack the corrective actions needed for conventional supervised learning. We propose

    vision-language-actionvlavla policyrobotwin
  531. arxiv:2609.34171 · physics.optics
    Monolithically integrated photonic neural network at driven-dissipative criticality
    Yuming Zhang, Ruoran Wang, Jingcheng Li, Hailong Zhou +3

    Artificial intelligence increasingly demands computing architectures capable of adaptive representation, temporal information processing, and robust operation under uncertainty. However, conventional photonic neural architectures typically rely on predefined computational operations and separate fun

    silicon photonic
  532. arxiv:2609.34170 · cs.RO
    RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models
    Yuhan Chen, Ke Yu, Pengfei Liu, Shuxun Wang +2

    Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult

    vision-language-actionvlamanipulation
  533. arxiv:2609.34167 · cs.CV
    Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction
    Juhyeon Park, Yeonwoo Kim, Peter Yongho Kim, Yansen Wang +4

    Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a

    benchmark
  534. arxiv:2609.34163 · cs.RO
    Reliability-Aware Sparse Route Memory for Round-Trip Vision-Language Navigation
    Bojun Long, Lingfan Bao, Tianhu Peng, Jingcheng Sun +1

    Vision-language navigation (VLN) is typically evaluated as a one-way task, although deployed robots may need to return after reaching a goal. We study continuous round-trip VLN and diagnose failures in directional observability, deviation recovery, and termination stability. We propose a reliability

    memory
  535. arxiv:2609.34160 · cs.AI
    RoutePrism: Tracing Construction Order Effects in Agent Memory
    Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang +4

    Processing the same records in a different order can discard different evidence, yet endpoint accuracy alone cannot reveal what changed or whether it mattered. We introduce RoutePrism, a diagnostic protocol that builds memory twice from the same source pool in two processing orders, then traces whic

    memoryagent memoryagent
  536. arxiv:2609.34159 · cs.LG
    WorldGraph: Graph-Native World Modeling
    Zezhong Ding, Yipeng Li, Xike Xie

    World models infer latent states of an environment to capture its underlying dynamics and predict future evolution. Many real-world environments, however, are inherently relational and observed as evolving graphs, where entities, relations, and their properties change over time. Prior graph-related

    world modelbenchmark
  537. arxiv:2609.34157 · cs.AI
    TableSeek: Structure-Preserving Agentic Evidence Seeking over Heterogeneous Table Corpora
    Jiaming Tian, Liyao Li, Wentao Ye, Haobo Wang +4

    Open-domain table retrieval seeks tables that contain sufficient evidence for answering a question or verifying a claim. Yet semantic relevance is often misleading: topically similar tables may lack the required facts, while answer-bearing evidence is often confined to a few cells whose meaning depe

    agentllm agentagenticbenchmark
  538. arxiv:2609.34151 · cs.AI
    Self-Evolving Agents via Likelihood-Guided Tool-Space Optimization
    Xuanqi Zhang, Ruinan Jin, Running Yang, Yuxuan Zhang +3

    Self-evolving agents can continually improve their behavior, while tools define the executable action space through which they interact with the environment. However, exposing the full tool library to model introduces substantial irrelevant context and can impair tool-use decisions. We study tool-sp

    tool-useself-evolvingbenchmark
  539. arxiv:2609.34149 · cs.CV
    Functional Hand Type Prior for 3D Hand Pose Estimation and Action Recognition from Egocentric View Monocular Videos
    Wonseok Roh, Seung Hyun Lee, Won Jeong Ryoo, Jakyung Lee +4

    Current methods for egocentric view action recognition often face challenges in perceiving dynamic hand movements relying solely on geometrical or physical information. In this work, we effectively address this problem by gaining insights into the correlation between functional hand configurations a

    benchmark
  540. arxiv:2609.34145 · cs.RO
    Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving
    Jiaxin Liu, Ruilin Yu, Liang Peng, Jingkai Wang +7

    Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, an

    retrieval-augmentedknowledge graph
  541. arxiv:2609.34144 · cs.CV
    CAST: Reconstruction-Coupled Acceleration of Interactive World Models
    Leyang Chen, Junyi Wu, Fanqing Kong, Shaoqiu Zhang +1

    Interactive world models must respond quickly to controls while preserving scene consistency. Existing acceleration methods can miss heterogeneous control responses and spatial transport when recovering skipped features. We observe that interaction-induced feature changes correlate with approximatio

    world model
  542. arxiv:2609.34143 · cs.CV
    Beyond Geometry: Benchmarking and Consistency Reasoning for 3D Logical Anomaly Detection
    Zhiqiang Qin, He Xie, Junfei Yi, Yang Yang +4

    Existing 3D industrial anomaly detection mainly targets local geometric deviations. In contrast, many industrial anomalies violate object-level design or assembly rules, which we define as 3D logical anomalies. To address these challenges, we introduce the Industrial Logical Anomaly Detection Datase

    benchmark
  543. arxiv:2609.34139 · cs.AI
    Same Tasks, Different Apps: Why Mobile GUI Agents Fail to Generalize?
    Tien Tran, Namho Koh, Daiki E. Matsunaga, Ayush Jain +1

    Mobile GUI agents deployed in real settings must work across different applications that support the same functionality. Most existing benchmarks test each task in only one app, so a high score can mean the agent understands the task, or only that it knows that particular app. We introduce AnyAppBen

    agentbenchmarkleaderboard
  544. arxiv:2609.34136 · cs.AI
    Waggle: Learning One Anonymous Local Law for Self-Organizing LLM Swarms
    Mingxi Zou, Wei Zhu, Zhuo Wang, Langzhang Liang +3

    As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable

    llm agentmulti-agentagent system
  545. arxiv:2609.34135 · cs.AI
    Evo2Team: When Do Evolved Skills Transfer? From Selection to Deployment
    Renxiang Wang, Jiaming Cui

    A skill bank that helps one multi-agent system may leave another's behavior unchanged. A transferred rule helps only when target agents act on it successfully. We study this path for routing and communication skills in Count-Frequency and AgentsNet, using teams of 4--32 agents and GPT and Qwen model

    multi-agentagent system
  546. arxiv:2609.34134 · cs.AI
    StateGuard: Analytical-State Management with Validity-Aware Intervention for Long-Horizon Data Agents
    Wenle Liao, Zhao Wang, Jingchao Zhang, Jiajie Jin +2

    LLM-based agents have shown strong capabilities in automated data analysis and are increasingly moving toward long-horizon, multi-stage analytical workflows. However, as the analytical process evolves, constraints, variables, and conclusions remain implicitly embedded in interaction histories, makin

    memorybenchmark
  547. arxiv:2609.34133 · cs.CV
    PrefLUT: Reusable and Refinable Personalized Color Editing from Pairwise Preferences
    Chuanzhi Xu, Langyi Chen, Chengkun Yue, Xuanhua Yin +4

    Photographic color editing is inherently personal: the same image can appear too warm, too muted, or already satisfactory to different users. Most lookup table (LUT) and reference-guided methods target a specified appearance rather than model persistent preferences from repeated user choices. To add

    evaluation protocol
  548. arxiv:2609.34132 · cs.AI
    From Attack Success to Attack Severity: Counterfactual Memory Attacks on LLM Agents
    Mingxi Zou, Langzhang Liang, Zhuo Wang, Yiyang Zhao +2

    As LLM agents increasingly rely on persistent memory for long-horizon and personalized behavior, they can retain and reuse information across interactions, but this also creates a lasting channel through which malicious memory writes can influence future behavior. Persistent-memory attacks are typic

    memorypersistent statepersistent memoryagent memoryagentllm agent
  549. arxiv:2609.34127 · cs.AI
    STITCH-RAG: Spatio-Temporal Influence Tracing over Topic Hypergraphs for Multi-Hop Retrieval-Augmented Generation
    Haodong Yang, Mengzhu Chen, Jia Cai

    Multi-hop retrieval-augmented generation requires a retriever to connect evidence distributed across documents while preserving a concise, faithful generation context. Existing indexes leave two complementary gaps: chunk-based RAG can break cross-passage evidence chains, whereas an unlabeled pairwis

    retrieval-augmentedragbenchmark
  550. arxiv:2609.34126 · cs.AI
    JET: Judge-Guided Evolution at Test Time for Agent Programs
    Yao Long Teng, Jiayi Cai, Bo An

    An agent's executable program governs how it uses tools, processes observations, and responds to failures. Evolving this program at test time can help adaptation, but deciding which changes to retain is difficult when true rewards are unavailable. Execution traces provide evidence of agent behavior,

    agentevaluator
  551. arxiv:2609.34125 · cs.CL
    Understanding Clinical Cognitive Dialogues Using Large Language Models
    Vishalakshi Arumugam, Dan Schumacher, Veronica Rammouz, Enrique Gonzalez Guerrero +2

    In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study

    benchmarkevaluation framework
  552. arxiv:2609.34124 · cs.CV
    SpatialSkill: Self-Evolving Skills for Cross-View Spatial Reasoning
    Ruifan Zuo, Guocheng Hu, Wanshui Gan, Junyi Wang +2

    Cross-view spatial reasoning requires a model to align different viewpoints into a coherent spatial representation, yet this ability remains challenging for vision-language models despite being natural to humans. Existing methods typically improve spatial reasoning by updating model weights, which k

    self-evolving
  553. arxiv:2609.34123 · cs.LG
    What Does a Stream Model Buy You in Flow Matching?
    Jian Xu

    Stream-level flow matching replaces the linear interpolant of conditional flow matching (CFM) by a Gaussian-process (GP) stream connecting each source--target pair, and reports lower sample error than \icfm{} on 2-Gaussian, MNIST and CIFAR-10 benchmarks. We ask what such a stream model actually cont

    benchmark
  554. arxiv:2609.34117 · cs.LG
    SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
    Gunho Park, Kyoungho Jeun, Juntaek Oh, Byeongjun Shin +2

    Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality

    benchmark
  555. arxiv:2609.34115 · cs.LG
    Forecast-Necessary Causal Discovery for Nonlinear Political Panel Data: Feedback, Functional Form, and the Dynamics of Democratization
    Michael Coppedge, Dmitry Zaytsev, Valentina Kuskova

    A non-significant coefficient in a dynamic panel model need not imply the absence of a relationship. It may instead reflect heterogeneous effects averaged toward zero, reciprocal dynamics overlooked by a recursive specification, or relationships masked by the omission of correlated covariates. Stand

    benchmark
  556. arxiv:2609.34112 · cs.AI
    Unknown is not normal: separating language-model extraction from rule-based decision logic for clinical risk scores
    Nicolás Vera Zúñiga

    Large language models (LLMs) are increasingly used to compute clinical risk scores from free-text notes. Notes are often incomplete, and treating undocumented findings as normal can silently misclassify patients. We test whether separating three-state extraction (present, absent or unknown, by an LL

    agentllm agent
  557. arxiv:2609.34105 · physics.optics
    Stimulated Electro-optic Scattering
    Violet Workman, Gaurav Bahl

    Stimulated Brillouin Scattering (SBS) couples light to acoustic waves and underpins applications ranging from sensing and signal processing to quantum photonics. In solids, this interaction is generally attributed to photoelasticity and the motion of dielectric boundaries. Here we show that piezoele

    quantum photonic
  558. arxiv:2609.34098 · cs.LG
    Beyond Correctness: Evaluating Semantic Knowledge in Cross-Table Transfer
    Seokyong Sheem, Hochang Lee, Suyeong Lee, Daekyum Kim

    Semantic knowledge is increasingly used to bridge heterogeneous schemas in tabular learning, but how much does that knowledge actually improve prediction? Studies in tabular learning commonly answer this question through semantic ablations that modify or suppress the supplied semantic knowledge. We

    evaluation framework
  559. arxiv:2609.34095 · physics.optics
    Tensor-Engineered Van der Waals NbOCl2 Resonant Metasurface for Polarization-entangled Bell State Generation
    Xin Zeng, Wenna Du, Yun-Kun Wu, Yuyang Zhang +19

    Polarization-entangled photon pairs are essential resources for quantum information technologies, yet realizing compact sources with intrinsically controllable entanglement remains challenging. Van der Waals (vdW) nonlinear materials such as NbOCl2 provide atomically thin platforms for quantum light

    manipulationquantum photonic
  560. arxiv:2609.34088 · cs.LG
    TRACE: Expert-Aligned ECG Representation Learning with Rigorous Benchmarking and Real-World Validation in Acute Cardiac Care
    Lovely Yeswanth Panchumarthi, Andrew Lu, Saurabh Kataria, Delgersuren Bold +10

    TRACE (Text-Reinforced Analysis of Cardio ECGs) is a multimodal electrocardiogram (ECG) representation model that learns clinically grounded signal embeddings for downstream cardiac classification. It is designed to address the limitations of existing CLIP-style training, which often struggles with

    benchmark
  561. arxiv:2609.34085 · cs.RO
    AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
    Haoran Zhu, Wancong Zhang, Yann LeCun, Anna Choromanska

    Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM

    world modelaction-conditionedbenchmark
  562. arxiv:2609.34083 · cs.LG
    Beyond One Epoch: Uncertainty-Weighted Sensitivity Regularization for Recommendation Models
    Richard Lettich, Shagun Gupta

    Recommendation models with sparse embeddings and a shared consumer often exhibit the one-epoch phenomenon: a second epoch lowers training loss while sharply degrading generalization. We present a view based on the violation of the prequential principle. On the first epoch, an example's label has not

    benchmark
  563. arxiv:2609.34082 · cs.CV
    K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings
    Yunfei Bai, Enrico Chionna, Akash Amol, Kawaljit Singh KC +1

    Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillatio

    self-improvingpost-training
  564. arxiv:2609.34079 · cs.AI
    GenoMorph: Pathway-Grounded Genomic Disease Reasoning via Adaptive Latent Computation
    Tanmoy Kanti Halder, Akash Ghosh, Arijit Roy, Sriparna Saha

    Large language models (LLMs) have demonstrated strong capabilities in biological reasoning; however, genomic disease inference remains largely dependent on memorized gene-disease associations rather than understanding biological pathways. This shortcut learning undermines robustness and generalizati

    benchmark
  565. arxiv:2609.34078 · cs.CV
    WhiteCon: Semi-Supervised Domain Adaptation Regression Through Whitening Transform and Dual Consistency
    Se Jin Sim, Seoung Bum Kim

    Domain adaptation is crucial for addressing distributional shifts that degrade model performance across domains. While most existing research has centered on classification, semi-supervised domain adaptation regression (SSDAR) for continuous-output tasks remains largely unexplored, particularly in p

    benchmark
  566. arxiv:2609.34077 · cs.LG
    MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
    Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui +2

    Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing all

    memorybenchmark
  567. arxiv:2609.34072 · cs.AI
    PhysFieldBench: Can Multimodal Models Understand Physical Fields?
    Yuezhou Ma, Huikun Weng, Jialong Wu, Chenyi Zhao +4

    Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoni

    post-trainingbenchmark
  568. arxiv:2609.34069 · cs.AI
    Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
    Piyush Jha, Aishik Ghosh, Vijay Ganesh

    The upgrade and rewriting of large scientific codebases has traditionally been a major challenge. While evolutionary search with large language models (LLMs) can port and accelerate legacy code, repair feedback in prompts alone does not prevent subsequent candidates from repeating the same errors. W

    agenticself-improving
  569. arxiv:2609.34064 · cs.LG
    Learning Perturbation Robust Policies for LLM Agents with Stable Optimization
    Pengxin Wang, Yuanzhe LI, Yuxin Ren, Huanrui Yang +1

    Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study

    llm agentpost-training
  570. arxiv:2609.34061 · cs.RO
    Quantile Head for Vision-Language-Action Models
    Xuan Wang, Yinan Wu, Haoran Duan, Jungong Han

    Vision-Language-Action (VLA) models integrate pretrained Vision-Language Models (VLMs) with action heads for robot control. Common action heads have distinct limitations: point regression provides only a point estimate of the action distribution, while standard flow-matching samplers require costly

    vision-language-actionaction headlibero
  571. arxiv:2609.34060 · cs.LG
    KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
    Hyesung Jeon, Hyeongju Ha, Seoyoung Lee, Beomseok Kang +1

    Prompt-specialized multi-agent systems enable multiple agents to share a model while performing complementary roles to solve complex tasks. However, agent-specific prefixes change the KV cache generated for the same shared context, causing each agent to repeatedly prefill the growing context and con

    memoryagentmulti-agentagent system
  572. arxiv:2609.34058 · cs.LG
    Do World Models Learn Global Understanding?
    Alexander Detkov, Matt Thomson

    AI systems often feel brittle and fragmented. A large language model (LLM) may correctly explain a concept but fail to apply it, or follow safety instructions in one context but not another. This behavior suggests a general failure to lift local information to a global understanding. To gain fundame

    embodiedworld model
  573. arxiv:2609.34054 · cs.LG
    PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction
    Hyesung Jeon, Hyeongju Ha, Jae-Joon Kim

    Multi-LoRA agent systems enable efficient role specialization by sharing a common backbone model. However, each agent repeatedly processes the growing shared trajectory and constructs its own KV cache, introducing substantial memory and computation redundancy in long-horizon tasks. Existing KV cache

    memoryagentagent systemagent benchmarkbenchmark
  574. arxiv:2609.34049 · cs.AI
    Thinking Outside the Box: Retention and Transmission of Information in Sliding-Window KV Inference
    Timothy DeLise, Seth Cromelin

    Sliding-window KV inference refers to processing a sequence incrementally while retaining only a fixed-size cache of recent key and value states. It can be applied to pretrained causal transformers at inference time without additional training, while its KV-cache memory remains fixed as more tokens

    memory
  575. arxiv:2609.34047 · cs.CV
    ARCH-B: Architectural Representation, Comprehension and Hierarchy Benchmark
    Kieran Sagar Parikh, Jose Luis Garcia del Castillo y Lopez

    Multimodal models increasingly interpret visual environments, but their ability to recognize the same building across photographs, floor plans, elevations, sections, and renderings remains poorly characterized. We introduce ARCH-B, a benchmark of 354 four-choice questions across 11 cross-representat

    benchmark
  576. arxiv:2609.34045 · cs.LG
    Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines
    Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya

    Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. Existing systems pool such capacity across open swarms anyone may join, which a group admitting only trusted machines cannot use. Bounding mem

    memory
  577. arxiv:2609.34044 · cs.CV
    SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
    Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer +11

    Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual i

    benchmark
  578. arxiv:2609.34041 · cs.LG
    Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
    Samuel Tetteh, Cody Fleming

    Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can pro

    benchmark
  579. arxiv:2609.34039 · cs.AI
    Large Language Models for Structured Clinical Data Analysis: Dual-Agent Grounding and Validation
    Erfan D. Dehkalani, Seetha Shankaran, Abbot R. Laptook, C. Michael Cotten +2

    Objective: To develop and characterize CLEAR-Med, a dual-agent framework for natural-language analysis of structured clinical data that separates SQL-based invocation from independent validation. Methods: CLEAR-Med uses one agent to translate a question into executable Structured Query Language (SQL

    agentagent frameworkbenchmark
  580. arxiv:2609.34036 · cs.LG
    UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents
    Wenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang +1

    On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncert

    agentic
  581. arxiv:2609.34035 · cs.RO
    3D Point Tracking with State Space Models
    Masahiro Ogawa, Qi An, Atsushi Yamashita

    Tracking any point of a dynamic scene in metric 3D - in absolute meters, not up to an unknown scale - underpins 3D and 4D reconstruction, robot navigation, and autonomous driving, where decisions are made in meters, not pixels. Our objective is a 3D point tracker accurate in those absolute terms and

    memorybenchmarkevaluator
  582. arxiv:2609.34034 · cs.LG
    ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling
    Matei-Ioan Stan, Oliver Rhodes

    A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities that have secured the Transformer's status as the de facto standard in sequence modelling. Any realisti

    memory
  583. arxiv:2609.34024 · cs.LG
    Jev in Medicine: A Benchmark Evaluation. Preliminary Results
    Alfredo Madrid-García, Beatriz Merino-Barbancho

    Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaM

    benchmark
  584. arxiv:2609.34019 · cs.LG
    SR4-Fit: A Unified Interpretable Rule-Based Machine Learning Framework for Informative and Trustworthy Decision-Making
    Shyam Sundar Murali Krishnan, Dean Frederick Hougen

    In many high-stakes applications, machine learning is dominated by black-box models that require post hoc explanations to justify their predictions. These explanations are often unreliable because they do not reflect the model's actual computations, limiting accountability and trust. A natural alter

    benchmark
  585. arxiv:2609.34018 · cs.RO
    Estimate, Don't Imitate: Reusing Differentiable State-Based Policies for Visuomotor Control
    Denis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz, Paul Mattes +2

    Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is train

    manipulationsim-to-real
  586. arxiv:2609.34017 · cs.AI
    Maat: Independent Deterministic Contract-Based Governance for Multi-Agent LLM Workflows
    Uliana Elina

    Large-language-model multi-agent systems (LLM-MAS) introduce a characteristic reliability problem: an error produced by one agent can be accepted as context by downstream agents and propagate across the workflow. Many proposed safeguards rely on learned or LLM-based judges whose verdicts are themsel

    agentmulti-agentagent systembenchmark
  587. arxiv:2609.34010 · cs.RO
    ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim
    Ivan Kapelyukh, Xiaohan Zhang, Stephen James, Laura Herlant +1

    We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses

    manipulationgrasp
  588. arxiv:2609.34006 · cs.RO
    TacGooseBumps (TacGB): Retrofitting Normal-Only Tactile Sensors with Shear Encoding for Learning Contact-Rich Manipulation
    Wenjie Li, Binyu Yang, Yuxin Chen, Ambrose Wang +1

    Contact-rich policies often fail because distinct physical states look alike yet require different actions. Cameras may not reveal whether a connector is aligned or fully seated, while many normal-only tactile sensors can miss the tangential interactions perpendicular to the grasping direction that

    manipulationtactilegrasp
  589. arxiv:2609.34004 · cs.LG
    RICE-Alpha: Reliability-Informed Correction with Event Graphs for LLM-Agent Stock Forecasting
    Tong Liu, Lanmiao Liu, Xiang Hu

    Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, information availability, and transition reliability are modeled. Existing LLM-based financial agents incorporate historical evidence, yet they provide limited

    memoryagent
  590. arxiv:2609.33998 · cs.CV
    MetaSampling: Making Frame Samplers Efficient for Long-Video Question Answering
    Ashim Dahal, Bikramjit Banerjee

    Frame selection is an important component of long-video question answering (VQA) with Multimodal Large Language Models (MLLMs). Existing frame-selection methods improve over simple top-$k$ embedding retrieval and uniform sampling, but are typically applied under a fixed global selection budget. We i

    benchmark
  591. arxiv:2609.33994 · cs.AI
    A packet-level digital hardware twin for commissioning megahertz diagnostic edge AI and plasma control system integration in tokamaks
    Semin Joung, Abhilasha Dave, Luca Scomparin, Filipp Khabanov +5

    High-bandwidth plasma diagnostics increasingly provide inputs to machine learning and signal-processing algorithms intended for real-time tokamak control, but the complete path from diagnostic sampling to control-system handoff is difficult to commission because of their sampling rates. We develop a

    memory
  592. arxiv:2609.33993 · cs.LG
    ASTRA: ADMM-Accelerated Topology Reconfiguration for Dynamic Satellite Constellations
    João Norberto, Ricardo Ferreira, Cláudia Soares

    Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Satellite Topology via Regret-Aware learni

    online learning
  593. arxiv:2609.33991 · cs.CV
    A Multi-Dataset Benchmark of YOLO-Based Weed Detection in Precision Agriculture
    Hristina Zdraveska, Vlatko Spasev, Ivica Dimitrovski, Ivan Kitanovski +1

    Weed detection is an important component of precision agriculture, enabling site-specific weed management and reducing unnecessary herbicide use. Although deep learning methods have achieved strong results for crop and weed detection, many studies rely on single-dataset evaluation, making it difficu

    benchmark
  594. arxiv:2609.33989 · cs.LG
    RewardExplainer: Learning Reward Model Explanations from Counterfactual Preference Feedback
    Jingyi He, Nier Wu, Shuang Liu, Xin Wang +2

    Reward models (RMs) are a key component of large language model post-training, providing reward signals for subsequent reinforcement learning. However, conventional discriminative RMs typically output only scalar scores, making it difficult to identify the response behaviors associated with their sc

    post-trainingbenchmark
  595. arxiv:2609.33987 · cs.CL
    Opera: A Verbal Critic Framework for Long-horizon Coding Agents
    Kai Mei, Zhiyuan Hu, Yutong Dai, Juntao Tan +6

    Long-horizon coding agents need timely corrections, yet feedback can be ineffective or even harmful when it misjudges ongoing work or fails to address the underlying problem. Existing critics focus on evaluating trajectories and generating feedback, but rarely track what happens after feedback is de

    benchmark
  596. arxiv:2609.33986 · cs.LG
    ICMAPE: In-Context Multiagent Pure Exploration
    Xinyi Hu, Alessio Russo, Aldo Pacchiano

    In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent

    multi-agentagent systembenchmark
  597. arxiv:2609.33984 · cs.LG
    From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting
    Chaoqi Zhang, Yu Wang, Haixu Tang

    Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingul

    benchmark
  598. arxiv:2609.33982 · cs.RO
    Test-Time Spatial Reasoning for Robot Manipulation Using Generative Real-to-Sim
    Ivan Kapelyukh, Yafei Hu, Ran Gong, Brandon May +5

    Spatial reasoning is fundamental to general robot intelligence, as it enables robots to complete long-horizon tasks involving multi-object interaction. We introduce Simify, a training-free, test-time framework that performs explicit spatial reasoning via massively parallel physics simulation. From a

    manipulationsim-to-real
  599. arxiv:2609.33981 · cs.LG
    Future Information-Directed Sampling for Bayesian Nonstationary Bandits
    Yichen Song, Alessio Russo, Aldo Pacchiano

    Exploration--exploitation is a central trade-off in bandit learning. While classical algorithms such as upper confidence bound methods and Thompson Sampling effectively balance this trade-off in stationary environments, their exploration strategies mainly reduce uncertainty about the current optimal

    benchmark
  600. arxiv:2609.33980 · cs.LG
    DynGraphAgentBench: A Benchmark for Agentic Lifecycle Control in Dynamic Graph Anomaly Detection
    Yuwei Han, Lingwei Wei, Wooseong Yang, Liangjie Huang +3

    Dynamic graph anomaly detection requires repeated decisions as graph structure and class prevalence drift, yet detector benchmarks usually score a fixed pipeline after current labels are known. We introduce DynGraphAgentBench, an executable benchmark for agentic lifecycle control under delayed feedb

    agenticbenchmark
  601. arxiv:2609.33974 · cs.CL
    Beyond Solo and Consistency: Vindicating Multi-Agent Debate via Conditional Progressive Pruning
    Ruosong Ye, Caiqi Zhang, Jiahao Li, Haijun Wu +8

    Large Language Model (LLM) based Multi-Agent Debate (MAD) is one of the most effective test time scaling techniques. Through multi-round communication, agents complement each other in knowledge and reasoning and solve tasks that no single member can solve. However, existing MAD frameworks fail to be

    agentmulti-agentbenchmark
  602. arxiv:2609.33973 · cs.RO
    FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving
    Yutong Liang, Quanquan Peng, Matthew Kim, Xiaolong Wang

    Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and re

    dexterousgrasp
  603. arxiv:2609.33969 · cs.CV
    Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors
    Ge Gao, Siyue Teng, Chanqgi Wang, Fan Zhang +4

    Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formu

    memory
  604. arxiv:2609.33967 · cs.LG
    ThinkNet: Compact Architecture Selection and Validation-Gated Ensembles for Subject-Independent MI-EEG Decoding
    Abdul Basit, Saim Rehman, Muhammad Shafique

    Practical assistive and rehabilitative brain--computer interfaces require subject-independent motor-imagery EEG (MI-EEG) decoders that generalize to new users under limited target-user data and constrained compute. However, held-out-subject performance can be overstated when test-subject information

    benchmark
  605. arxiv:2609.33965 · cs.LG
    Greenpixie's AI Token Methodology: Assessing the Energy, Water and $\mathrm{CO_2\text{-}eq}$ Impact of AI Tokens for Open and Closed Weight Models
    Joshua Horswill, Ross Hunter, Matt Clifford, James Hall

    We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a

    embodiedbenchmark
  606. arxiv:2609.33955 · cs.AI
    Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents
    Kaiwen Luo, Ming Gao

    Many LLM-as-a-judge evaluations score fixed outputs under a fixed task definition. Production multi-turn business agents instead require a maintained measurement system: correctness depends on business-specific facts and procedures, outcomes emerge across turns, and failures must be attributed to ei

    agenthuman-in-the-loop
  607. arxiv:2609.33947 · cs.LG
    Simple Diffusion Language Models Are More Effective Few-Step Generators Than Reported
    Hasan Amin, Ming Yin, Rajiv Khanna

    Diffusion language models (DLMs) promise fast parallel generation, yet high-quality samples often require large number of refinement steps, which diminishes their advantage in practice. This has led to massive interest in and rapid development of new methods for effective few-step generation. We sho

    benchmark
  608. arxiv:2609.33944 · cs.RO
    ReSync: Re-Aligning the Two Clocks of Asynchronous World-Action Models
    Xi Lin, Feihong Zhang, Yulong Shi, Yanghong Mei +5

    Jointly generating future video and actions has become a standard recipe for world-action models, and the strongest systems denoise the two streams on separate schedules: actions are decoded in few steps so control stays fast, while the video stream runs longer to keep the predicted future sharp. Th

    benchmark
  609. arxiv:2609.33940 · cs.LG
    Behavioral Monitoring of JEPA World Models with Jacobian Centroids
    Thomas Walker, Randall Balestriero, Richard Baraniuk

    Detecting failures in World Model (WM)-based planning requires monitoring whether the model is behaviorally aligned with the current task, which in turn requires studying its internal representations. Here, we show that centroids---sub-component Jacobian row-sums---effectively identify the behaviora

    world model
  610. arxiv:2609.33937 · cs.CV
    Test-Time Generalized Category Discovery
    Shambhavi Mishra, Omprakash Chakraborty, Julio Silva-Rodriguez, Ismail Ben Ayed +2

    Test-Time Adaptation (TTA) and Generalized Category Discovery (GCD) are traditionally treated as disjoint problems: the former adapts models to domain shift assuming all test classes are known, while the latter discovers novel categories assuming labeled training data for known classes. However, rea

    benchmark
  611. arxiv:2609.33935 · cs.CV
    Video, Ergo Genero: Unifying Video Tasks via Spatiotemporal Analogy
    Chia-Hsiang Kao, Belinda Zeng, Bharath Hariharan, Menglin Jia

    Adapting video models to new tasks typically requires dedicated data curation and fine-tuning. While visual analogy provides a training-free alternative by specifying tasks in-context, it remains restricted to the image domain. To explore whether analogy-based methods can unify diverse video tasks a

    manipulationevent camera
  612. arxiv:2609.33927 · cs.LG
    Optimizing the Phi-2 Small Language Model for Real-time Chatbot Applications Using Parameter-Efficient Fine-Tuning (PEFT) with QLoRA Quantization
    PhanTan Khanh Nguyen, Ashfaq Ali Shafin, Khandaker Mamun Ahmed

    This study explores the optimization of the Phi-2 Small Language Models (SLMs) for real-time chatbot applications through Parameter-Efficient Fine-Tuning (PEFT) and Quantized Low-Rank Adaptation (QLoRA). QLoRA specifically refers to the integration of PEFT with LoRA alongside a 4-bit quantization pr

    memory
  613. arxiv:2609.33920 · cs.AI
    HyperMCTS: Hypergraph-Augmented MCTS for Long-Horizon LLM Agents
    Tingsong Xiao, Nithish Balachandar Moudhgalya, Chandrayee Basu, Lichao Wang +4

    Long-horizon tasks require large language model (LLM) agents to coordinate decisions under constraints that span an entire solution. Monte Carlo Tree Search (MCTS) offers a promising approach to test-time scaling by exploring alternative action trajectories, but model computation and environment int

    llm agent
  614. arxiv:2609.33915 · cs.LG
    Training Witnesses: Trusting the Training without Trusting the Trainer
    Houjun Liu, Pratyusha Sharma

    Progress in machine learning cannot outpace our ability to verify it. With an explosion in papers today, every scientific claim rests initially on trust in the trainer, leading to uneven evaluation, baselines, and forestalling of reliable progress. Traditionally, the burden of verification falls on

    leaderboard
  615. arxiv:2609.33910 · cs.AI
    When Consent Outlives Context: Residual Authority Replay in Long-Lived Agents
    Zhihao Zhang, Chao Wang, Rujia Li, Qingze Wang +2

    LLM agents increasingly rely on user approval to authorize security-sensitive actions at runtime. Such approvals are granted within a specific task and execution context. In long-lived agents, authorization decisions may need to persist across tasks or sessions. We find that this continuity can outl

    llm agent
  616. arxiv:2609.33906 · cs.LG
    JIVE: Jacobian-Informed Volume Expansion for Diverse Generative Sampling
    Guangxun Zhang, Brian Cai, Boxuan Zhang, Chao Chen +1

    Generative models often suffer from mode collapse and limited sample diversity. While prior works attempt to mitigate this by jointly generating a batch of samples and repelling their trajectories, these heuristics do not explicitly maximize the diversity of the resulting endpoints. We introduce JIV

    benchmark
  617. arxiv:2609.33905 · cs.CL
    SlopBench: How Well Can We Rank Language Models by Slop? A Multi-Domain Benchmark of Repetitive AI Writing
    Dhruv Roongta, Harsha Gaddipati, Anh Tuan Huynh

    SlopBench asks which models produce the stiff, repetitive prose readers call AI slop, a question detectors leave open once they have classified a text as machine-written. We evaluated eighteen models on 112 hand-written tasks in email, social posts, essays, and workplace chat, sampling each model on

    benchmark
  618. arxiv:2609.33901 · cs.LG
    Finite Probes Suffice: Identifiability and Universality for Weight-Space Learning
    Soutrik Sarangi, Yonatan Sverdlov, Adir Dayan, Haggai Maron +1

    Learning properties of neural networks has recently attracted growing interest, with existing approaches operating either directly on network parameters or through probe-based representations of network behavior. While probing methods have shown strong empirical performance, their theoretical founda

    benchmark
  619. arxiv:2609.33899 · cs.CL
    NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models
    Ziwei Chen

    We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final t

    benchmark
  620. arxiv:2609.33893 · cs.LG
    MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models
    Zhi Wen Soi, Giulio Segalini, Jian-Jia Chen, Lydia Chen

    Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the aud

    benchmark
  621. arxiv:2609.33889 · cs.LG
    Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding
    Jungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo +1

    At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. Activation sparsity trims the first term and KV-cache sparsity the second, yet their reported speedups are hard to c

    long context
  622. arxiv:2609.33887 · cs.LG
    Faster Block-Diffusion Serving with Distribution-Free Risk Guarantees
    Jungseob Lee, Dongyub Jude Lee, Chanjun Park, Sugyeong Eo +1

    Block-diffusion language models are served at hand-picked operating points, such as acceptance thresholds, buffer depth, schedule, checkpoint and precision, and each point is chosen by its mean benchmark accuracy. However, a mean does not tell an operator how often a faster configuration fails on pr

    benchmark
  623. arxiv:2609.33885 · cs.LG
    Prospective Interpretation Risk: Principled Communication Control Between LLMs
    Wanrong Yang, Rehan Deen, Julian Ma, Yuheng Fan +6

    Large language model (LLM) agentic systems increasingly rely on models communicating with one another, yet existing uncertainty and multi-agent methods rarely estimate how a particular receiver will interpret a message before it is sent. This matters in heterogeneous systems, where capable receivers

    multi-agentagentic
  624. arxiv:2609.33882 · cs.RO
    DexTaG: Tactile-as-Guidance in Reinforcement Learning for Dexterous Manipulation
    Han Yang, Yian Wang, Yunlong Song, Zhenjia Xu +1

    Glove-based motion capture is emerging as a scalable approach to collecting dexterous-hand demonstration data. However, due to the kinematic gap between the human and robot hand, the recorded human motions cannot be executed directly on the robot, especially for contact-rich tool-use tasks involving

    manipulationdexteroustactilegrasptool-use
  625. arxiv:2609.33872 · cs.RO
    Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation
    Sichao Liu, Zekun Wang, Lixuan Tang, Yiming Li +5

    Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions

    manipulationrobot policyevaluation framework
  626. arxiv:2609.33871 · cs.CL
    Population Physics, Population Problems: Safety and Emergence in LLM Societies
    Adrian de Wynter

    The collective behaviour of large language model (LLM) societies is not the sum of their individual outputs. It yields statistically distinct, sometimes-unpredictable phenomena, for which the tools we use to study single agents may not scale. Due to recent incidents involving autonomous agentic syst

    autonomous agentmulti-agentagenticagent system
  627. arxiv:2609.33870 · cs.AI
    When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents
    Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur, Ece Kamar +1

    LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable

    agentllm agentpost-trainingbenchmark
  628. arxiv:2609.33867 · cs.AI
    R$^2$ Flow: Recursive Self-Improvement via Recursive Skill Evolution
    Mingda Zhang, Qiang Huang, Yanjin Li, Zijia Wang +3

    LLM-based agents can improve themselves across tasks by reusing and revising the skills they orchestrate into executable procedures. Flow-based training fits this loop: it samples procedures in proportion to reward, and the flow through each skill credits it for the next library revision. Three obst

    self-improvement
  629. arxiv:2609.33857 · cs.LG
    PI-NOMT: Physics-Informed Neural Optimal Mass Transport for Brain Fluid Dynamics
    Mehmet Emin Acar, Vahit Bugra Yesilkaynak, Helene Benveniste, Gozde Unal

    Recovering hidden transport mechanisms from sparse spatiotemporal observations is a fundamental inverse problem in scientific machine learning. In brain tracer imaging, dynamic contrast-enhanced MRI (DCE-MRI) provides time-resolved measurements of tracer concentration, while the underlying velocity

    post-trainingbenchmark
  630. arxiv:2609.33855 · cs.LG
    Program-Verified Self-Evolution for Vision-Language Models
    Ahmed Heakl, Sungik Choi, Moontae Lee, Salman Khan

    Self-evolving vision-language models train on questions they generate from unlabeled images. Since these questions have no gold answers, prior methods label them by majority vote over sampled answers or by a model judge. In a human evaluation, we find that 24\% of majority-vote labels and 18\% of mo

    scene graphself-evolvingbenchmark
  631. arxiv:2609.33854 · cs.CV
    ReDrive: Shaping Representations with World Modeling for End-to-End Driving
    Yueting Zhu, Shaoyu Chen, Yuehao Song, Hui Sun +3

    Driving policies require capabilities of scene understanding and future evolution prediction. To achieve this goal, current end-to-end models typically construct complex perception-planning pipelines or introduce world models that explicitly predict future states, resulting in a complex system archi

    world model
  632. arxiv:2609.33845 · cs.AI
    How code helps different tasks? A decompositional lens on LLM post-training
    Zheng Yu, Yiwei Li, Yishen Chen, Xiang Li +3

    Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decomposi

    post-trainingpost training
  633. arxiv:2609.33844 · cs.LG
    ViBR-WM: Visual Bayesian Regression for World Modeling
    Jifan Li, Ning Ning

    Modeling temporal dependence and uncertainty is central to forecasting with world models. The Visual Bayesian Regression World Model combines visual features, physical histories and known covariates through interpretable regression, within a modular architecture supporting trend, seasonal and cycle

    world model
  634. arxiv:2609.33836 · cs.RO
    DeltaSeek: Toward Active Perception in Evolving Construction Environments
    Sanjay Acharjee, Md Nazmus Sakib

    Construction environments evolve continuously, causing large geometric changes that degrade static mapping and registration performance. This necessitates active perception, where robots deliberately select sensing configurations to resolve the environment's current state. We present DeltaSeek, an i

    benchmark
  635. arxiv:2609.33832 · cs.RO
    Achieve What You Imagined: Learning to Align Actions with Visual Plans
    Yuheng Qiao, Ziran Wei, Xiaohan Wang, Daqiang Guo +4

    World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before produc

    manipulationaction headworld modelaction-conditioned
  636. arxiv:2609.33822 · cs.AI
    Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory
    Jayant Parashar, Eugene F. Douglass, William C. Bastian, Suchendra M. Bhandarkar

    An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces in

    memoryagentbenchmark
  637. arxiv:2609.33818 · cs.CV
    Augmenting Visual Anomaly Detection with Automated Interpretability
    Antonio De Santis, Arsenio Leo, Marco Brambilla

    Visual anomaly detectors identify deviations from known-normal data, but their anomaly signals may mix evidence of actual anomalies with benign visual variation. We investigate whether automated interpretability can augment visual anomaly detectors by identifying and intervening on different compone

    benchmark
  638. arxiv:2609.33814 · cs.LG
    Annealed Sinkhorn with Momentum: Certified Unregularized Optimal Transport in Linear Memory
    Samuel J. K. Chin, Maximilian Schiffer

    We characterize Bregman Douglas-Rachford splitting (BDRS) for unregularized discrete optimal transport and develop an anytime primal-dual certificate in linear memory. We first establish that BDRS coincides with warm-started Inexact Proximal point method for exact Optimal Transport (IPOT) using a si

    memorybenchmark
  639. arxiv:2609.33812 · cs.LG
    Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
    Eduardo Ariño de la Rubia, Szilard Pafka

    Repeated runs of the same coding agent are known to give different benchmark scores. We ask what that variation means for a team running an agent on its own task, by intensive replication on one machine-learning task: an agent improves the training code of an XGBoost classifier for airline delays, a

    agentbenchmark
  640. arxiv:2609.33810 · cs.LG
    Controlling Speaking Rate in Autoregressive TTS via Activation Steering
    Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi +3

    Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-

    benchmark
  641. arxiv:2609.33807 · cs.RO
    CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation
    Yiheng Lyu, Xueying Jiang, Wenhao Li, Shijian Lu +1

    How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning,

    embodiedmanipulationgraspagenticcode-as-policybenchmark
  642. arxiv:2609.33803 · cs.LG
    Diffusion Reward Models
    Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze WangZiqing Qiao +11

    Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonabl

    rlhfbenchmark
  643. arxiv:2609.33802 · cs.LG
    Autonomous phase discovery
    Shiyu Zhou, Yuxuan Zhang, Sebastian Wetzel, Roger Melko +1

    Understanding quantum phases of matter has long relied on physicists' intuition and mathematical tools such as symmetry and topology. Remarkably successful as these approaches have been, they provide no universal way to explore a Hamiltonian space whose organizing principle is not known in advance.

    benchmark
  644. arxiv:2609.33791 · cs.LG
    Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
    Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi Wang +10

    Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in t

    post-trainingbenchmark
  645. arxiv:2609.33780 · cs.LG
    Selecting Diverse SFT Traces Improves Post-RL Generalization
    Dylan Zhang, Mingyuan Wu, Jinning Li

    Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to

    benchmark
  646. arxiv:2609.33778 · cs.AI
    Evidence-Inference Reconstruction: When The Evidence Is Recalled But The Reasoning Goes Wrong
    Megan Diehl, Ser-Nam Lim

    Modern multi-hop LLM agents are equipped with built-in mechanisms to detect errors in intermediate reasoning steps. Such errors trigger corrective actions from these agents, which mostly follow the paradigm of retrying the steps or the reasoning trajectories. Not only are these retries expensive, we

    llm agentagentic
  647. arxiv:2609.33773 · cs.AI
    Learning Strategies to Break Judges
    Guruprerana Shabadi, Aaditya Naik, Rajeev Alur, Mayur Naik

    As AI agents surpass human performance, it becomes exceedingly hard for system designers to evaluate them directly and understand their failure modes. Consequently, agents themselves are being deployed extensively to evaluate, judge, and provide feedback on model traces. But this raises an important

    agentai agentagentic
  648. arxiv:2609.33772 · cs.AI
    Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents
    Weiyi Xu, Xiaowen Yang, Wen Da, Hang Xu +6

    Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instr

    agentagent benchmarktool usetool-usepost-trainingbenchmark
  649. arxiv:2609.33769 · cs.CV
    M3-Score: Fidelity, Memorization and Coverage as Separate Axes for Evaluating Generative Radiology Image Models
    Sathiyamohan Nishankar, Pubudu Sanjeewani, Asanka Perera

    Quantitative evaluation of generative models for radiology remains challenging. Clinically relevant structures are often small and infrequent, feature spaces learned from natural images may represent them poorly, and a single summary score cannot distinguish limited fidelity from limited diversity.

    evaluation framework
  650. arxiv:2609.33768 · cs.LG
    DEALS: Decentralized Expertise-Aware Load Serving for Multi-Agent LLM Systems
    Jingjuan Huang, Wenbin Wang, Yanchuan Yin, Alvaro Velasquez +1

    Multi-agent systems (MAS) have recently emerged as an effective approach for coordinating large language model (LLM)-based agents to solve complex tasks through structured interactions. In practice, MASs often handle a stream of heterogeneous and complex tasks, requiring agents to decompose each tas

    agentmulti-agentagent system
  651. arxiv:2609.33765 · cs.RO
    Principal Steering Subspaces for Online Adaptation of Frozen Generative Robot Policies
    Jialeng Ni, Nathan Zhao, Kunpeng Song

    Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise

    vision-language-actionvlavla policyhumanoid
  652. arxiv:2609.33764 · cs.LG
    Beyond Fixed Features: Architecture-Dependent Sensitivity to Node Representations under Heterophily
    Priyanath Maji, Sidharth Gaur, Rajavinoth Paul Durai

    Graph Neural Networks (GNNs) perform well on homophilic graphs but struggle in heterophilic settings, where connected nodes often carry dissimilar labels. Existing evaluations typically compare architectures under a fixed node-feature representation, leaving unclear whether conclusions about heterop

    benchmark
  653. arxiv:2609.33763 · cs.AI
    SecProbe: Adaptive Evaluation of Coding Agents on Cybersecurity Vulnerabilities
    Xiaonan Luo, Yue Huang, Kehan Guo, Ping He +8

    Assessing cybersecurity vulnerability awareness in coding agents requires evaluations that reveal capability gaps and remain informative as models evolve. Static benchmarks offer fixed coverage and difficulty, while scarce vulnerable repositories and costly expert authoring limit their renewal at sc

    agentbenchmark
  654. arxiv:2609.33762 · cs.LG
    EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
    Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li +6

    LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsist

    memoryagentllm agent
  655. arxiv:2609.33759 · cs.LG
    Positions Are Not Facts: The Mismatch Between KV Caches and Memory
    Changhai Zhou, Yuhua Zhou, Shiyang Zhang, Jun Gao +4

    When a fact changes, how should a language model update the history stored in its key-value (KV) cache? Hiding the old record is cheap, but it may still contain needed details or answer questions about the past. We compare hiding whole records, hiding only replaced values, and deleting old text and

    memory
  656. arxiv:2609.33758 · cs.CV
    ENet-GP: Unified Document Image Restoration
    Sujal Burad, Aakanksha, A. N. Rajagopalan, Sumit Shekar

    Reliable document digitization in uncontrolled capture settings is challenging because real images exhibit multiple interacting degradations rather than a single isolated distortion. Documents thus captured are affected simultaneously by geometric distortions, like page warping, as well as photometr

    benchmark
  657. arxiv:2609.33757 · cs.LG
    YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
    Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu +31

    Symbolic models make melody, harmony, rhythm, and form explicit but typically stop before a finished recording; audio models produce complete songs while leaving composition implicit. We introduce YuE2, which unifies symbolic and audio music generation at frontier quality through symbolic planning.

    agenticbenchmark
  658. arxiv:2609.33754 · cs.LG
    Collaborative Synthetic Data for Privacy-Preserving Financial Fraud Detection Across Organizational Silos
    Simeon Allmendinger, Domenique Zipperling, Burhanettin Bahadir Kibar, Niklas K{ü}hl

    Organizations seek analytical value from AI, yet relevant data are often fragmented across organizations and constrained by privacy. This is acute in financial fraud detection, where rare fraud cases and imbalanced local datasets limit decision-relevant analytics. Federated learning enables collabor

    benchmark
  659. arxiv:2609.33748 · cs.RO
    AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
    Rui Wang, Xiangyu Wang, Donglin Yang, Yibo Li +3

    World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less

    manipulationrobotwin
  660. arxiv:2609.33746 · cs.LG
    PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention
    Kunming Shao, Jierun Chen, Yanli Wang, Ruoyu Wang +3

    At each decoding step a language model attends over the key-value (KV) cache of every earlier token, so at long context the attention call is bounded by memory bandwidth. Sparse attention reads only a subset of keys chosen by a cheap score estimate, and most methods give the unread tokens zero weigh

    memorylong context
  661. arxiv:2609.33738 · cs.CL
    The Effects of Incremental Instruction Delivery on Language-Model Creative Writing
    Anshuman Singh, Abrar Eyasir, Haseeb Yaqoob, John Manavalan

    Large language models are increasingly used as interactive writing tools, where users develop stories, revise ideas, and introduce new requirements across multiple turns rather than specifying a complete brief upfront. Yet most evidence on multi-turn instruction degradation comes from tasks with obj

    benchmark
  662. arxiv:2609.33737 · cs.RO
    MomWorld: Momentum-Aware Latent World Model for Long-Horizon Autonomous Driving
    Ziying Song, Shengkai Zhang, Lei Yang, Haozhuang Chi +5

    Long-horizon planning enables autonomous vehicles to anticipate scene evolution and potential risks, supporting safe and stable decisions in complex interactions. However, existing methods struggle to propagate motion trends from observed history into the future. Long rollouts based on a single late

    world model
  663. arxiv:2609.33729 · physics.app-ph
    Wave-Domain Semantic Equalization Using a Practical Dynamic Metasurface Antenna with Strong Mutual Coupling
    Yassine Sghaier, Emilio Calvanese Strinati, Philipp del Hougne

    Semantic mismatch between independently trained AI-native agents in heterogeneous networks can impair semantic communications. Hybrid analog-digital semantic equalization can align the incompatible latent representations without retraining the semantic transceivers. We study a practical realization

    benchmark
  664. arxiv:2609.33728 · cs.LG
    ALDER: Discovering the Laws of a World by Acting in It
    Teng Cao, Yu Deng, Quentin Delfosse, Kristian Kersting

    Reliable world models should not only predict future states but express how actions change the world in an explicit, transparent and testable form, such as equations. Yet methods that rely on a fixed set of trajectories cannot distinguish equally good competing hypotheses, while searches over a fixe

    world modelbenchmark
  665. arxiv:2609.33725 · cs.LG
    Reliable Replay through Spatial Coherence in Online Continual Learning
    Haixiang Sun, Jiefu Zhang, Yinghao He, Yang Xu +3

    Continually adapting models to new tasks requires retaining earlier knowledge under limited memory and computation. Experience replay addresses this challenge, but priorities based on individual loss increases overlook how related memories respond to the same update and can overemphasize isolated re

    memory
  666. arxiv:2609.33717 · cs.AI
    Self-Designed Evaluators and Warm Memory for Long-Horizon Agents
    Saeid Asgari, Emre Kiciman, Leonardo de Oliveira Nunes, Ranveer Chandra

    A tool-using language-model agent deployed over a long stream of tasks receives no reward, so it cannot tell whether it succeeded, cannot safely retry, and cannot label the experience it needs to improve. We present SelfSuite, in which the agent's own base model, given only the world's public materi

    memoryagentagenticbenchmarkevaluator
  667. arxiv:2609.33716 · cs.CV
    Revisiting Diffusion Fine-Tuning for Unsupervised Domain Adaptation
    Xuan Qi, Yi Wei, Daniele Berardini, Vito Paolo Pastore +1

    Diffusion-based unsupervised domain adaptation (UDA) improves cross-domain transfer by generating target-specific synthetic data for downstream adaptation. Existing methods are largely designed for single-target adaptation: when a model trained on one labeled source domain must be adapted to multipl

    benchmark
  668. arxiv:2609.33715 · cs.LG
    Theory Guided and Interpretable Neural Operator Design for Partial Differential Equation Learning
    Zeyuan Song, Zheyu Jiang

    Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the

    benchmark
  669. arxiv:2609.33713 · cs.AI
    BIRD: Distilling Decision Boundaries into Rationales for MLLM Adaptation
    Anglin Liu, Yanlin Wu, Ruichao Chen, Yuting Zhang +5

    Adapting general-purpose multimodal large language models (MLLMs) to specialized domains requires learning domain-specific decision criteria, which often hinge on subtle visual distinctions between otherwise plausible answers. Rationale augmentation aims to expose such evidence through additional ob

    self-improving
  670. arxiv:2609.33707 · cs.RO
    Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse
    Futa Waseda, Shuhei Kurita, Isao Echizen

    Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial

    vision-language-actionvlalibero
  671. arxiv:2609.33702 · cs.AI
    Understanding Confabulation and Rethinking Reconstruction in Activation Explanations
    Gert Lek, Zixuan Xia, Pin-Yu Chen, Lydia Y. Chen

    Natural Language Autoencoders (NLAs) produce unsupervised text explanations of a model's activations: a verbalizer describes an activation and a reconstructor learns to recover it from this text. Under the established point-reconstruction NLA training recipe, explanations become more useful for pred

    evaluation framework
  672. arxiv:2609.33699 · cs.AI
    SpecRead: A Benchmark for Measuring Whether Language Models Understand Hardware Specifications
    Feilian Huang

    Existing benchmarks for large language models (LLMs) in hardware design evaluate downstream artifacts such as generated RTL, assertions, or testbenches. When a model fails such a benchmark, the failure is ambiguous: it may have misread the specification, or it may have understood the specification a

    benchmark
  673. arxiv:2609.33691 · cs.CL
    When Do Agents Help? Embedding, LLM and Agentic Alignment of Classical Texts and Their Translations
    Máté Metzger

    Classical texts aligned with their translations support machine translation, retrieval and computational research, but evidence comparing alignment workflows is scattered. This study compares seven systems on 452 texts in Pali, Sanskrit, Mishnaic Hebrew and Tibetan, comprising 9,833 human-aligned un

    agentautonomous agentagentic
  674. arxiv:2609.33688 · cs.LG
    TopoMamba: A Load-Support Relation-Guided Multi-Directional State-Space Model for Topology Optimization
    Bin Lou, Yuxuan Cheng, Huaizhi Zong, Junhui Zhang +1

    Deep learning has emerged as an efficient alternative for predicting high-performance material distributions in topology optimization. Existing methods struggle to accurately capture load-transfer information, limiting out-of-distribution generalization, while their model architectures often incur h

    benchmark
  675. arxiv:2609.33683 · cs.CV
    MAD-Guard: Controlled Study of Autoregressive Generation versus Direct Decision Interfaces for Closed Multimodal Forensic Tasks
    Hao Chen

    When should multimodal foundation models generate tokens, and when should they directly output a decision? We present MAD-Guard, a controlled study of output-decision interfaces for closed multimodal forensic tasks. Once a multimodal representation is computed, is autoregressive generation necessary

    manipulationbenchmark
  676. arxiv:2609.33678 · cs.AI
    SWE-Game: Can Coding Agents Build the Games We Want?
    Xiaoyu Chen, Lai Wei, Jin Wang, Xiangyu Zou +7

    We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-

    agentbenchmarkevaluator
  677. arxiv:2609.33676 · cs.AI
    Auditing Agent Actions through Query-Conditioned Attribution
    Yifan Liu, Praveen Venkateswaran, Abdulhamid Adebayo, Dong Wang

    LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for

    agentllm agentbenchmark
  678. arxiv:2609.33672 · cs.CL
    Reset Is Not Recovery: Evaluating Recoverability from False Conversational Context via Sycophancy Hysteresis
    Adi Shnaidman

    Grounded language models are usually evaluated by adding relevant context, but multiturn dialogue also contains unsupported user claims that may contaminate later factual answers. We study post-pressure recoverability: whether a model returns to clean-context behavior after a user repeatedly advocat

    benchmark
  679. arxiv:2609.33668 · cs.CV
    When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling
    Ping Guo, Zhiqi Huang, Xinran Li

    Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degra

    benchmark
  680. arxiv:2609.33667 · cs.LG
    Scalable Attribution and Control of Model Behavior During Training
    Sleem Abdelghafar

    Attributing and controlling model behavior during training requires identifying each example's contribution quickly enough to act before the next update. However, examples in the same training batch can produce similar behavioral changes, making their individual contributions difficult to distinguis

    evaluator
  681. arxiv:2609.33666 · cs.CL
    One Model Is Not a Crowd: Multi-LLM and Aspect-Conditioned Diverse Comment Generation
    Nafis Irtiza Tripto, Delvin Ce Zhang, Mahjabin Nahar, Dongwon Lee

    Human communication on the internet is shaped by diverse perspectives, most visibly expressed in online comment spaces. As large language model (LLM)based AI agents begin to inhabit these spaces, a key question arises: whether synthetic comment threads can capture the diversity inherent in human dis

    ai agent
  682. arxiv:2609.33665 · cs.AI
    CompoWorld: Compositional Environment Scaling for General Agents
    Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu +8

    Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We in

    world modelbenchmark
  683. arxiv:2609.33662 · cs.LG
    Audit-First VAPO: Risk-Certified Selective Updates under Imperfect Verification
    Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang +4

    Imperfect verifiers can assign a harmful update direction even when clipping and regularization bound its magnitude. We introduce Audit-First VAPO, which separates discrete directional admission from continuous magnitude control. An observation-only accept-appeal-abstain policy uses a finite seconda

    benchmark
  684. arxiv:2609.33658 · cs.AI
    AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents
    Tianzhuo Yang, Zirui Mi, Yantao Huang, Guoxi Zhang +3

    Safety alignment for large language models (LLMs) in conversational settings is largely framed around whether to answer or refuse a request. In agentic settings, however, the same models must decide whether to act as permission-critical evidence emerges during execution. This creates a distinct chal

    agentllm agentagenticevaluation framework
  685. arxiv:2609.33655 · cs.LG
    Towards Eliminating Catastrophic Forgetting in the Curriculum Learning of Math Reasoning Tasks
    Zengyan Yang, Yangyang Wu, Kai Huang, Pengfei Lyu +2

    Curriculum learning has found broad application across numerous domains. Nevertheless, its effectiveness is intrinsically curtailed by catastrophic forgetting, driven by the shifts in model parameter distributions between curriculum tasks. In this paper, we investigate the phenomenon of catastrophic

    curriculum learningbenchmark
  686. arxiv:2609.33653 · cs.RO
    Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies
    Duo Wu, Haifeng Wang, Rongwei Lu, Jinghe Wang +5

    Reinforcement learning (RL) enables generalist robot policies to improve through trial-and-error interaction, yet its effectiveness is fundamentally constrained by sparse task rewards. Existing general-purpose reward models typically alleviate this issue by learning task progress from expert demonst

    manipulationlibero
  687. arxiv:2609.33650 · cs.LG
    FuseAlign: Forced Alignment in the Wild
    Mithilesh Vaidya, Stephen Bailey, Sumukh Badam, Matthew Bendel +1

    Word-level forced alignment estimates when each transcript word occurs in an audio recording. It underpins text-based media editing, subtitling, speech-data curation, and phonetic analysis. Existing evaluations understate the difficulty of forced alignment by relying on short, clean speech, perfect

    benchmark
  688. arxiv:2609.33647 · cs.RO
    InfraVLA: Extending Vision-Language-Action Navigation with Infrastructure Cameras
    Lukas Vierling, Benjamin Ramtoula, Luke Robinson, Ronald Clark +1

    Many indoor environments in which robots operate, such as warehouses, offices, and hospitals, already have cameras installed. They observe parts of the building that the robot cannot see from where it stands, yet navigation policies, including recent vision-language-action (VLA) models, do not use t

    vision-language-actionvlaquadruped
  689. arxiv:2609.33646 · cs.CV
    Probe to Act: Elevating Browser-Use Agent via Active Visual Probing
    Keliang Li, Heng Wang, Chen Hu, Daxin Jiang +2

    Browser-use agents require seamless alignment between structured web metadata and visual information, while preserving relevant context across long interactions. Existing interfaces often rely on either screenshot-level action prediction or static Set-of-Marks overlays, leaving the model to resolve

    agentbenchmark
  690. arxiv:2609.33642 · cs.AI
    Learning to Learn from Context: Synthetic Training from Perturbed Public Documents
    Haoyi Wu, Yang Xiao, Yusong Sun, Wenyang Hui +3

    Real-world tasks often require large language models (LLMs) to learn from complex task-specific context rather than pretrained parametric knowledge. This capability remains a weakness of LLMs, while human annotation for such task contexts is expensive and difficult to scale. Public high-quality docu

    long-context
  691. arxiv:2609.33640 · cs.LG
    SafeMol: Dual-Modality Safety Alignment for Molecular Multimodal Models
    Xinmiao Wang, Ruijie Wang, Menghui Wang, Jiawei Chen +4

    Molecular multimodal models support diverse understanding and generation tasks but may introduce safety vulnerabilities when handling hazardous molecules. In this work, We reveal substantial jailbreak vulnerabilities under both text-only and graph-conditioned settings. Our analysis further shows tha

    benchmark
  692. arxiv:2609.33639 · cs.AI
    Trajectory Unlearning on LLM-based Agents
    Yingdan Shi, Ren Wang

    Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an

    agentautonomous agentagenticbenchmark
  693. arxiv:2609.33637 · cs.RO
    Hierarchical Multi-agent Reinforcement Learning for Warehouse Robot Coordination under Communication Loss
    Weihao Sun, Gehui Xu, Andreas A. Malikopoulos

    In this paper, we propose a hierarchical multi-agent reinforcement learning framework for coordinating robot teams in warehouse environments under communication loss. We partition the robot team into groups, with centralized coordination within each group and distributed coordination across groups.

    multi-agent
  694. arxiv:2609.33634 · cs.AI
    Safety Reconstructed: Generative Modeling via Masked Diffusion Builds Strong Safety Guardrails
    Gert Lek, Abele Malan, Chaoyi Zhu, Pin-Yu Chen +2

    Guard models are the last line of defense between a language model and a harmful output, yet their training objective is surprisingly narrow. Existing guards learn to predict a single verdict token from a conversational context, concentrating supervision on a single target. The consequences are stru

    benchmark
  695. arxiv:2609.33630 · cs.LG
    lapanda: A Matrix-Free Differentiable Solver for Nonconvex Constrained Optimization Layers
    Yuankun Chen, Zifei Nie, Kangyu Lin, Ján Drgoňa +1

    Differentiable optimization brings the structural guarantees of mathematical optimization to network pipelines, allowing them to be trained end-to-end. However, its application remains challenging for nonconvex constrained problems, as existing differentiable solvers often suffer from limited modeli

    memorybenchmark
  696. arxiv:2609.33628 · cs.LG
    Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
    Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang +1

    Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such a

    curriculum learning
  697. arxiv:2609.33627 · cs.CV
    StoryEngine: A State-Grounded Agentic Framework for Video Storytelling
    Yingrui Wang, Zeqing Wang, Yeying Jin

    Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of

    agenticbenchmark
  698. arxiv:2609.33618 · cs.AI
    ParaAgent: Reinforcing Parallel Acting in Open-World Tool Environments
    Shengbin Yue, Hongru Wang, Siyuan Wang, Xiaoxin Chen +2

    Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them withou

    multi-agentbenchmark
  699. arxiv:2609.33609 · cs.LG
    You Only Edit Once: Incentivizing In-Context Capability of LLMs via Local Demonstration Refinement
    Jiarong Wen, Qi Wang, Yun Qu, Yixiu Mao +7

    In-context learning (ICL) is crucial for boosting the inference performance of large language models (LLMs). However, the effectiveness of ICL in LLMs is greatly influenced by the choice of demonstration sets. Exhaustive searches over these sets are combinatorial, and existing selectors often rely o

    benchmark
  700. arxiv:2609.33608 · cs.AI
    Learning Transferable Reaction Mechanisms from Visual Chemical Knowledge
    Yujian Yuan, Jiaxin Xu, Xin Cai, Yufan Chen +3

    Reaction mechanisms describe the step-by-step transformations underlying chemical reactions and are central to reaction analysis and synthesis. Learning-based models have achieved strong performance on established mechanism-prediction benchmarks, but transferring them to unseen chemistry remains cha

    benchmark
  701. arxiv:2609.33603 · cs.CV
    ViCoR: Reliable Molecular Structure Extraction via Spatially Aligned Verification and Executable Revision
    Yujian Yuan, Xin Cai, Yufan Chen, Jiaxin Xu +4

    Reliable optical chemical structure recognition (OCSR) is essential for building high-quality chemical data from scientific literature, yet even small recognition errors can propagate into chemical databases and downstream models. In practice, recognized structures often require manual inspection an

    benchmark
  702. arxiv:2609.33601 · cs.AI
    JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
    Kaicheng Yang, Kaisen Yang, Chunyu Liu, Xianglong Yan +7

    Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization.

    post-training
  703. arxiv:2609.33597 · cs.LG
    The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces
    Georgios Feretzakis, Alexandros Papaspyridis

    Side-channel evaluators routinely inspect leakage tests while acquisition is still running, and extend or stop the campaign based on what they see. Fixed-horizon screening such as the Welch $t$-test with threshold $|t|>4.5$ gives no error guarantee for this monitored decision rule. We study anytime-

    evaluator
  704. arxiv:2609.33595 · cs.RO
    Beyond One-Step Accuracy: State-Affine Latent Transition for Reliable Visual Planning
    Boyuan Zhang, Yingjun Du, Xiantong Zhen, Ling Shao

    Joint-embedding world models enable visual planning by learning action-conditioned dynamics in latent space. Yet they are commonly trained for one-step prediction on encoded states, while planning recursively applies the learned transition to its own predictions. One-step accuracy therefore does not

    world modelaction-conditioned
  705. arxiv:2609.33593 · cs.CV
    LoopLUT: 3D Lookup Tables with Progressive Region Refinement for Real-Time 4K Image Enhancement
    Yang Ye, Jiajun Ma, Chen Wu, Wei Wang +4

    Color enhancement of 4K images must meet a quality target under a tight compute budget. Three-dimensional lookup tables (3D LUTs) dominate real-time enhancement because they decide at low resolution and apply a per-pixel lookup at full resolution. A single global LUT, however, is spatially invariant

    benchmark
  706. arxiv:2609.33589 · cs.LG
    TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs
    Zihan Lin, Xiaohan Wang, Jie Cao, Jiajun Chai +3

    Efficient exploration often remains a central bottleneck in reinforcement learning with verifiable rewards (RLVR). Although temperature control and test-time scaling strategies can increase rollout diversity of large language models (LLMs), they either expand the sample budget at rollout time or lea

    benchmark
  707. arxiv:2609.33582 · cs.CV
    Fill2SR: Repurposing Inpainting Diffusion Transformers for Real-World Super-Resolution
    Xingfu Yi, Xiaoxue Yu

    Recent real-world image super-resolution (SR) methods often adapt text-to-image (T2I) backbones with ControlNet-style branches or spatial conditioning tokens, which increases memory and computes with resolution and often constrains training to a fixed scale. We propose Fill2SR, which repurposes a ma

    memorybenchmark
  708. arxiv:2609.33581 · cs.CV
    ForeFly: A Dual-Horizon World Action Model for Aerial Vision-Language Navigation
    Kunhui Wang, Xintong Zhang, Junyu Gao, Changsheng Xu

    Aerial Vision-Language Navigation (AVLN) requires UAVs to maintain reliable instruction following over long trajectories in complex 3D environments. However, existing AVLN approaches are predominantly reactive or limited to single-horizon prediction, overlooking complementary future cues across diff

    benchmark
  709. arxiv:2609.33579 · cs.AI
    OpenFC: Learning Verification Policies towards Open-Search Fact Checking
    Xinming Wang, Kaixiang Qiu, Yansong Lin, Chunji Lv +4

    Open-search fact checking is not merely retrieval followed by classification, but a sequential decision problem in which every query, source visit, and stopping decision reshapes the evidence available for verification. Yet existing systems often distribute these decisions across predefined pipeline

    tool-usebenchmark
  710. arxiv:2609.33575 · cs.RO
    SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models
    Tianfu Li, Haoxuan Xu, Wenbo Chen, Haitian Li +6

    Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often

    vision-language-actionvlavla modelmanipulationworld modelaction-conditioned
  711. arxiv:2609.33570 · cs.LG
    HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
    Xinrui Chen, Mengyang Li, Ou Wu, Ji Zhang

    Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existi

    memory
  712. arxiv:2609.33566 · cs.LG
    Correct then Forecast: Observer State-Space Models for Time Series Forecasting
    Alexis-Raja Brachet, Guillaume Clavier--Frémond, Abdelhakim Ziani, Pierre-Yves Richard +1

    Time series forecasting requires extrapolating the dynamics of an observed process beyond the last available measurement. Yet recurrent forecasting models typically treat observations as inputs that directly control their latent dynamics. It leads to a regime change when these observations become un

    latent dynamicsbenchmark
  713. arxiv:2609.33565 · cs.AI
    Dr. Free: You Don't Need Difficulty Rewards for Self-Evolving Search Agents
    Zhipeng Qian, Zihan Liang, Yufei Ma, Jie Ma +9

    A central limitation of current data-free self-evolution methods for training search agents is their reliance on difficulty-based proposer rewards. These methods reward a proposer for generating questions that challenge a co-evolving solver, using solver difficulty as a proxy for question quality. Y

    knowledge graphself-evolvingbenchmark
  714. arxiv:2609.33563 · cs.LG
    MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
    Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt

    World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal f

    world modelaction-conditionedmulti-agent
  715. arxiv:2609.33558 · cs.CV
    PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization
    Liang Peng, Shizhuo Mu, Bohan Tan, Wenyuan Wang +5

    3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object a

    memorybenchmark
  716. arxiv:2609.33551 · cs.RO
    FoLD: Force-Informed Learning for Dexterous Articulated Object Manipulation
    Haowei Shen, Tingai Li, Yumeng Liu, Wenyuan Guang +5

    Transferring human demonstrations to dexterous robots remains challenging because differences in hand morphology and contact dynamics often cause retargeted motions to fail at producing the intended object behavior. We present \textbf{FoLD}, a framework for learning dexterous manipulation of articul

    manipulationdexterousbenchmark
  717. arxiv:2609.33550 · cs.LG
    Sparsity by Default: The Theory and Practice of ARD in Gaussian Process Regression for Variable Selection
    Jia Cai

    Automatic relevance determination (ARD) is the standard device for input selection in Gaussian process (GP) regression. By giving the covariance kernel a separate lengthscale for every input and learning those lengthscales by maximizing the marginal likelihood, ARD lets the data decide which coordin

    embodied
  718. arxiv:2609.33548 · cs.LG
    TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge
    Houssem Sifaou, Prabodh Katti, Bipin Rajendran, Osvaldo Simeone

    Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family o

    memory
  719. arxiv:2609.33546 · cs.RO
    Steer2Grasp: Inference-Time Embodiment-Aware Steering for Diverse Physically Feasible Grasp Diffusion
    Vignesh Vembar, Ayush Kaura, A Padmaprabhan, Siddharth Sinha +4

    Current grasp diffusion models provide rich priors for generation, yet their object-centric approach can violate the kinematic and collision constraints imposed by the embodiment and the environment. Existing embodiment-aware methods primarily perform local corrections around generated grasps throug

    grippergrasp
  720. arxiv:2609.33536 · cs.LG
    Does Execution Require Target KV Fidelity? A Mixed-Fidelity KV Runtime for LLM Serving
    Jiantong Jiang, Yue Yang, Peiyu Yang, Feng Liu

    Large language model (LLM) serving is increasingly constrained by the GPU memory consumed by key-value (KV) caches. Existing compression, eviction, and offloading techniques alleviate this pressure, but serving runtimes typically treat only the configured target KV representation as execution-ready.

    memory
  721. arxiv:2609.33534 · cs.CL
    ManiEdit: Sequential Unstructured Knowledge Editing for Language Models from a Manifold Perspective
    Rui Liu, Chenheng Zhang, Haoxuan Li, Zhouchen Lin

    Large language models (LLMs) inevitably generate some incorrect or outdated content, necessitating efficient and precise mechanisms for continual knowledge updates. However, existing model editing methods struggle to sequentially edit unstructured long-form knowledge, suffering from severe edit forg

    benchmark
  722. arxiv:2609.33530 · cs.LG
    E-CONAN (Entailment, CONtradition And Neutral) Diagnostics Dataset Investigating Linguistic Phenomena in Arabic Natural Language Understanding
    Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi

    Natural Language Understanding (NLU) plays a crucial role in various applications, yet its performance suffers from weaknesses in handling the complexities of human languages, ranging from lexical ambiguity to high-level reasoning difficulties. Analyzing errors across diverse linguistic phenomena is

    benchmark
  723. arxiv:2609.33524 · cs.AI
    EverMine: Dissecting the Self-Evolution of Research Capabilities in Long-Horizon Alpha Research
    Siyuan Li, Jiangfeng Zhang, Rui Yao, Weihua Qiu +2

    Self-evolving agents aim to turn research feedback into reusable skills, tools, and research rules. Whether these accumulated capabilities continue to improve later research requires controlled evaluation. Long-horizon alpha discovery provides a state-dependent setting: once a new factor enters the

    self-evolving
  724. arxiv:2609.33521 · cs.LG
    LLM4Trust: Exploring the Capabilities of Large Language Models for Trust Evaluation
    Jie Wang, Yanbo Sun, Zheng Yan, Jiahe Lan +1

    Trust evaluation plays a critical role in cybersecurity by supporting risk mitigation and decision-making. A variety of trust evaluation methods have been proposed, with learning-based approaches offering high accuracy and automation. However, they often require substantial ground truth, suffer from

    benchmark
  725. arxiv:2609.33517 · cs.MA
    TRACE: Governing Memory Validity in Evolving Multi-Agent Systems
    Wenjun Xiong, Shengtao Zhang, Shangding Gu, Bo Tang +4

    Persistent memory lets language-model agents carry information across long-running collaborations, but leaves a lifecycle question open: what may a returning agent still act on once the shared state has changed? A memory can be correctly retrieved, relevant to the current task, and faithful to its s

    memorypersistent memoryagentmulti-agentagent system
  726. arxiv:2609.33509 · cs.AI
    What Happens During Autonomous Deep Research After the User Steps Away?
    Yimin Liu, Yijia Zhang, Yanmin Li, Tangwen Luo +3

    In autonomous deep research, a user provides a task and relevant background, then leaves the agent to conduct an extended investigation without further human intervention. We study how this initial user information is reflected in intermediate actions and how these actions relate to final recommenda

    agentevaluatorevaluation framework
  727. arxiv:2609.33505 · cs.AI
    When Evidence Changes the Subject: Subject-Typed Claim Licensing for Learned Routing
    Jian Chen, Zixuan Yuan

    Modern learned systems increasingly combine learned components with search, repair, or external solvers. Benchmarks often measure the resulting end-to-end system, while scientific claims may concern only one component, creating an attribution problem: evidence can fail to support the requested compo

    benchmark
  728. arxiv:2609.33503 · cs.AI
    RelaxKV: Recomputation Guided by the Query with Sparse Context Attention for Efficient KV Cache Reuse
    Ruoling Qi, Yirui Liu, Xuaner Wu, Yuxin Jin +4

    Cross-request KV caching reduces the prefill cost of Retrieval-Augmented Generation (RAG), but conventional prefix caching severely limits cache reuse across requests. Position-Independent Caching (PIC) removes this constraint by reusing independent chunks, but their KV states miss cross-chunk inter

    retrieval-augmented
  729. arxiv:2609.33497 · cs.LG
    Hamiltonian JEPA: Action-Conditioned World Models with an Inherited Control State
    Tamim Zoabi, Ameen Ali, Lior Wolf

    Planning from pixels needs more than a latent space that is stable and predictable. The state the planner scores must also be organized by how actions move the system. Joint-embedding predictive architectures (JEPAs) avoid pixel reconstruction by predicting future representations, but existing actio

    world modelaction-conditionedbenchmark
  730. arxiv:2609.33496 · cs.LG
    Chameleon: Dynamic Format Adapter for Efficient Diffusion
    Arnab Sanyal, Sandeep Chinchali

    Post-training quantization (PTQ) is the standard way to run modern diffusion models on memory-constrained accelerators, yet every existing diffusion PTQ scheme fixes the $\mathit{number\ format}$ in advance and only tunes the scale, zero point, or per-layer bit-width. At a fixed bit-width the best f

    post-training
  731. arxiv:2609.33495 · cs.CL
    LLMs Trust Their Own: Identity-Dependent Conformity in Multi-Agent Systems
    Liron Soffer, Ravid Shwartz-Ziv, Chen Shani

    Large language models (LLMs) are increasingly deployed in multi-agent settings, where agents observe and influence one another, making social influence a key dimension of AI behavior and safety. We investigate whether LLMs' responses depend on the social identity of other agents, beyond the effect o

    manipulationmulti-agentagent system
  732. arxiv:2609.33492 · cs.AI
    Federated Multi-Modal Human Activity Recognition using Multi-Agent Reinforcement Learning
    Debasmita Dey, Tanmay Sen, Himel Mallick

    Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences i

    agentmulti-agent
  733. arxiv:2609.33485 · cs.CL
    DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
    Guanzheng Chen, Viet Dac Lai, Subhojyoti Mukherjee, Branislav Kveton +4

    While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural entanglement in monolithic architectures, where the massive search burden of contextual grounding exha

    long-contextlong context
  734. arxiv:2609.33484 · cs.RO
    AMBIT: Anticipatory Multimodal Body Recruitment for Bimanual Tracking on a Humanoid
    Hanlong Li, Sihan Tan, Takeshi Ashizawa, Benjamin Yen +1

    A humanoid with 5-DoF arms cannot track generic bimanual end-effector trajectories with its arms alone; pelvis and waist motion must be recruited, but which motion, and when, is not uniquely determined. On a Unitree R1 in fixed double support, the set of dynamically valid recruitment strategies (pel

    humanoid
  735. arxiv:2609.33482 · cs.LG
    How Synthetic Labels Improve Conformal Prediction: A Perspective on Conditional Coverage
    Qianyi Chen, Bo Li

    Conformal prediction provides distribution-free finite-sample marginal coverage, but post-hoc calibration data may be too scarce to learn how uncertainty varies across inputs. Meanwhile, abundant covariates can often be labeled cheaply by domain models or general-purpose language models. We study wh

    benchmark
  736. arxiv:2609.33477 · cs.AI
    Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
    Yirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin +5

    Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. Full-attention KV caches are token-addressable, whereas linear-attention layers maintain recurrent states that cannot be rolled back to arbi

    long-context
  737. arxiv:2609.33470 · cs.AI
    LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs
    Haochen Luo, Yifan Li, Binh Minh An, Xiaolong Luo +3

    Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamental

    llm agentmulti-agentagent systemevaluation framework
  738. arxiv:2609.33467 · cs.LG
    A Cheap Verifier is Good Enough: LLM Post-training is Robust to Erroneous Rewards
    Andreas Plesner, Curtis Northcutt, Francisco Guzmán, Anish Athalye

    When post-training large language models on tasks with semi-verifiable rewards, there are many factors (training steps, base model size, training order, data quality, verifier accuracy, etc.) that practitioners must contend with to maximize model performance. Yet, it remains unclear how well verifie

    post-training
  739. arxiv:2609.33464 · cs.RO
    VIDEAS: Distilling Explicit Action Semantics from Demonstration Videos for World Models via Prior-Guided Simulation
    Jianan Wang, Haoquan Zhai, Siyang Zhang, Bin Li +6

    World models learn internal representations of environment dynamics to predict future states, enabling agents to optimize action plans without physical interactions. However, developing world models that genuinely internalize underlying causal physical laws to explicitly reason about action precondi

    embodiedworld model
  740. arxiv:2609.33463 · cs.CL
    Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
    Cunchun Li, Haonan He, Yifan Gao, Minglei Li +3

    Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-rewe

    benchmark
  741. arxiv:2609.33462 · cs.CV
    SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera
    Shriram Damodaran, Soumyaratna Debnath, Cheston Tan, Lin Wang

    Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images

    embodiedai agentbenchmark
  742. arxiv:2609.33460 · physics.app-ph
    Bare-Die Antiferromagnetic Computing
    Yu Liu, Zhuoting Han, Zexin Feng, Peixin Qin +14

    Semiconductor electronic devices are increasingly constrained by fundamental quantum tunneling effects and charge-based mechanisms, which severely limit further miniaturization, write-speed scaling, and environmental robustness of silicon-based technologies. These limitations are particularly prohib

    memory
  743. arxiv:2609.33457 · cs.LG
    A Free Knob: Decoupling Calibration and Predictive Skill in Threshold-Based Evaluation
    Md Tanveer Hossain Munim, Bijoy Ahmed Saiem, Al-Amin Sany, Tanzima Hashem

    Many dense-prediction benchmarks evaluate rare events by pooling prediction and target over spatial blocks, thresholding each, and scoring the contingency table. At a fixed rare operating point, the max-pooled Critical Success Index (CSI) confounds spatial discrimination with amplitude calibration:

    benchmark
  744. arxiv:2609.33455 · cs.AI
    What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation
    Zzizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo

    On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted

    benchmark
  745. arxiv:2609.33451 · eess.SY
    DAOCP: a dual active set solver for optimal control problems
    Alberto Zaupa, Samuel Erickson, Mikael Johansson

    We present DAOCP, a dual active set solver for linear quadratic optimal control problems with stage-wise equality and inequality constraints. Active set methods are leading Model Predictive Control benchmarks for full-body robotics, but existing solvers operate on dense QPs, while typical problem di

    benchmark
  746. arxiv:2609.33446 · cs.AI
    HESP: Separating What to Probe from When to Stop in Local LLM Alert-Triage Agents
    Zhuowen Liu, Zhixuan Wang

    Security operations centers receive far more alerts than analysts can investigate, and organizations that cannot send their telemetry to hosted models must automate triage with small open-weight LLMs on their own hardware. Current LLM agents leave the investigation procedure to the model, and small

    llm agent
  747. arxiv:2609.33443 · cs.CL
    Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
    Seonghyeon Go, Yongwoo Kim, Hyeonjin Cha, Jaeho Shin

    Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving the

    benchmark
  748. arxiv:2609.33441 · cs.CL
    MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation
    Ruihong Zeng, Jonathan Tonglet, Preslav Nakov, Iryna Gurevych

    Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what

    benchmark
  749. arxiv:2609.33440 · cs.AI
    MAC-Net: A Multi-Task Deep Learning Framework for Modeling Cognitive Function From Task-Based fMRI
    Md. Tanvir Rahman, Nabil Anan Orka, Asaduzzaman Khan, Mohammad Ali Moni

    Objective cognitive assessment from neural signals supports neurorehabilitation, but individual-level prediction from task-based fMRI (tfMRI) remains difficult because neural features coexist with substantial demographic and scanner-related variation. We present the Multi-task Activation and Contras

    benchmark
  750. arxiv:2609.33439 · cs.AI
    Raven: The Harness of Harnesses for Composable Agentic Intelligence
    EverMind AI

    As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits t

    agentai agentmulti-agentagenticagent system
  751. arxiv:2609.33436 · cs.LG
    SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval
    Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung

    Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture comb

    memorymemory architecturepersistent state
  752. arxiv:2609.33430 · cs.AI
    APEX: An Extensible Model for Agent-Assisted Production Scheduling
    Felix J. Grumbach, Stefan Görlitz

    Production scheduling requires realistic models that reflect operational constraints and efficient methods that balance competing goals. Putting these methods into use also requires data integration, model adaptation and specialist expertise. We present APEX, an extensible production scheduling fram

    agentbenchmark
  753. arxiv:2609.33429 · cs.AI
    Graph-Guided Repository Environment Construction
    Jianying Pan, John Zhang, Hongyu Zhang

    Coding agents now increasingly rely on execution to validate their solutions, making the construction of reliable execution environments a critical enabling capability. However, repository environment construction is challenging because execution requirements are fragmented across repository artifac

    benchmark
  754. arxiv:2609.33426 · cs.CL
    TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment
    Zhenyu Lei, Zihan Chen, Yaochen Zhu, Shangbin Feng +4

    Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out cha

    benchmark
  755. arxiv:2609.33419 · cs.CV
    TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
    Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang +2

    Comparisons in video self-supervised learning often evaluate complete training recipes rather than isolating the method itself: architecture, objective, data exposure, schedule, scale, and decoder capacity can all vary at once. This makes it hard to identify which choices yield motion-prioritized re

    v-jepa
  756. arxiv:2609.33417 · cs.CL
    Grounding Memory Summarization in Utility Intent
    Zhenyu Lei, Mingjia Shi, Xingbo Fu, Haoyu He +2

    Existing summarizers for memory systems are typically optimized for human-facing criteria such as faithfulness, which misaligns with their true objective: preserving the evidence needed to support future queries. We show that conditioning summarization on query-answer pairs substantially improves an

    memory
  757. arxiv:2609.33414 · cs.CV
    TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models
    Shuning Wang, Zhiheng Wu, Xun Zhou, Chongyang Cui +7

    Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anc

    benchmark
  758. arxiv:2609.33412 · cs.CV
    Resolving State-Representation Mismatch: State-Space Visual Reasoning for Open-Loop VLA Planning
    Junhao Xiao, Haoxiang Zhao, Menghao Fang, Jinkui Zhang +7

    Despite rapid progress in vision-language-action (VLA) models, existing reasoning paradigms still face a fundamental \emph{state-representation mismatch} in open-loop planning. Given only an initial observation, models must internally simulate action-conditioned state transitions, whereas text-, pix

    vision-language-actionvlaaction-conditioned
  759. arxiv:2609.33411 · cs.AI
    MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
    Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma +2

    Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving be

    benchmarkevaluator
  760. arxiv:2609.33410 · cs.LG
    FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward
    Sriman Achanta

    Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization reference as it scans keys. Earlier contributions therefore remain provisional and may require resca

    long-contextlong context
  761. arxiv:2609.33409 · cs.CL
    Dense Is Not Enough: Hierarchical Supervision Allocation for Long-Horizon On-Policy Distillation
    Yuhao Sun, Binrui Wu, Zhuoer Xu, Ming Wen +4

    On-policy distillation (OPD) transfers the capabilities of a large language model to a smaller student by providing teacher supervision on the student's own rollouts. In long-horizon agentic tasks, however, uniform token-level matching can allocate supervision poorly: a large local discrepancy need

    agentagenticagent benchmarkbenchmark
  762. arxiv:2609.33408 · cs.LG
    StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks
    Mustafa Bora Çelik, Ceren Çelik, Orhan Gazi

    In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90\% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses toward maximum uniform entropy ($H=2.584$ bits) as sparsity increases, failing to capture long-range gai

    persistent state
  763. arxiv:2609.33407 · cs.LG
    Let CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure Priors
    Emma Lei Hovmand, Jonas Elsborg, Melih Kandemir, Arghya Bhowmik

    De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical prior that should be improved by likeli

    evaluator
  764. arxiv:2609.33403 · cs.AI
    DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
    Yupeng Xie, Zhenyang Wang, Liangwei Wang, Jiayi Zhu +2

    Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack na

    multi-agent
  765. arxiv:2609.33402 · cs.CV
    VaME: Exploring Variational Latent Reasoning for Multimodal Embeddings
    Peixi Wu, Mingzhou Jiang, Feipeng Ma, Biao Yang +10

    Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing appr

    benchmark
  766. arxiv:2609.33401 · cs.AI
    Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
    Yixuan Liu

    Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models expose typed decisions with probabilities that software can use to allow, block, or escalate inputs, but whether these probabilities support reliab

    agent
  767. arxiv:2609.33399 · cs.CV
    SciGen-Verifier: A Multimodal Reasoner for Explainable Verification in Scientific Image Generation
    Jiali Chen, Zhengteng Lin, Zuqi Wang, Shirong Lin +5

    In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet ve

    benchmark
  768. arxiv:2609.33398 · cs.AI
    COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement
    Siwei Chen, Xinping Bao, Xinyu Cai, Yuan Cao +2

    Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the ext

    self-improvementonline learning
  769. arxiv:2609.33397 · cs.AI
    CoViST: Visual Token Compression via Composable States
    Qi Zhang, Xiandong Meng, Ronggang Wang, Siwei Ma

    Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. T

    benchmark
  770. arxiv:2609.33392 · cs.LG
    AutoHGNN: Robust and Efficient Neural Architecture Search for Hypergraph Neural Networks
    Sirui Li, Pietro Liò b, Xinsheng Li, Baisong Liu +1

    Hypergraph neural networks have achieved significant success in recent years. However, manual architecture crafting is labor-intensive and often fails to capture complex, higher-order relations, making the automation of hypergraph neural network structure design crucial. To improve the automation an

    benchmark
  771. arxiv:2609.33385 · cs.CL
    OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading
    Jingyuan Xiao, Jiayue Wang, Yitao Hu, Xinning Wang +6

    Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a large expert parameter footprint that exceeds memory-constrained GPU capacity. Expert offloading is a na

    memory
  772. arxiv:2609.33384 · cs.CV
    PulseQuant: Propagation-Guided Subspace Correction for 4-Bit Video Diffusion Transformers
    Yutong Wang, Xingtong Ge, Enhuai Liu, Yunke Wang +3

    Quantization errors in video diffusion transformers can be amplified or attenuated by subsequent denoising updates, making local reconstruction error an incomplete predictor of final impact. We introduce PulseQuant, a 4-bit post-training quantization method that combines trajectory sensitivity with

    post-training
  773. arxiv:2609.33379 · cs.AI
    NLPG: Natural-Language Policy Gradients for Self-Evolving Language Agents
    Xu Liu, WenZhang Wei, Jun Cao, Dehua Peng +3

    Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typically rely on scalar rewards or repeatedly m

    agenttool useself-evolvingbenchmark
  774. arxiv:2609.33378 · cs.RO
    Recursive Harness Distillation across Agents for Robot Manipulation
    Seungyeon Kim, Junhoo Lee, Minkyu Kim, Baekseung Kim +1

    A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventi

    vision-language-actionmanipulationgr00tagent
  775. arxiv:2609.33377 · cs.LG
    Optimal Transport Dropout for Structured Predictive Uncertainty
    Giacomo Lorenzon, Francesco Regazzoni

    Deterministic neural networks and neural operators provide point predictions with no intrinsic measure of reliability. Yet, predictive uncertainty may stem from irreducible outcome variability, finite data, or limitations of the chosen model class. Monte Carlo dropout offers a computationally conven

    benchmark
  776. arxiv:2609.33371 · cs.LG
    API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary
    Patrick Kenney, Hadi Ahmadi, Denis Lusson, Donald Nguyen +1

    Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where it may persist in conv

    memoryagentllm agentagenticagent system
  777. arxiv:2609.33370 · cs.LG
    The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection
    Farhan Shahriyar Hossain, Taufikur Rahman Fuad, Md Abrar Jahin, Md Rizwan Parvez

    Open-set graph anomaly detection trains on a few labeled anomalies from one class and must also find anomaly classes that were never labeled. Published results share three conventions: the test score is read at the best epoch on the test set, baseline numbers are copied from earlier papers, and most

    benchmark
  778. arxiv:2609.33362 · cs.CL
    From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS
    Kangxiang Xia, Xinfa Zhu, HangRui Hu, Kexin Huang +7

    Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation,

    agenticiterative refinementbenchmark
  779. arxiv:2609.33359 · cs.CV
    When Does Geometric View Synthesis Help Wine Label Retrieval? A Public One-Shot Benchmark Across Self-Supervised and Vision-Language Backbones
    Yueh-Cheng Huang

    Geometric view synthesis can expand a single wine-label photograph into a training set, but its value with pretrained image encoders is unclear. We study this on a public WineSensed-derived benchmark of 1,000 classes, one enrollment photograph per class, and 4,295 real queries. With the earlier DINO

    benchmark
  780. arxiv:2609.33357 · cs.AI
    DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?
    Nan Huang, Mario Tapia-Pacheco, Kun Zhou, Yiming Huang +3

    Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical ta

    ai agentbenchmark
  781. arxiv:2609.33356 · cs.AI
    Long-Horizon Analog Design Bench: Benchmarking Agents on Hours-Long Analog and Mixed-Signal Circuit Design Tasks
    Analog Design Bench Team

    Coding agents now sustain hours-long, tool-driven loops, yet their ability to carry long-horizon analog and mixed-signal circuits to electrical specification remains unmeasured. We introduce Analog Design Bench, a long-horizon agentic benchmark of 50 transistor-level design tasks contributed by 17 c

    agentagenticbenchmark
  782. arxiv:2609.33354 · cs.RO
    Traceable Human-to-Humanoid Sign Language Benchmarking
    Ao Liu, Shengeng Tang, Lechao Cheng, Yanbin Hao +2

    Sign data collection is costly, and teleoperation scales poorly, motivating reuse of large video corpora. Humanoid signing requires converting video-derived human motion into robot trajectories while preserving linguistic motion cues. Errors from fitting, human-motion repair, retargeting, robot geom

    humanoidteleoperationbenchmark
  783. arxiv:2609.33352 · cs.LG
    How Much Imprecision is Enough Imprecision in my Classifier? A Practical Elicitation Procedure
    Victor F. Lopes de Souza, Sébastien Destercke, Abdelhak Imoussaten

    Set-valued classifiers, whether derived from precise probabilities and an adapted cost function, from convex sets with a robust inference mechanism, or from conformal methods, are routine options to obtain more robust, trustworthy predictions. However, there is a lack of operational tools to measure

    benchmark
  784. arxiv:2609.33351 · cs.LG
    QuPID: Quantum Parameter-Efficient Input-Dependent Retrieval Adaptation for Medical RAG
    Hyojun Ahn, Emily Jimin Roh, Soohyun Park, Walid Saad +2

    Fidelity-based quantum retrieval ranks candidates by the fidelity between query and archive states. Applying a shared input-independent unitary after fixed state encoding leaves that fidelity unchanged, so training the circuit cannot alter the ranking. Quantum parameter-efficient input-dependent ret

    ragbenchmark
  785. arxiv:2609.33350 · cs.LG
    KoopCell: Koopman-Based Generative Model for Learning Single-Cell Dynamics from Distribution Snapshots
    Wanfeng Lu, Yutong Zhang, Keyi Zhou, Chenxin Ge +2

    Learning population dynamics from temporally sparse, unpaired distribution snapshots is a fundamental challenge in developmental biology. Recent approaches based on neural differential equations and flow matching can interpolate between observed population snapshots, but may struggle to extrapolate

    latent dynamicsmemory
  786. arxiv:2609.33347 · cs.LG
    MultiEcho: An Experimental Science of Learned Worlds
    Meng Zhu, Airui Zhang

    World models can be studied as experimental systems with response laws of their own. We introduce MultiEcho, a framework for estimating these laws through controlled counterfactual interventions, delimiting their applicability, and separately testing their physical correspondence. Across nine simula

    world model
  787. arxiv:2609.33342 · cs.LG
    CalibHyper: Chance-Corrected Relational Hypergraphs for Few-Shot Molecular Property Prediction
    Linyu Li, Zhi Jin, Yuanpeng He, Dongming Jin +6

    Molecular property prediction is central to drug development and materials discovery, but experiments are costly and labeled data are scarce. Context-aware methods use auxiliary assay labels to support few-shot prediction, and recent work supervises property relations with label agreement. However,

    benchmark
  788. arxiv:2609.33338 · cs.CV
    OPERA: A Unified Omnimodal Progressive Spatio-Temporal Reasoning Agent for Referring Video Segmentation
    Jingchen Ni, Yuji Wang, Shannan Yan, Haoru Li +2

    Referring video segmentation with heterogeneous multimodal queries---spanning text, audio, and reference images---demands both robust cross-modal understanding and precise spatio-temporal reasoning. We propose OPERA (Omnimodal Progressive spatio-tEmporal Reasoning Agent), a unified reasoning agent b

    agentbenchmark
  789. arxiv:2609.33337 · cs.RO
    Safe Score Matching: Diffusion Policies with Hamilton-Jacobi Reachability for Online Safe Reinforcement Learning
    Boyang Li, Matthew Kim, Sylvia Lee Herbert

    Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates tha

    benchmark
  790. arxiv:2609.33336 · cs.LG
    Beyond Conservatism: Recoverability-Conditioned Exploration for Model-Based Imitation Learning
    Xuanlin Chen, Ziyue Wang, Xunlan Zhou, Yuan-yih Shang +2

    Model-based imitation learning (MBIL) improves real-environment interaction efficiency by optimizing policies on imagined rollouts from a learned world model. However, the gap between model-induced and real-environment occupancies makes policy learning sensitive to model error. Conservative MBIL mit

    manipulationworld model
  791. arxiv:2609.33335 · cs.AI
    Does Learning to Predict the World Help Agents Act? Auditing World-Model Post-Training
    Xinyu Che, Hang Yan, Yanchen Liu, Haochen Liu +4

    Predicting how an environment will change before acting is a natural route to better decision making for agents. Recent post-training methods therefore require agents to predict the next observation and turn that prediction into a reward or a direct supervision signal, which is called world model. E

    world modelagentpost-trainingworld-model post-training
  792. arxiv:2609.33326 · cs.AI
    ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces
    Jerry Wang, Haibo Jin, Xiaopeng Yuan, Peng Kuang +1

    Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what

    long-contextmulti-agentagent system
  793. arxiv:2609.33325 · cs.CV
    VisionHOPE: Visual Backbones as Self-Modifying Learning Systems
    Siran Peng, Tianshuo Zhang, Tianyu Fu, Weisong Zhao +7

    Visual backbones have evolved from Convolutional Neural Networks (CNNs) with local aggregation to Vision Transformers (ViTs) with global interactions, State-Space Models (SSMs) with input-dependent state transitions, and Test-Time Training (TTT) layers that adapt an inner learner while processing an

    memory
  794. arxiv:2609.33323 · cs.LG
    Agentic Multi-Turn Reasoning: A Fairness Approach
    Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill, Jackson Cothren +2

    Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit

    memoryagentictool usebenchmark
  795. arxiv:2609.33322 · cs.AI
    Robust Hierarchical Structures for Agentic Document Analysis
    Ruiying Ma, Yiming Lin, Aditya G. Parameswaran

    Large Language Models (LLMs) enable us to better understand text documents, including PDFs and Word documents. However, LLMs, as well as more modern LLM agents, i.e., those with tool-calling abilities, typically treat such documents as plain text, ignoring the fact that they are often organized hier

    llm agentagentic
  796. arxiv:2609.33319 · cs.AI
    PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning
    Kecheng Liang, Haoyang Liu, Zexin Chen, Zirong Liu +4

    A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the under

    benchmark
  797. arxiv:2609.33314 · cs.LG
    ZeroGAR: Benchmarking the Adversarial Robustness of Zero-Shot Graph Models
    Zhongjian Zhang, Xiao Wang, Busheng Zhang, Bo Yan +4

    Zero-shot graph models (ZGMs), which learn transferable knowledge from source graphs and directly apply to unseen target graphs without any adaptation, have achieved promising performance and attracted considerable attention. Despite their proliferation, existing ZGMs are predominantly evaluated on

    manipulationbenchmark
  798. arxiv:2609.33311 · cs.RO
    SocialHumanoid: Towards Expressive Humanoid Behavior via One-Step Co-Speech Motion Generation
    Chengqun Yang, Tengjie Zhu, Liang Xu, Fulong Liu +9

    Humanoid robots are increasingly expected to serve as embodied social agents that communicate naturally with humans through face-to-face interaction. During such communication, humanoid robots require body behaviors that are synchronized with speech, affectively expressive, and suitable for real-tim

    embodiedhumanoidwhole-body control
  799. arxiv:2609.33310 · cs.RO
    CompliantWBC: Whole-Body Compliance for Heavy Humanoids via Force Latent Estimation and Residual Impedance Targets
    Tan-Dzung Do, Cuc T. Trinh, Tuan Dat Phuong, Chien Le +3

    Whole-body compliant control is essential for deploying heavy humanoids under high payload in human-centric environments. Most prior force-aware learning-based pipelines focus on end-effector resistance, per-link upper-body springs, or end-effector stiffness modulation, leaving arbitrary-site pertur

    humanoid
  800. arxiv:2609.33306 · cs.CV
    LoopTrack: A Simple Baseline for Parameter-Efficient Transformer Tracking
    Liang Peng, Chenxiao Li, Libo Zhang, Xingping Dong +1

    Current Transformer-based tracking methods typically stack multiple Transformer blocks with separate parameters to model interactions between the target template and the search region for target localization. These trackers often incur substantial parameter overhead from stacked blocks, making their

    memory
  801. arxiv:2609.33304 · cs.CV
    Relevance Does Not Imply Applicability: Experience Activation for Personal GUI Agents
    Fuyao Zhang, Xuan Wang, Zherui Li, Jiaming Zhang +2

    Personal Graphical User Interface (GUI) agents rely on interaction history to infer what a user wants from ambiguous instructions and to anticipate recurring routines. Existing approaches retrieve task-relevant history and append it to the policy's context, implicitly assuming that experience releva

    agent
  802. arxiv:2609.33303 · cs.LG
    BITS: Rethinking Fair and Comprehensive Evaluation for Irregular Time Series Forecasting
    Kangjia Yan, Linfeng Wang, Tianen Shen, Xiangfei Qiu +6

    Despite recent progress in irregular time series forecasting, the field still lacks a unified benchmark for fair and comprehensive evaluation. Existing evaluations are often conducted on a limited set of datasets with inconsistent experimental protocols and predominantly error-based metrics, renderi

    benchmark
  803. arxiv:2609.33301 · cs.AI
    Hesitation-Aware On-Policy Distillation for Diffusion Language Models
    Jianguo Huang, Lipeng Wan, Yanchen Deng, Bo An

    Diffusion large language models (dLLMs) generate text by iterative unmasking. At each denoising step, a dLLM proposes a token at every masked position, but the decoder commits only a confident subset of these proposals. Trace-based on-policy distillation (TOPD) builds on this process by matching the

    benchmark
  804. arxiv:2609.33299 · cs.RO
    AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents
    Cunhao Zhu, Yifeng Wang, Dongliang Xu, Yunzhong Hou +2

    World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as

    embodiedaction-conditionedembodied agentbenchmark
  805. arxiv:2609.33298 · cs.LG
    Direct Hidden-State Alignment: Mapping and Controlling Preference Expression in LLMs
    Fansheng Zhang, Shengran Guo, Zexiao Wang, Liang Yuan +2

    In many settings, post-training need not create the target behavior from scratch: the base model can already produce it, but not reliably. This shifts part of preference alignment from capability acquisition to behavioral expression. We ask how a specified preference is represented in native model c

    post-training
  806. arxiv:2609.33296 · cs.CL
    BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages
    Priyanka Dasari, Yuvrajsinh D. Bodana, Vandan Mujadia, Arafat Ahsan +2

    Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks foc

    benchmarkllm-as-judge
  807. arxiv:2609.33295 · cs.AI
    TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces
    Dehai Min, Daoan Zhang, Yiming Zeng, Huayi Zhang +12

    An agent can complete a task while exhibiting undesirable behavior during execution. Developers need tests for the specific behaviors encountered in deployment, beyond fixed benchmark suites. We present TraceDance, an agent system that constructs targeted benchmarks from deployment traces for user-s

    agentagent systemtool useself-improvementbenchmark
  808. arxiv:2609.33290 · cs.CL
    Calibration, Not Answer Selection: Distilling Internal Confidence in Reasoning Models
    Yadong Xi, Rongsheng Zhang, Tangjie Lv, Ziyang Luo +1

    Reinforcement learning with binary correctness rewards trains correctness, not calibrated confidence. The confidence that reasoning models verbalize is systematically overconfident, and the problem is not merely one of scale: verbalized confidence tracks how willing a model is to commit to an answer

    benchmark
  809. arxiv:2609.33289 · cs.AI
    Learning to Sell: Reinforcement Learning for Strategic Large Language Model Agents in Multi-Product Markets
    Shuze Daniel Liu, Claire Chen, Jiuqi Wang, Thorsten Joachims

    Autonomous large language model (LLM) agents operating in multi-product markets must make sequential decisions under information asymmetry and resource constraints. We develop a machine learning approach for training such agents to act effectively as sellers in a multi-item bargaining environment, w

    agentpost-training
  810. arxiv:2609.33288 · cs.RO
    Informative Viewpoint Selection for Episodic-Memory Embodied Question Answering using Omnidirectional Images
    Kaname Kitamura, Asako Kanezaki

    Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omn

    embodiedagent
  811. arxiv:2609.33286 · cs.CV
    InfoEdit: Probing Global Layout Reasoning in Infographic Editing
    Cheng Yang, Chufan Shi, Huijuan Wang, Bo Shui +5

    Multimodal foundation models edit natural photographs at production quality, yet the same models struggle with structured visual content such as infographics. Unlike photographs, infographics encode information through logical relations; editing one element often requires surrounding elements to be

    benchmarkevaluation protocol
  812. arxiv:2609.33279 · cs.LG
    Domain Generalization under Sampling Pattern Shifts in Irregular Time Series
    Changhun Kim, Joohyung Lee, Kwanhyung Lee, Donghwee Yoon +2

    Irregularly sampled multivariate time series (ISMTS) are prevalent in real-world applications, where both observation times and available measurements can vary substantially across domains. While recent models increasingly exploit such sampling information for prediction, its robustness under sampli

    benchmark
  813. arxiv:2609.33276 · cs.AI
    ChronoFlow: Hierarchical Flow Matching for Irregular Time Series Generation
    Changhun Kim, Sunguk Jang, Jeongjun Lee, Juhwan Choi +4

    Recent advances in generative modeling have substantially improved time series generation, yet most existing methods either assume a regular temporal grid or focus on feature dynamics under a given sampling structure. This makes them illsuited for generating irregular time series in their native for

    benchmark
  814. arxiv:2609.33270 · cs.AI
    Structured Sparse Memory for Recurrent Reasoning
    Zixuan Zhao, Samuel Wheeler, Neil Getty, Xiaotian Duan +2

    Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a comp

    memory
  815. arxiv:2609.33269 · cs.RO
    Q-WAM: 4-Bit Quantization of World Action Models with Action-Subspace Protection
    Arash Akbari, Arman Akbari, Jingwu Luo, Yuhao Lei +6

    World Action Models (WAMs) jointly generate video and robot actions through iterative diffusion and perform strongly in robotic manipulation. However, their prohibitive compute and memory costs pose substantial deployment challenges. Post-training quantization (PTQ) can reduce these costs, but exist

    manipulationhumanoidrobotwinmemorypost-trainingbenchmark
  816. arxiv:2609.33268 · cs.AI
    LSTMem: Hierarchical Long Short-Term Online Memory for Large Language Models
    Xianglong Shi, Ruijie Yang, Sirui Zhao, Shukang Yin +3

    Large language models increasingly serve as long-horizon assistants and agents, where they must both accumulate information across interactions and make the relevant parts available when later requests depend on them. Existing compact online memories typically use a single persistent state both to a

    memorypersistent statebenchmark
  817. arxiv:2609.33264 · cs.RO
    VehDyn: A Driving World Model Benchmark for Vehicle Dynamics
    Tianyi Wang, Wangsheng Du, Jiazhou Chen, Tianyi Zeng +11

    Video world models are emerging as data engines, action planners, and generative simulators for autonomous driving, but existing benchmarks primarily assess visual fidelity and coarse physical plausibility, providing limited evidence on whether generated driving futures obey realistic vehicle kinema

    world modelbenchmarkevaluation framework
  818. arxiv:2609.33261 · cs.CV
    EngIntervene: Benchmarking Multimodal Engineering State Understanding and Design Intervention Reasoning
    Jinchang Zhang, Yingda Tao, Jiakai Lin, Guoyu Lu

    Multimodal engineering benchmarks largely evaluate static understanding, such as recognizing components, interpreting diagrams, or answering technical questions. This leaves a missing middle between engineering perception and full design generation: whether a model can use an understood system state

    benchmark
  819. arxiv:2609.33260 · cs.LG
    CORTEX: A Verified Experience Layer for Generalist Agents
    Garapati Keerthana, Manik Gupta

    An agent can solve a task today and face the same task under new facts, tools, or governing knowledge tomorrow. Most agent systems can retrieve relevant text or recall prior conversations, but they lack a principled way to decide when a previous solution is still valid, when it must be adapted, and

    agentagent system
  820. arxiv:2609.33258 · cs.RO
    PORTER: Edge-Cloud Residency for Persistent 3D Scene Graph Memory
    Yue Chang, Yifan Tian, Jiajing Peng, Dazhi Huang +4

    Recent task-driven and just-in-time 3D Scene Graph (3DSG) methods reduce per-task representations by constructing or activating only task-relevant information. Yet sparse per-task working sets do not bound onboard memory usage over a robot's lifetime: as tasks change, payloads accumulated for earlie

    memoryscene graph
  821. arxiv:2609.33256 · cs.RO
    ActionGround: Training-Free Runtime Refinement of Frozen VLA Policies
    Namai Chandra, Madhur Thareja, Shriram Damodaran, Addison Lin Wang

    Vision-Language-Action (VLA) models map visual observations and language instructions directly to robot actions, but they do not explicitly represent the phase structure of manipulation tasks or the rigid-body dynamics governing execution. We present ActionGround, a neuro-symbolic, training-free run

    vision-language-actionvlavla policymanipulationopenvlalibero
  822. arxiv:2609.33254 · cs.LG
    BERT4DTI : BERT-based Model for Predicting Drug-Protein Interactions
    Thanina Hamitouch, Khadidja Henni, Abdelkrim Arie, Amina Selma Haichour +2

    Understanding how drugs interact with protein targets is fundamental to drug discovery, drug repurposing and the early identification of promising therapeutic candidates before costly experimental testing. Sequence-based DTI models face three practical limitations: labelled interactions are scarce a

    benchmark
  823. arxiv:2609.33248 · cs.LG
    Feedback-Robust AI for Patient Knowledge Graphs
    Mohammed Sameer Syed

    Patient knowledge graphs from bedside monitoring should type their relations and state whether the data support their signs. In anesthesia and intensive care, clinicians titrate drugs and ventilation in response to the physiology, so temporal relations mix the patient's response with the clinician's

    knowledge graph
  824. arxiv:2609.33244 · cs.AI
    ActiveMem: Dynamic Latent Memory Trees for Long-Horizon Agents
    Song-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan

    Large Language Model (LLM) agents increasingly rely on external memory to support long-horizon reasoning and decision making. Existing memory systems typically retrieve historical trajectories or summaries as independent context fragments, overlooking the procedural dependencies underlying multi-ste

    memoryexternal memoryagentagent benchmarkbenchmark
  825. arxiv:2609.33243 · cs.AI
    CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents
    Song-Li Wu, Jingyi Wang, Zhaocheng Du, Weinan Gan +1

    Code agents require long-horizon decision-making over complex interaction trajectories. However, existing reinforcement learning (RL) approaches typically optimize behavior at the token level, creating a mismatch between low-level generation and high-level behavioral reasoning. This limitation leads

    agentagenticbenchmark
  826. arxiv:2609.33239 · cs.LG
    Schrödinger--Föllmer Actor--Critic: Diffusion Policy Improvement with Finite-Sample Analysis
    Yuling Jiao, Lican Kang, Jerry Zhijian Yang, Jincheng Ying

    Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schrödinger--Föllmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (K

    diffusion policy
  827. arxiv:2609.33237 · cs.RO
    SurgFlow: 3D Object-Centric Contact Flow for Surgical Robot Manipulation
    Changwei Chen, Xiao Liang, Yinuo Yang, Nicole Shen +5

    Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should m

    manipulationhumanoidgrasp
  828. arxiv:2609.33236 · cs.CV
    PARSEE-VAD: Efficient Training-Free Online Video Anomaly Detection via Proposition-Aware Reasoning and Streaming Evidence Escalation
    Ji Wang, Shuangqing Zhang, Guo-Sen Xie, Fang Zhao

    Training-free online video anomaly detection (VAD) with frozen multimodal language models faces two coupled challenges: extracting reliable current-window semantics under causal and computational constraints, and maintaining temporal continuity without repeatedly transmitting high-dimensional histor

    benchmark
  829. arxiv:2609.33235 · cs.RO
    Learning with Object-centric Representations of Tactile Interactive Perception for Robot Manipulation
    Xinyi Yang, Zilin Si, Zhuowei Xu, Zeynep Temel +1

    Implicit object properties that are difficult to directly infer from vision, such as material, container contents, or softness, can be revealed through tactile sensing and exploratory interactions. However, because tactile signals are transient and sparse, extracting informative tactile events and e

    manipulationtactile
  830. arxiv:2609.33233 · cs.AI
    Inspire: Benchmarking Scientific Literature Search for Open Research Problems
    Jianrong Ding, Zhengyan Shi, Jianyuan Zhong, Kai Qiu +4

    Scientific literature search often begins with an open research problem rather than a known target paper or a fixed candidate set. We introduce INSPIRE, a benchmark for evaluating agents that search prior literature to make progress on solution-redacted research problems. Each instance pairs a resea

    agentbenchmark
  831. arxiv:2609.33232 · cs.LG
    MTLiquid: Enabling Efficient Multi-Task Learning using Liquid Neural Networks for Lightweight Healthcare Monitoring Systems
    Rachmad Vidya Wicaksana Putra, Fahad Abdul Rauf, Muhammad Shafique

    Continuous-time sensing and monitoring with timely and accurate decision-making are critical for many real-world applications. In healthcare monitoring systems, physiological signals are often available or sampled at irregular time intervals, hence requiring continuous-time processing to provide acc

    memory
  832. arxiv:2609.33230 · cs.RO
    AevaScenes: An FMCW LiDAR Dataset and Benchmark for Long-Range Perception
    Gautham Narayan Narasimhan, Heethesh Vhavle, Kumar Bhargav Viswanatha, James Reuther +1

    FMCW LiDAR measures per-point radial Doppler velocity alongside range, providing a motion cue unavailable in conventional time-of-flight sensors. Exploiting this signal at long range remains understudied. We present an FMCW LiDAR dataset of 575 sequences (57.5K frames) with over 8 million annotated

    benchmark
  833. arxiv:2609.33226 · cs.CL
    Beyond Memory Construction: Rethinking Memory Access for LLM-based Conversational Agents
    Donghua Cai, Yongheng Deng, Yifei Wang, Zijun Shen +1

    Memory is a core component of conversational agents, enabling coherent and context-aware behavior over long interactions. Recent approaches commonly rely on LLM-based memory construction, where raw interactions are rewritten into structured memory units and later retrieved via a RAG pipeline. While

    memoryragrag pipeline
  834. arxiv:2609.33221 · cs.LG
    RMB: Reward Model Boosting Mitigates Reward Hacking
    Jiabin Fan, Dezhi Ye, Yongchang Hao, Lili Mou

    Reinforcement Learning from Human Feedback (RLHF) is a powerful technique for aligning large language models (LLMs) with human preference. However, it often suffers from the reward hacking issue, where policy optimization improves the proxy reward model while actually degrading performance with resp

    rlhf
  835. arxiv:2609.33218 · cs.RO
    Scope-WM: Scoped Computation for Efficient Visual World Models
    Chunzheng Li, Zesheng Jia, Hongda Zhang, Jiaying Tang +3

    Visual world models enable robotic planning by predicting future observations, but dense latent-state propagation and sample-intensive trajectory optimization incur high inference latency and peak memory usage, limiting real-time deployment on resource-constrained platforms. Existing sparse world-mo

    world modelaction-conditionedmemory
  836. arxiv:2609.33217 · cs.CV
    RepFlow: Reciprocal Supervision Improves Generation and Representation in Flow Models
    Weili Zeng, Feng Tian, Shengqi Liu, Yichao Yan

    Generative models learn visual structure through denoising, yet their internal states are entangled with both noise level and network depth, making it difficult to obtain a stable visual representation from the generator itself. We introduce RepFlow, which learns such a representation from the gener

    post-training
  837. arxiv:2609.33212 · cs.CL
    CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification
    Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan +2

    Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-S

    evaluation framework
  838. arxiv:2609.33208 · cs.CV
    WorldAgent: Verification-Guided Agentic Physical World Construction
    Caoliwen Wang, Mengdi Wang, Yige Chen, Zejia Wu +12

    Constructing complex physical worlds from language requires coordinating extensive 3D environments, detailed structures and objects at different spatial scales, and interacting physical processes under both stated goals and implicit physical constraints. We present WorldAgent, an agentic framework f

    agentic
  839. arxiv:2609.33207 · cs.LG
    MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers
    Rachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, Muhammad Shafique

    Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. However, huge parameter counts and complex multi-head self-attention (MHSA) operations make it challenging to achieve high energy efficiency in SViT infere

    memory
  840. arxiv:2609.33205 · cs.LG
    SMORE: Stability-Promoting Mesh-Agnostic Model Reduction for Time-Dependent PDEs
    Yangyuan Li, Weichao Li, Shaowu Pan

    High-fidelity simulations of time-dependent partial differential equations (PDEs) are computationally expensive, motivating data-driven reduced-order surrogates for many-query tasks such as uncertainty quantification, design optimization, data assimilation, and optimal control. However, existing sur

    latent dynamics
  841. arxiv:2609.33204 · cs.CL
    Turning Speech Language Models into Multilingual Listeners
    Tolúlopé Ògúnrèmí, Dan Jurafsky, Chris Manning, Ahmet Üstün +1

    Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present MULTISPEECHQA, a large-scale, synthetic

    benchmark
  842. arxiv:2609.33200 · cs.LG
    Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning
    Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil, Md. Faiyaz Abdullah Sayeedi +2

    On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding

    benchmark
  843. arxiv:2609.33197 · cs.RO
    TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation
    Yongsheng Zhao, Han Gao, Baoping Cheng, Jingyao Tang +6

    Vision-Language-Action (VLA) models provide a unified framework for grounding high-level semantic information into low-level robot actions, enabling scalable robotic manipulation across diverse tasks. However, existing VLA models lack explicit mechanisms to disentangle the states and intents of the

    vision-language-actionvlavla modelmanipulation
  844. arxiv:2609.33196 · cs.AI
    Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries
    Haiquan Hu, Yuzhu Liang, Weicheng Tang, Yanzeng Li +2

    Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boun

    benchmarkleaderboard
  845. arxiv:2609.33186 · cs.LG
    Offline Policy Evaluation via Mixed Bellman Residuals and Adaptive Critic Representations
    Amitakshar Biswas, Yuhan Li, Ruoqing Zhu

    Evaluating a target policy using data generated by a different behavior policy remains a fundamental challenge in reinforcement learning. While most existing work relies on the standard one-step Bellman residual, we consider a convex combination of one-step and two-step residuals with a fixed mixing

    policy evaluation
  846. arxiv:2609.33183 · cs.LG
    Identifying Temporal Features within Transcoders for Time Sensitive Factual Recall
    Sanjay Govindan, Yang Song, Maurice Pagnucco

    Large Language Models (LLMs) suffer from temporal misalignment, often due to the contradictory nature of their training corpora. While current mitigation strategies rely on computationally expensive fine-tuning or context-heavy retrieval augmented generation (RAG), the internal mechanisms governing

    retrieval augmented
  847. arxiv:2609.33181 · cs.AI
    SeOPD: Self-Evolving LLMs via Online Policy Distillation from Self-Generated Chain-of-Thought
    Xiaoshu Chen, Xiangyu Wong, Sihang Zhou, Ke Liang +1

    Recent advances in online policy self-distillation (OPSD) have demonstrated that large language models (LLMs) can improve their capabilities by leveraging external privileged information (PI), such as manual annotations or feedback from external environments. However, obtaining accurate annotations

    self-improvementself-evolving
  848. arxiv:2609.33180 · cs.LG
    Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks
    Xiaojing Sun, Yuhan Zeng, Zihua She, Xiao Wang

    As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite r

    self-improvementbenchmarkevaluation framework
  849. arxiv:2609.33177 · cs.RO
    DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model
    Tianyun Jiang, Wenrui Bao, Bingxin Xu, Yu Tian +1

    World-Action Models (WAMs) offer visual foresight for robotic manipulation, but pixel-space models repeatedly reconstruct entire future scenes, incurring high computational cost and spatio-temporal redundancy. In physical manipulation, consecutive frames often share most of their visual context; the

    manipulationliberoworld model
  850. arxiv:2609.33172 · cs.RO
    Dynamic Manipulation with World-Action Models via Counterfactual Planning
    Sunwoo Park, Wonbin Lee, Seonghyun Jin, Youngmin Kim +2

    World-Action models (WAMs) trained on static demonstrations often fail to manipulate moving targets even when they possess the required manipulation skills. We attribute this failure to target-response collapse: as execution advances, the policy becomes increasingly biased toward the learned continu

    manipulation
  851. arxiv:2609.33171 · cs.LG
    Perturb-and-Solve: Efficient Learned-Operator Conditioning for Latent Diffusion Inverse Problems
    Abduragim Shtanchaev, Arip Asadulaev, Luiza Labazanova, Aidar Alimbayev +2

    Latent diffusion models serve as powerful priors for solving inverse problems in image restoration, such as deblurring, inpainting, and super-resolution. Current methods have a trade-off between generality and efficiency. Solvers that are restricted to a fixed set of degradation operators are fast a

    memory
  852. arxiv:2609.33169 · cs.RO
    When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability
    Xingjian Li, Yi Han, Jianhua Z. Huang

    Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matter

    memory
  853. arxiv:2609.33165 · cs.RO
    Beyond Tasks: A Vision for Reproducing an Animal-like Behavioral Substrate Using Modern Robot Learning Techniques
    Samiyuru Menik, Hemadri Jayalath

    Recent advances in robot learning have produced increasingly capable embodied agents. Yet comparatively less attention has been given to a more basic form of competence that animals exhibit continuously: the ability to remain situated, responsive, and behaviorally coherent as physical, environmental

    embodiedembodied agent
  854. arxiv:2609.33161 · physics.app-ph
    Uncertainty Quantification for the Fission Matrix Method: A Rigorous Mathematical Framework and Computationally Efficient Alternatives
    Valerio Mascolino

    The fission matrix (FM) method recasts neutron transport in matrix form, enabling fast, interpolation-based reactor calculations from a pre-computed database of Monte Carlo-derived coefficients. Propagation of the underlying Monte Carlo statistical uncertainty to the FM eigenvalue and eigenvector is

    memorybenchmark
  855. arxiv:2609.33158 · cs.CV
    FOCUS: Benchmarking Retinal Model Generalization from Foundation Vision Encoders to Multimodal LLMs
    David Restrepo, Chenwei Wu, Luis Filipe Nakayama, Miguel L. Martins +3

    Progress in AI-based retinal image analysis has advanced with foundation models, yet evaluating their reliability remains challenging. Performance reported on a single dataset does not capture how models behave under dataset shift, across clinical definitions, or for different patient subgroups. Thi

    benchmarkleaderboard
  856. arxiv:2609.33157 · cs.RO
    TimelyDAgger: Timing-Aware Expert Querying for VLA Policy Improvement
    Zhixuan Zhao, Peiyan Li, Enhao Zhang, Yueran Tao +9

    DAgger improves robot policies by aggregating expert supervision from states visited during policy execution. Robot-gated DAgger automates expert queries, allowing the robot to decide when to request expert takeover. While existing gates emphasize detecting the need for assistance, takeover timing a

    vision-language-actionvlavla policypost-trainingevaluation framework
  857. arxiv:2609.33155 · cs.CL
    Where Do Test-Time Scaling and Training Fall Short in Individual Stance Prediction?
    Yuyang Zhao, Xuan Liu, HaoYang Shangm Haojian Jin

    Test-time scaling and post-training have improved LLM performance in coding and mathematical reasoning, but their effectiveness for individual stance prediction remains unclear. We study this question by predicting a person's stance in a new discussion from their history. We evaluate widely used tes

    post-training
  858. arxiv:2609.33153 · cs.AI
    What Does a Skill Actually Do? Estimands and Evaluation Validity for Tool and Skill Use in LLM Agents: A Critical Review
    Shuyang Zhang

    Large language model (LLM) agents increasingly draw on external tools and reusable skills selected at run time from libraries that hold thousands of entries. Reports that a retriever, router, or skill library "improves" an agent may refer to retrieval recall, the success change from enabling a libra

    agentllm agent
  859. arxiv:2609.33150 · cs.CL
    Generalization Dynamics of LM Pre-training
    Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

    People typically assume that LMs stably mature from pattern-matching parrots to generalizable intelligence during pre-training. We build a toy eval suite and show this mental model is wrong: throughout pre-training, LMs frequently and suddenly hop between parrot-like and intelligence-like computatio

    eval
  860. arxiv:2609.33147 · cs.LG
    CFLoRA: Federated Fine-tuning of LLMs with Complementary Factors for Error-free Aggregation
    Yanan Ma, Qiyuan Chen, Zihan Fang, Xianhao Chen +1

    Federated low-rank adaptation (LoRA) enables collaborative fine-tuning of large language models without centralizing private client data. Its factorized update, however, creates a structural mismatch in federated averaging: averaging the two LoRA factors separately does not equal averaging their pro

    benchmark
  861. arxiv:2609.33146 · cs.AI
    LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks
    Euntae Choi, Sumin Song, Sungjoo Yoo

    An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budg

    agentllm agentagenticbenchmark
  862. arxiv:2609.33144 · cs.LG
    Beyond the Training Horizon: Mechanisms and Limits of Length Generalization in Looped Transformers
    Jia Liang, Xi Jin, Liangming Pan

    Looped Transformers can generalize to reasoning chains longer than those encountered during training, but the computations enabling this behavior and limiting its extent remain unclear. We mechanistically compare two looped-Transformer configurations, which we call the Matched-Recurrence Looped Tran

    self-correction
  863. arxiv:2609.33143 · cs.LG
    CAME: Company-Aware Evidence-Memory Experts for Interpretable Quarter-Ahead Revenue Forecasting
    Ya-Wen Wu, Meng-Fen Chiang, Kuang-Da Wang, Wen-Chih Peng

    Quarter-ahead revenue forecasting requires company-scale numerical accuracy, strict temporal validity, and company-specific interpretation of narrative disclosures. LLMs can distill textual evidence but can produce scale-misaligned forecasts, whereas history-based anchors are stable but miss forecas

    memory
  864. arxiv:2609.33141 · cs.AI
    On Device Agentic Operation Caches -- Classifier-Centric NL-to-Action Generation
    Moghis Fereidouni, Anthony Arnold, Sumit Gulwani, Mark Marron +1

    Agentic AI is increasingly being embedded in software applications to provide natural language interfaces to features and functionality. In most cases these agents are powered by enterprise (100+ billion parameter) or frontier class large language models that require substantial computational resour

    agentic
  865. arxiv:2609.33138 · cs.RO
    Multi-Modal Non-Prehensile Estimation of Physical Parameters via Press-and-Pull Tipping
    Steven M. Hyland, Jing Xiao, Cagdas D. Onal

    Recovering physical properties of unknown objects through non-prehensile interaction is challenging because no single manipulation primitive reveals all relevant parameters. Planar pushing couples mass and friction, while conventional tipping cannot recover friction and may fail entirely when low-fr

    manipulationgrasp
  866. arxiv:2609.33131 · cs.LG
    ILP-BO: Integer Linear Programming-Based Black-Box Optimization
    Hyakka Nakada, Shu Tanaka

    Black-box Optimization (BO) is a powerful framework for optimizing expensive objective functions or unknown functions with a limited number of evaluations. A central step of standard BO such as Bayesian optimization is the optimization of a surrogate-based acquisition criterion, which is commonly pe

    benchmark
  867. arxiv:2609.33127 · cs.LG
    Policy Plasticity Matters in Offline-to-Online Reinforcement Learning: Refitting Offline Policies for Online Adaptation
    Yuheng Huang, Yunpeng Qing, Yixiao Chi, Yilun Kong +1

    Offline-to-Online Reinforcement Learning (O2O RL) has emerged as a practical paradigm that pre-trains the policy using static offline datasets and subsequently adapts the policy through online interactions. Existing O2O methods primarily address the transition through value calibration, while genera

    online learning
  868. arxiv:2609.33125 · cs.RO
    Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface
    Zhizhen Zhang, Yuxia Fu, Zijian Wang, Helen Huang +1

    Co-training offers a straightforward way to build a multi-task vision-language-action (VLA) policy, but can fall short of the performance achieved by training each task independently. The challenge is to retain these task-specific gains in a multi-task policy without joint post-training. Combining i

    vision-language-actionvlamanipulationgr00tliberopost-training
  869. arxiv:2609.33123 · cs.AI
    Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring
    Zhixiang Zhang, Zesen Liu, Wai Ip Lai, Hongxu chen +1

    Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or

    agentself-evolvingbenchmark
  870. arxiv:2609.33120 · cs.LG
    Mycelium: A Generalizable Cross-Grid Multi-Task Model for Electrical Distribution Systems
    Zhengyang Wei, Shourya Bose, Helgi Hilmarsson, Elena Carnio +1

    Electrical distribution grid operations require inference across heterogeneous networks from sparse, noisy, and incomplete time series measurements. In this work, we identify challenges and explore solutions towards a unified model that can perform diverse tasks grounded in the physics of the electr

    benchmark
  871. arxiv:2609.33119 · cs.AI
    MedRouter: Demystifying Knowledge Differences Across Medical LLMs for Routing-Based Reasoning
    Lang Cao, Binghang Lu, Yuhao Shen, Yue Guo

    Medical question answering spans diverse specialties and modalities, and individual medical large language models (LLMs) exhibit distinct strengths across tasks and domains. This heterogeneity suggests that combining specialists may enable broader coverage of medical questions than relying on any si

    agenticbenchmark
  872. arxiv:2609.33117 · cs.LG
    ECG-Scroll: A Long-Horizon, Streaming Benchmark and Agent Environment for Interpretation of Ambulatory Electrocardiograms
    Haitao Li, Chenglin Li, Zhengyao Ding, Ziyu Li +2

    Multimodal large language models (MLLMs) can now interpret a standard ten-second, twelve-lead electrocardiogram (ECG) with clinically grounded, reward-verified reasoning. Real cardiac monitoring is different. Ambulatory (Holter) and telemetry recordings span hours to days and are read as they stream

    memoryagentllm agenttool usebenchmark
  873. arxiv:2609.33115 · cs.AI
    Modular Discovery of General Game-Playing Algorithms with Large Language Models
    Zun Li, John Schultz, Marc Lanctot, Daniel Hennes

    General Game Playing across arbitrary games from rules alone remains challenging due to differing algorithmic requirements across game classes and strict decision-time constraints. Rather than hand-designing search heuristics for specific domains, can we leverage Large Language Models (LLMs) to disc

    multi-agentbenchmark
  874. arxiv:2609.33114 · cs.LG
    Can Tabular Foundation Models Amortize Statistical Inference?
    Kai Ye, Shijin Gong, Hongyi Zhou, Valentina Zangirolami +1

    For decades, statistical inference has largely been developed one problem at a time. Given a scientific target, such as a treatment effect or a regression function, statisticians design a problem-specific estimator together with a procedure for quantifying its uncertainty. This paper proposes a diff

    post-trainingbenchmark
  875. arxiv:2609.33113 · cs.AI
    ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding
    Tao Long, Weili Shi, Hussein Mozannar, Maya Murad +1

    As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent work. Through a formative study (N=14), we identified PILOT: five supervisory practices for Planning, I

    agent
  876. arxiv:2609.33112 · cs.LG
    Simulation-Free Learning of GP-SDEs from Irregular Observations
    Zhidi Lin, Yuhao Liu, Ying Li, Edwin Fong +1

    Gaussian process stochastic differential equations (GP-SDEs) provide a flexible Bayesian model for unknown continuous-time state dynamics with uncertainty quantification, but learning and inference from noisy and irregular observations remain computationally challenging. To address this issue, we pr

    benchmark
  877. arxiv:2609.33110 · cs.LG
    D-JEPA: Design-Recoverable JEPA Representation with Swappable Physics Decoders
    Nitin Nagesh Kulkarni, Aashwin Anand Mishra, Yin Yu, Peter Lyu

    Joint-Embedding Predictive Architectures (JEPAs) provide a framework for learning compact representations without directly reconstructing high-dimensional observations. However, in parameterized physical systems, learned representations can entangle geometry with operating conditions and task-specif

    benchmark
  878. arxiv:2609.33109 · cs.CV
    Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models
    Tuo Liang, Disheng Liu, Nengbo Wang, Vipin Chaudhary +1

    Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orient

    benchmark
  879. arxiv:2609.33104 · cs.RO
    VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning
    Zhenghao Xiao, Minting Pan, Nantian He, Dongzhan Zhou +1

    While action-conditioned video prediction provides an intuitive world model for robotics, purely data-driven predictors often suffer from compounding errors and physically implausible hallucinations in long-horizon rollouts, severely undermining downstream action planning. We propose VPTwin, a Real-

    manipulationworld modelaction-conditioned
  880. arxiv:2609.33102 · cs.MA
    ORBIT: A Framework for Multi-Agent Safety and Security Evaluations
    Ben Hagag, William L. Anderson, Srija Chakraborty, Christian Schroeder de Witt

    Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, from

    agentmulti-agentagenticbenchmarkevaluation framework
  881. arxiv:2609.33101 · cs.RO
    Evolving Dexterous Robots from Scratch
    Zihan Guo, Shuzhe Zhang, Muhan Li, Peiyang Li +1

    Little is known about how to manually design agents capable of dexterous manipulation. Some design principles have been inferred from close examination of how animals manipulate objects, but these structures and behaviors have so far resisted biomimicry and may not be optimal for artificial machines

    manipulationdexterousmanipulator
  882. arxiv:2609.33097 · cs.RO
    Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation
    Zhihao Chen, Yiyuan Ge, Ziyang Wang, Pu Cao +1

    Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deploy due to their large parameter counts and computational requirements. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence select

    benchmark
  883. arxiv:2609.33096 · cs.RO
    Humanoids for Robot-Assisted Surgery: Bimanual Base Placement and Tool-Mount Optimization via Capability Maps
    Peihan Zhang, Zekai Liang, Florian Richter, Nikita Thareja +3

    Rapid advances in humanoid robotics have motivated growing interest in the application of humanoids for healthcare and clinical tasks. However, it remains unclear how close contemporary humanoids are to meeting the kinematic demands of robot-assisted laparoscopic surgery. In this work, we address th

    humanoid
  884. arxiv:2609.33095 · cs.RO
    REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening
    Peizhen Li, Longbing Cao, Yang Zhang

    Generating responsive listener facial motion is an important task for embodied conversational AI. Two modeling challenges are central: accounting for the timing of speaker cues while maintaining continuity with the listener's ongoing motion, and capturing locally variable facial events alongside the

    embodiedhumanoid
  885. arxiv:2609.33093 · cs.LG
    How Linear Attention Remembers
    Kichang Lee, JaeYeon Park, Songkuk Kim, JeongGil Ko

    Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memor

    memory
  886. arxiv:2609.33090 · cs.CV
    OneSign: Unifying Sign Language Understanding Tasks with One Model
    Shiwei Gan, Yafeng Yin, Xiao Liu, Desibieer Tuerdaken +2

    SLU encompasses a diverse set of tasks, including ISLR, CSLR, and SLT. Although these tasks share basic semantic and linguistic foundations, they are typically addressed with task-specific architectures and training pipelines, which hinders knowledge sharing and requires costly pretraining and finet

    benchmark
  887. arxiv:2609.33077 · cs.LG
    Parameter-Efficient 3D Segmentation of Liver and Liver tumors: Depthwise factorization Scales Better Than Dense Convolution with Spatial Dimensionality
    Adham M. Alkhadrawi, Mohammed A. B. Mahmoud

    Three-dimensional dense convolutional networks are the strongest performers on volumetric medical image segmentation, but their parameter counts scale poorly: moving a dense k x k convolution to k x k x k multiplies its weights by k. We observe that depthwise separable factorization does not share t

    benchmark
  888. arxiv:2609.33075 · cs.AI
    QureRadEmbed: Structuring Radiological Similarity through Attribute and Reasoning Supervision
    Janhavi Prabhu, Sahil, Shivam Ashok Shukla, Manoj Tadepalli

    Radiological similarity depends on disease relationships and on fine details such as laterality, lobe, severity, size, and certainty. Broad biomedical similarity can overlook these qualifiers, particularly when several attributes vary together. We introduce QureRadEmbed, a 4B radiology-aware encoder

    benchmarkevaluator
  889. arxiv:2609.33066 · cs.LG
    Zero-Storage Procedural Neural Synthesis via Boundary Dynamics: Formal Verification in Lean 4 and Bare-Metal Gauntlet Validation
    Volkan Dağlı, Zerrin Dağlı, Dağhan Dağlı

    Contemporary neural inference architectures rely on dense floating-point weight matrices stored in high-bandwidth memory (VRAM), incurring severe memory-wall bottlenecks and preventing native execution inside deterministic virtual machines like the Ethereum Virtual Machine (EVM). Verifying terminati

    memory
  890. arxiv:2609.33061 · cs.AI
    LLM sequential decision making under uncertainty in biochemical domains
    Mattias Akke, Soojung Yang, Jurgis Ruža, Sathya Edamadaka +1

    Large language models (LLMs) are increasingly used to drive scientific discovery. Understanding how LLMs make decisions from new data and memory of the literature is vital before trusting them to design experiments under tight experimental budgets. However, their decision strategies are invisible in

    memorybenchmark
  891. arxiv:2609.33053 · cs.RO
    SwingRL: Adaptive Observation Reinforcement Learning with World-Model Prediction for Cable-Suspended Hoisting Control
    Guangming Wang, Xiaoyu Zhang, Yucheng Xin, Wanli Ma +7

    Cable-suspended hoisting is widely used to move heavy or bulky payloads that cannot be handled conveniently by rigid pick-and-place systems, for example in crane-assisted construction. Robotic hoisting using flexible cables is challenging because payload motion is underactuated, external disturbance

    world model
  892. arxiv:2609.33051 · cs.LG
    SketchSSM: Write to the Full State, Read from a Compact Sketch
    Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer +2

    Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires

    benchmark
  893. arxiv:2609.33049 · cs.RO
    Is Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via Perturbation
    Bumgeun Park, Donghwan Lee

    Behavior cloning (BC) is vulnerable to covariate shift during closed-loop execution, where small prediction or execution errors can drive the robot toward states poorly covered by the demonstration data. We focus on action-sequence planning, where a policy predicts a finite-horizon sequence of actio

    manipulation
  894. arxiv:2609.33044 · cs.CL
    Pinned and Still Unstable: Within-Judge Verdict Variance and the Noise Floor of LLM-as-Judge Leaderboards
    Krishna Chytanya Ayyagari

    Modern LLM evaluation assumes that pinning a judge to a fixed model snapshot and decoding at temperature zero yields reproducible verdicts. We show this assumption fails as a property of how LLM-as-Judge is operationalized on cloud serving infrastructure, not of any particular model family. Across f

    benchmarkllm-as-judgeleaderboard
  895. arxiv:2609.33039 · cs.AI
    Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States
    Difan Jiao, Ashton Anderson

    Language model agents can now perform sophisticated sequences of actions via tools and harnesses, which has increased the scope of the damage they can cause. Guard models, however, are mainly built for content moderation and thus are not well-suited to detecting this agentic risk. To address this, w

    agentictool usebenchmark
  896. arxiv:2609.33037 · cs.LG
    Byzantine-Robust Federated RAG via Aligned Calibration and Fixed-Membership Conformal Prediction
    Prasanjit Dubey, Xiaoming Huo

    Retrieval-augmented generation (RAG) lets language models answer questions more accurately by consulting relevant documents. Many valuable collections, such as medical records, cannot be pooled because of privacy rules. Federated RAG leaves each collection with its owner, or node, which scores candi

    retrieval-augmentedrag
  897. arxiv:2609.33030 · cs.RO
    What Must a World Model Distinguish for Planning?
    Rongzhe Wei, Hans Hao-Hsun Hsu, Peizhi Niu, Yifan Li +1

    World models simulate the consequences of action candidates, but good planning need not preserve every physical distinction required for accurate prediction. We formalize this gap through a hierarchy of mechanism, response, and decision sufficiency. Given a candidate set, the planning query determin

    world modelaction-conditioned
  898. arxiv:2609.33023 · cs.AI
    SRE-Marathon: A Continuous, Change-Driven Benchmark for Autonomous Site Reliability Agents
    Yifang Tian, Yingjian Bai, Yifeng He, Zichun Chong +3

    Benchmarks for site reliability engineering (SRE) agents are typically episodic: one fault is injected, the agent receives an incident task, and its response is scored. Production operation is not. Incidents surface through noisy alerts, overlap in time, and often originate from code or configuratio

    agentbenchmark
  899. arxiv:2609.33020 · cs.CV
    Residual Diffusion Implicit Models
    João Guerreiro, Pedro Tomás, Helena Aidos, Jacinto C. Nascimento

    Diffusion models achieve state-of-the-art results across multiple tasks. However, in inverse problems, standard initialization from pure Gaussian noise misaligns the generative process with real-world degradations. More recent methods such as diffusion bridges impose strict endpoint constraints and

    benchmark
  900. arxiv:2609.33017 · cs.AI
    Trust and Task Completion in the World of Consumer AI Agents
    Jeroen Olieslagers, Eduardo Pujol, Gal Zahavi, Lukas Ingemarsson +1

    Action agents do things for people. They send email, spend money, and call businesses while the user is busy with something else, so a mistake can turn into an action before anyone notices. They fail their users in two ways. They break trust when they do something the user never agreed to, or hold b

    ai agent
  901. arxiv:2609.33014 · cs.AI
    TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference
    Tzu-Heng Huang, Jet Lin, Eric Lin

    Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rar

    benchmark
  902. arxiv:2609.33013 · cs.AI
    The Epistemics of Agent Memory: Measuring, and Governing, the Consolidation Decision in Long-Horizon LLM Agents
    Sasank Annapureddy, Anjaneya Prasad Thamatani

    Long-horizon LLM agents must convert accumulated experience into durable memory, deciding what to keep, compress, abstract into reusable skills and rules, or forget. We report a four-phase research program on this consolidation problem whose central finding is a shift in what is measured: from how m

    agent memoryagentllm agentbenchmark
  903. arxiv:2609.33012 · cs.AI
    When Pair Count Is Not the Sample Size: What All-Pairs Agent Comparisons Estimate
    Wei-Jung Huang

    When an agent benchmark compares every pair of leaderboard entries, the number of comparisons can look much larger than the independent evidence behind them: A versus B and A versus C both reuse A. Whether this reuse affects inference depends on what the analysis is meant to describe. If the board a

    agentagent benchmarkbenchmarkleaderboard
  904. arxiv:2609.33010 · cs.AI
    Model-Aware Data Selection from In-and-Out Information Interplay
    Yifan Wang, Xiaomin Li, Yuexing Hao, Dongwon Jung +6

    LLMs are effective representations that assimilate vast amounts of knowledge during pretraining, but post-training is necessary for models to reliably access this knowledge and "know what they know." We observe an interesting rank equilibrium between knowledge stored in the weights and the data stre

    post-training
  905. arxiv:2609.33007 · cs.RO
    CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning
    Shivam Aarya, Zhang Xi-Jia, Chengyue Huang, Junhyun Kim +5

    Robot learning has largely relied on human-teleoperated demonstrations to acquire effective learnable behaviors. However, human-operated data collection processes can be unintuitive, difficult to scale, and inherently asynchronous. We explore an alternative: distilling physical behavior from general

    manipulationdiffusion policyfranka
  906. arxiv:2609.33000 · cs.RO
    TriDrive: Joint Driver, Vehicle, and Road Modeling for Forecasting and Driver Monitoring
    Yuhang Wang, Jingxin Yang, Chuheng Wei, Yuechen Guo +3

    Predicting how drivers, vehicles, and road scenes interact and evolve together is central to driver monitoring. Prior work models in-cabin activity or traffic-conditioned driver motion in isolation, motivating joint driver, vehicle, and road modeling with real-time on-vehicle evaluation. We introduc

    v-jepabenchmark
  907. arxiv:2609.32996 · cs.CV
    Oracle Gaps in Reliability Coverage: Sampling Noise or Policy Specialization?
    Mert Onur Cakiroglu, Mehmet Dalkilic, Hasan Kurban

    Policies trained from the same base model can appear to solve different problems. An oracle that chooses the best policy for each problem may therefore appear much stronger than any single policy. Selecting the largest estimated success rate also selects favorable sampling errors. We study this effe

    post-training
  908. arxiv:2609.32993 · cs.AI
    X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
    Sitao Cheng, Xunjian Yin, Zhiyuan Sun, Yuxuan Li +3

    Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits unused, and flat training uses each scarce trajectory less f

    agent
  909. arxiv:2609.32991 · cs.LG
    What Should Data Teach? Moving Bottlenecks Across Circuit, Store, and Use
    Yixiao Chen, Ke Cheng, Jiangtao Guan, Shuo Huang +4

    What should data teach a language model at a particular point in training? A circuit view reveals three distinct bottlenecks: forming a computation, making its required content available, and selecting among available routes. A shared diagnosis-to-data principle connects them: localize the missing o

    long-context
  910. arxiv:2609.32990 · cs.AI
    Certified Long-Horizon Code Agent Evolution via Validation-Gated Skill Optimization
    Yifan Wang, Hao Cheng, Xiaomin Li, Yuexing Hao +12

    Long horizon agent self-evolution without model weight updates is essential for enabling deployed agents to accumulate reusable skills and improve over time. Prior self-evolution work has focused primarily on short-horizon tasks, while repository-level software engineering remains unexplored despite

    agentself-improvement
  911. arxiv:2609.32984 · cs.CV
    ReVision3D: Attribution-Guided Recursive Self-Improvement for 3D Medical Perception
    Ho Hin Lee, Yuyin Zhou, Yannan Yu, Shi Gu +1

    Recursive self-improvement (RSI) offers a promising path for overcoming the limited visual capability of current medical imaging agents. Yet applying RSI to volumetric imaging remains difficult: failures can arise from acquisition, perception, training recipe, or downstream inference, while self-gen

    self-improvement
  912. arxiv:2609.32980 · cs.LG
    Feasible Flow Matching for Graph Reconstruction via Within-Sampling Primal-Dual Guidance
    Haoming Chen, Nicolas Zilberstein, Santiago Paternain, Santiago Segarra

    Graph reconstruction from partial observations often comes with structural side information, such as degree bounds, triangle counts, or an edge-density band. Prior-Informed Flow Matching (PIFM) reconstructs graphs by transporting a local prior toward the graph distribution, but it provides no mechan

    benchmark
  913. arxiv:2609.32976 · cs.LG
    Adaptive Ensemble Selection for Noisy Labels on Tabular Data
    Faizaan Ali, Inwon Kang, Oshani Seneviratne

    Incorrect or corrupted labels in tabular datasets can significantly degrade supervised learning performance, particularly when mislabeling is subtle and not easily detectable from feature space alone. In the context of automated or AI-augmented data science workflows, robust detection of such label

    benchmark
  914. arxiv:2609.32966 · cs.LG
    Self-Confirming Superposition Traps in Reinforcement Learning
    Dai Shi, Andi Han, Feng Chen, Yiqun Duan +2

    Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirmin

    world modeldreamerv3
  915. arxiv:2609.32965 · cs.AI
    Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
    Hongyi Du, Tianyi Zhang, Weijia Zhang, Yi Yang +8

    Multiple agents may often conflict in an organization: for example, one coding agent changes an interface in a repository, but another continues to develop on the old version where existing tests become stale. A conversation can resolve the episode, but when the participants change, what makes the l

    agentmulti-agentbenchmark
  916. arxiv:2609.32964 · cs.AI
    The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining
    Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He +4

    Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally

    benchmark
  917. arxiv:2609.32961 · cs.LG
    Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
    Ritul Satish, Prasoon Sinha, Akiho Kawada, Neeraja J. Yadwadkar

    As LLM agents tackle longer tasks, they increasingly compress growing histories of reasoning, actions, and tool outputs. Compression can reduce token use, but it also changes the information available for later decisions. Existing agentic harnesses bundle decisions about what to compress, when to co

    context compressionagentllm agentagentic
  918. arxiv:2609.32960 · cs.AI
    Environmental Impact of Generative and Agentic AI: An in-Depth Analysis and Green Solutions
    Abderaouf Bahi, Amel Ourici, Ibtissem Gasmi

    The proliferation of generative and agentic artificial intelligence (AI) systems has introduced computational demands whose environmental consequences are substantial yet underexamined. This paper examines the environmental footprint of modern AI systems across energy consumption, carbon emissions,

    embodiedagentic
  919. arxiv:2609.32958 · cs.RO
    Learning Geometry-Aware Virtual Fixtures From Sparse Demonstrations
    Maximilian Mühlbauer, Marcella Piacentino, Raffaella Mancino, Paolino De Risi +9

    In many teleoperation applications, collecting a large number of demonstrations as required for traditional probabilistic learning from demonstration (LfD) approaches may not be feasible. To still give operators the ability to intuitively create trajectories as Virtual Fixtures (VFs), we propose to

    teleoperation
  920. arxiv:2609.32949 · cs.LG
    Optimal Nonparametric Dynamic Pricing with Censored Demand and Adversarial Inventory
    Mengxiao Zhang, Yingfei Wang, Haipeng Luo

    We study online dynamic pricing with censored demand, where an arbitrary inventory level is revealed before pricing and may adapt to past observations, while demand follows an unknown, price-dependent distribution that is stationary over time. For a horizon of $T$ rounds, Xu et al. [2026] achieved $

    benchmark
  921. arxiv:2609.32942 · cs.LG
    The Key Handoff: Retrieval in Hybrid Language Models
    Kaan Kale, Oguzhan Baser, Sriram Vishwanath

    A two-hop question makes a language model retrieve twice: once to produce a bridge entity, and once to retrieve with it. Transformers resolve that entity in their early layers. Hybrid models replace most of the attention with a recurrent state, so where the key becomes usable, and where it is spent

    memory
  922. arxiv:2609.32940 · cs.LG
    TwinS-GCN: Spectral conjugate for Spectral Graph Convolutional Networks
    Chun Hei Michael Chan, Flavia Petruso, Dimitri Van De Ville

    Graph convolutional networks propagate information by repeated local aggregation through a graph shift operator; i.e., a $K$-layer network reaches $K$ hops neighborhood. On the one hand, such spreading can lead to oversmoothing. On the other hand, long-range dependencies demand the depth. Transporti

    benchmark
  923. arxiv:2609.32939 · cs.LG
    Theory of Scene: Breaking the Symmetry Trap in Multi-Agent LLM Coordination
    Liangqi Yuan, Wenzhi Fang, Shiqiang Wang, Christopher G. Brinton

    Multi-agent systems built on large language models (LLMs) are largely homogeneous, as their agents behave alike even across distinct LLMs. We show that when such agents act concurrently without communication, they collide on targets they must split and diverge on targets they must take together, a d

    agentmulti-agentagenticagent systembenchmark
  924. arxiv:2609.32935 · cs.RO
    Communication-Aware Heterogeneous Graph Learning for Decentralized Multi-Human Multi-Robot Task Allocation
    Ziqin Yuan, Ruiqi Wang, Baijian Yang, Byung-Cheol Min

    Multi-human multi-robot (MH-MR) teams combine robotic autonomy with human expertise, but effective task allocation requires coordinating scarce, dynamically available human support with distributed robot execution. Limited robot-robot and human-robot communication further complicates this coupling b

    multi-agentbenchmark
  925. arxiv:2609.32934 · cs.LG
    The Impact of Stochasticity on the Rashomon Effect in Machine Learning
    Andrea Apicella, Francesco Isgrò, Andrea Pollastro, Roberto Prevete

    Neural network training is inherently stochastic, with factors such as weight initialization leading to distinct models despite comparable predictive performance. This phenomenon is commonly associated with the Rashomon effect, which describes the existence of multiple near-optimal models for the sa

    benchmark
  926. arxiv:2609.32933 · cs.LG
    Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs
    Nam Phuong Tran, Trinh Ha Mai Huynh, Tuyen Pham Le, Van-Truong Nguyen +2

    In safety-critical applications, deployment uses a single policy, whose performance and constraint satisfaction should hold directly rather than only for an average or mixture of training policies. This motivates last-iterate guarantees in constrained reinforcement learning. Recent progress has esta

    policy evaluation
  927. arxiv:2609.32929 · cs.LG
    Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linear Propagation
    Kiarash Banihashem, MohammadTaghi Hajiaghayi, Mahdi JafariRaviz, Silvio Lattanzi +1

    Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures withou

    benchmark
  928. arxiv:2609.32927 · cs.LG
    The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning
    Cristina V. Lopes, Yuangang Li, Justin Tian Jin Chen, Alberto Krone-Martins +2

    Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning bene

    manipulation
  929. arxiv:2609.32924 · cs.AI
    Diagnosing Sampled LLM Reasoning in Formal Geometry: Coverage, Realization, and Validity Evidence
    Xiao Yue, Guangzhi Qu

    Repeated sampling can reveal a correct numerical answer without yielding either a reliable system output or a supported derivation. We present Coverage, Realization, and Validity Evidence (CRV), an evaluation protocol for sampled large language model (LLM) reasoning over formal geometry states. Cove

    evaluation protocol
  930. arxiv:2609.32922 · cs.AI
    Precision As You Need: Stochastic Computing Is a Dense Adaptive Quantizer
    Haoran Jin, Kangqi Zhang, Jirong Yang, Barry Lyu +3

    Matrix multiplications dominate the inference cost of modern transformer-based vision models, yet existing efficiency techniques such as post-training quantization and mixed-precision inference are largely limited to the small set of fixed-width formats (INT4, INT8, BF16, and FP16) supported by conv

    post-training
  931. arxiv:2609.32921 · cs.LG
    Adaptive Latent Capacity for World Models
    Idan Achituve, Lior Dikstein, Idit Diamant, Arnon Netzer +1

    We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over

    world model
  932. arxiv:2609.32918 · cs.RO
    Vision-Language Agents for Active Perception in Optics Laboratories
    Ryan Lopez, Sachin Vaidya, Seou Choi, Serena Landers +1

    Vision-language models (VLMs) are increasingly being used in scientific workflows, but their ability as agents to directly control laboratory experiments from visual feedback remains underexplored. This capability is important because many laboratory tasks do not naturally provide dense, pre-defined

    agent
  933. arxiv:2609.32917 · cs.AI
    Planner-as-Router: Joint Plan-Time Model Routing for Cost-Efficient Multi-Agent Workflows
    Vivek Kumar Singh, Preeti Priyam, Gautam Bhowmick

    Running large language model (LLM) agents in production gets expensive fast. A frontier model (the largest, most capable tier) is accurate but can cost 25 times what a small model costs per token, and the gap compounds once a workflow chains several calls together. Planner-as-Router (PaR) attacks th

    multi-agentagenticbenchmark
  934. arxiv:2609.32915 · cs.AI
    AgentTell: Behavioural Side-Channel Leakage in Browser-Use Agents
    Asif Shahriar, Md Nafiu Rahman, Sadif Ahmed, Farig Sadeque +1

    Browser-use agents often carry information in their context as they move between websites. While it may be necessary for task completion, it also creates a privacy risk, especially when the information contains a private fact regarding the user. For example, an agent may learn a user's affiliation a

    memoryagentbenchmark
  935. arxiv:2609.32912 · cs.AI
    Progression- vs Automata-based Anticipatory Monitoring of LTL over Finite Traces (Extended Version)
    Sarah Winkler, Toryn Klassen, Sheila McIlraith, Marco Montali

    When safety-critical systems are developed from a known internal specification, their correctness can be established by model checking. In the frequent case where such a specification is unknown or inaccessible, runtime verification presents an attractive alternative, e.g., to ascertain that autonom

    agentic
  936. arxiv:2609.32902 · cs.CL
    Linger and Lose: Knowledge Collapse in Low-Bit Language Models
    Prashanna Mani Paudel, Shivanand Venkanna Sheshappanavar

    Training language models with ternary weights is commonly judged by loss and downstream accuracy, which record only a modest cost relative to full precision. We show that these metrics can conceal a much larger failure. We instead measure knowledge capacity, the factual bits stored per parameter, on

    post-training
  937. arxiv:2609.32900 · cs.AI
    Constraints Are Graphs, Not Chains: Exact Decoding for Diffusion Language Models
    Jianchang Su, Wei Zhang

    Diffusion language models (dLLMs) predict masked positions in arbitrary order, but their exact constrained decoders still encode constraints as sequential languages, whose state must track every unresolved dependency between positions. For relational constraints this encoding grows exponentially: fo

    benchmark
  938. arxiv:2609.32898 · cs.LG
    When Less Compute Is More: Adaptive Early Exit Improves Pretrained Outlier Detection
    Tianyang Zhou, Leman Akoglu

    Pretrained tabular foundation models process every dataset at a fixed depth, with inference costs growing with dataset size. To address this, we present the first study of depth-adaptive early-exit for pretrained outlier detection models. While early-exit is typically motivated by efficiency, we unc

    benchmark
  939. arxiv:2609.32894 · physics.app-ph
    Composition-Driven Metal-to-Semiconductor Transition and Enhanced Phonon Transport in B-C substituted Clathrate
    Ghulam Hussain, Dario Massa, Rajibul Islam, Magdalena Birowska +2

    Establishing chemical design rules that simultaneously control the electronic structure and thermal transport is a long-sought goal for heat-management and energy materials. Here, we demonstrate that a single B-to-C substitution changes the electron count and simultaneously reconstructs the bonding

    benchmark
  940. arxiv:2609.32890 · cs.LG
    A Function-Level Vulnerability Score Measures Flag Rate More Than the Model: Protocol Effects on Paired Benchmarks
    Maciej Cichoń, Bartłomiej Dmitruk

    Language models are increasingly evaluated as vulnerability detectors, and scores reported for similar models differ widely between papers. We measured how much of that difference evaluation protocol accounts for, with model outputs held fixed. In a paired test, a model must flag a vulnerable functi

    benchmarkevaluation protocol
  941. arxiv:2609.32886 · cs.AI
    StraTune: Adaptive Selection of Revision Operators for Self-Evolving LLM Skills
    Zeping Liu, Yan Li, Ni Lao, Gil Wolff +1

    Large language models (LLMs) can learn reusable textual skills from execution feedback without updating their parameters, but effectively deciding how to revise these skills remains a key challenge. Existing methods typically rely on a fixed revision operator, a search strategy and the revision form

    self-evolvingbenchmark
  942. arxiv:2609.32885 · cs.AI
    Can LLMs Predict the Future? A Brier Score Analysis of Prediction Markets
    Yuanbo Li, Zekun Li, Xiaoyan cong

    We study whether model upgrades improve probability estimates for prediction-market questions. Our Resolved Market Forecasting (RMF) benchmark contains 3,000 resolved binary questions across nine domains, on which we evaluate six Claude and Qwen model variants using a question-only, zero-shot protoc

    benchmark
  943. arxiv:2609.32876 · cs.LG
    Multimodal LLMs Outperform Pathology Foundation Models in Cross-Domain Histological Similarity
    Yishu Zhang, Yun Li, Daiwei Zhang

    State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outper

    benchmark
  944. arxiv:2609.32874 · cs.LG
    Machine learning for the LHC physics program: a 2025-2026 stocktake
    Jesse Thaler

    The first sentence of this abstract--and the introduction to these proceedings--was authored by a human, but the bulk of this document was generated by an agentic AI system. In this talk, I take stock of machine learning (ML) for the LHC physics program over the twelve months from May 2025 to May 20

    ai agentagentic
  945. arxiv:2609.32870 · cs.LG
    Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning
    Xing Han, Yuxin Wang, Chen Chen, Wei Dai +7

    Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be

    memoryself-playself-evolving
  946. arxiv:2609.32868 · cs.AI
    ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations
    J. Paul Liu, Uthpala Herath, Andrew Petersen

    Traditional scientific computing requires researchers to translate computational intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler state and application logs. We present ASCEND (Autonomous Scientific Computing Engine and Novel Discov

    agentai agentbenchmark
  947. arxiv:2609.32867 · cs.CL
    Are You Sure You're Sure? Two Confounds in a Sycophancy Benchmark
    Atharv Gupta, Akshat Jindal, Lavanya Nigam, Aryan Sood

    Sycophancy is a language model's tendency to cave when a user pushes back, abandoning a correct answer for the user's. Several benchmarks now measure it by scripting an objection and recording how often the model caves. Because that objection is a prompt template, whatever else the template varies i

    benchmark
  948. arxiv:2609.32863 · cs.CV
    SV2V-RSim: A Comprehensive Benchmark for Self-Selective V2V Cooperative Perception with Near-Realistic Data
    Yulu Wu, Chao Wei, Jujun Cheng, Zhangkai Ni +5

    Vehicle-to-Vehicle (V2V) cooperative perception enhances autonomous driving by enabling vehicles to share information beyond their direct line of sight. However, existing V2V datasets are limited by a small number of participating agents, static collaborator selection strategies, and a significant d

    sim-to-realagentbenchmark
  949. arxiv:2609.32862 · cs.RO
    RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents
    Jingsong Liang, Shuhao Liao, Shizhe Zhang, Diyuan Hou +14

    A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alon

    embodiedliberomemoryagentagenticembodied agent
  950. arxiv:2609.32856 · cs.CV
    PolyTopoBench: A Benchmark for Complex Vector Polygon Generation from Remote Sensing Imagery
    Zeping Liu, Ni Lao, Weiwei Sun, Gil Wolff +4

    Vector polygon generation converts visual inputs, e.g., remote sensing (RS) images, into vectorized polygonal geometries, supporting applications such as autonomous driving, vector map construction, and remote sensing. Early pipelines predict raster masks and post-process them into polygons, which p

    benchmarkevaluation framework
  951. arxiv:2609.32855 · cs.RO
    FINE: Future-Informed Navigation Encoding for Data-Efficient Vision-Language Navigation
    Khang H. Nguyen, Hoang Pham Quang Nguyen, Ha Phuong Nguyen, Khanh Dinh Binh +5

    Adapting vision-language navigation (VLN) policies to new environments is expensive because every additional route and instruction requires an embodied demonstration. Yet standard observation-to-action training uses only a small fraction of the information already contained in each trajectory. In pa

    embodiedworld modelagent
  952. arxiv:2609.32852 · cs.CL
    ARSM: Auto-Regressive State Machine for Agentic Reasoning Compression
    Xiafeng Man, Siyuan Ye, Xiaosong Ma

    While Large Language Model (LLM)-based agents demonstrate strong capabilities in long-horizon tasks by interleaving reasoning with external environment interactions, the continuous accumulation of context rapidly creates a critical memory bottleneck. Existing memory compression methods rely on task-

    memoryautonomous agentagentic
  953. arxiv:2609.32846 · cs.LG
    SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals
    Yavuz Yarici, Ghassan AlRegib

    Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal

    benchmark
  954. arxiv:2609.32838 · cs.LG
    HamiFormer: Dual-Expert Diffusion Fields with Affine Symplectic Maps
    Haoxiang Huang, Xiang Liu, Shuwei Wang, Jingheng Ma +1

    Predicting smooth dynamics and collisions requires modeling continuous evolution and abrupt state changes. We introduce HamiFormer, a dual-expert diffusion field combining whole-window denoising with residual-corrected Hamiltonian propagation. Their mixed-state feedback attenuates the direct contrib

    iterative refinement
  955. arxiv:2609.32837 · cs.RO
    Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation
    Xuesong Li, Shuai Chen, Feng Li, Zhongliang Jiang +2

    Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditi

    world modelscene graph
  956. arxiv:2609.32835 · cs.AI
    FinancialAuditBench: Benchmark Construction under Differential Privacy Using Real-World Priors
    Jerry Huang, Sarvesh Babu, Matt Van Buren, Alexander Wang +4

    As AI agents are becoming widely adopted in the financial services industry, careful measurement is essential to understand where they can be reliably deployed and where oversight and professional review remain necessary. Such measurement, however, is constrained by limited access to proprietary or

    ai agentbenchmark
  957. arxiv:2609.32831 · cs.LG
    UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
    Wanqi Yang, Yuexiao Ma, Mei Xie, Xiawu Zheng +1

    Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods ar

    long-context
  958. arxiv:2609.32827 · cs.AI
    Improving LLM Collaboration via Multi-Agent Preference Learning
    Shuo Liu, Xinzichen Li, Tianle Chen, Christopher Amato

    Several works have explored multi-agent reinforcement learning (MARL) in LLM collaboration. However, constructing reliable rewards is difficult in practice, as complete and accurate metrics are often unavailable and hard to aggregate. Preference learning provides an alternative by learning from comp

    agentmulti-agentagent systemtool use
  959. arxiv:2609.32825 · cs.AI
    The Decomposition Tax: LLM Pipelines Lose Up to 40 Accuracy Points at Their Own Interfaces
    Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao

    A four-stage LLM pipeline gives up as much as 40.5 accuracy points at its own interfaces (gemma-3-12B on MATH-500, Holm-corrected p = 1.66e-19; the largest tax in the primary family). We hold model, problem, stages, stage prompts and completion budget fixed, vary only whether each stage can still se

    benchmark
  960. arxiv:2609.32813 · cs.LG
    USAI-Quant: A Quantitative Reasoning Benchmark for Vision-Language Models in Built Environments
    Dongdong Wang, Qingqi Song, Yuzhou Chen, Deepak Balakrishnan +2

    Large vision-language models (VLMs) have emerged as a powerful paradigm for urban and spatial AI. However, current state-of-the-art large VLMs still struggle with quantitative reasoning on remote sensing imagery. Existing benchmarks and algorithms are predominantly based on qualitative Visual Questi

    benchmark
  961. arxiv:2609.32810 · cs.LG
    OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories
    Anqi Li, Zhixuan Ge, Yixuan Duan, Jiarong Qian +8

    Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchmark with 611 patient cases and 19,157 discussion turns across

    benchmark
  962. arxiv:2609.32809 · cs.AI
    Overwhelmed by Choice: Studying LLM Decision Making at Scale
    Yu-Chi Lin, Aryan Seth, Anshul Aravind, Eugene Lee +3

    Multiple-choice and candidate-selection evaluations are widely used to assess LLM reasoning and decision-making, yet most benchmarks contain relatively small candidate sets. It remains unclear whether conclusions drawn from these settings remain valid as the candidate space scales. We systematically

    long-contextbenchmark
  963. arxiv:2609.32805 · cs.LG
    Decision-Sufficient State Representations: Measuring and Reducing Write-Time Regret
    Bingyu Shen, Boyang Li

    Long tasks produce more history than an LLM agent can hold in its context, and more than it uses reliably even when the history fits. A growing line of work therefore has agents carry a short written state instead: at every step a writer rewrites the state, and a reader acts from the state alone. St

    memoryagentllm agent
  964. arxiv:2609.32803 · cs.AI
    Nutri-ATLAS: Embodied Agent for Tabulated Lookup and Assistance for Smarter nutrition
    Uttej Kallakuri, Boxun Hu, Ankur A. Butala, Najim Dehak +1

    Generative and Agentic IoT systems offer a promising foundation for digital healthcare applications that combine sensing, personalized reasoning, and autonomous interaction in real-world environments. Nutrition assistance is a natural use case, but existing Large Language Model (LLM)-based systems a

    embodiedknowledge graphagentagenticembodied agent
  965. arxiv:2609.32802 · cs.LG
    Re-derivability Decides What a Staged Agent Pipeline Recovers After an Upstream Fault
    Tianqi Bu, YuXuan Peng, Junteng Tu, Henghui Xiao

    One variable sets what an upstream fault costs a staged pipeline of language-model agents: re-derivability, how much of what a stage needs it can rebuild from the original problem. Grounding an inspector agent in that problem is worth +0.608 [+0.517, +0.700] to +0.358 over a blind one on four open-w

    agent
  966. arxiv:2609.32801 · cs.RO
    PlanGuard: A Guardrail for Multi-Step Plan Safety in Embodied Agents
    Junchi Chen, Changtao Miao, Yuxiao Xiang, Zhenchao Jin +8

    Embodied task planners may produce multi-step plans whose subtask dependencies and interactions with the environment create physical risks during execution. Yet existing safeguards overlook such compositional risks, as general-purpose guardrails focus on semantic harm and embodied safety detectors a

    embodiedembodied agent
  967. arxiv:2609.32798 · cs.MA
    Adaptive and Resilient Dual-Layer Resource Slicing for Hovering Aerial Backhaul Networks
    Chuan-Chi Lai, Jen-Hsiang Li

    This paper investigates adaptive and resilient dual-layer resource slicing in hovering aerial agent (HAA)-assisted backhaul networks for heterogeneous 5G/6G services, including enhanced mobile broadband (eMBB), ultra-reliable and low-latency communications (URLLC), and massive machine-type communica

    agent
  968. arxiv:2609.32795 · cs.AI
    AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks
    Woojung Song, Hoyeol Yang, Jeonghoon Shim, Sungjib Lim +3

    Large language model (LLM) agents assist users with everyday tasks that can be completed in many reasonable ways. Even when their answers are useful, how agents carry out these tasks may not match users' preferences and needs. For example, agents differ in whether they ask clarifying questions or se

    agentbenchmark
  969. arxiv:2609.32792 · cs.LG
    Understanding and Exploiting Anisotropy in Post-Training
    Samyak Jha, Harshvardhan Saini, Yizhen Liao, Yiming Tang +1

    LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry dispro

    post-training
  970. arxiv:2609.32791 · cs.AI
    $T^5$: Twin-Critic Training for Token-Level Thoughts in Reinforcement Mid-Training
    Nan Qiao, Yebin Yang, Weinong Wang, Shuning Wang +9

    Reinforcement mid-training lets language models learn internal thoughts from unlabeled text, but efficient token-level credit assignment remains challenging. Existing group-relative methods require costly repeated generation. Learned critics offer single-rollout feedback, but accurate return predict

    benchmark
  971. arxiv:2609.32790 · cs.LG
    FedHV: Low-Overhead Hypervolume Weighting for Federated Multi-Objective Optimization
    Amirardalan Dehghanpour, Seyed Mohammad Azimi-Abarghouyi, Christopher G. Brinton

    Task-wise federated multi-objective optimization (FedMOO) trains a shared model for competing prediction objectives under heterogeneous data, partial participation, and communication constraints. Existing methods commonly derive task weights from gradient or update geometry. This requires task-speci

    benchmark
  972. arxiv:2609.32789 · cs.CV
    Hierarchical Frequency-Domain Compression of Implicit Geometric Representations for Large-Scale Point Clouds
    Manlin Yao, Jiabin Liu, Guan Wang, Haixu Liu +1

    Large-scale point cloud representations of complex geome tries incur prohibitive computational and memory costs, necessitating compressed implicit representations. To ad dress this, we propose a unified framework comprising im plicit geometric field representation, hierarchical frequency domain comp

    memory
  973. arxiv:2609.32787 · cs.AI
    CLAIRE: A Schema-Grounded Hybrid Workflow for Healthcare Administrative Form Completion
    Garapati Keerthana, Manik Gupta

    Healthcare administrative staff transfer structured information from electronic health records, referrals, claims systems, provider rosters, and work queues into dynamic forms. We developed and evaluated CLAIRE (Clinical Language and Agentic Intelligence for Reasoning and Entry), a hybrid workflow t

    agenticbenchmark
  974. arxiv:2609.32785 · cs.LG
    Learning When to Recur: Token-Adaptive Recursion for Imbalanced Ophthalmic Domain Incremental Learning
    Nanxi Yu, Kang Li, Ye Du, Xiaowei Hu +2

    Domain incremental learning is essential for adapting ophthalmic deep learning models to sequential clinical domains while preserving diagnostic expertise. Existing domain incremental learning methods predominantly address the domain shift induced by style variations. However, they often overlook th

    benchmark
  975. arxiv:2609.32783 · cs.RO
    CollisionGAT: Controller-Agnostic One-Step Collision Screening for Multi-Agent Motion
    Alan Debbas, Edwin Meriaux, Gregory Dudek

    Before a team of robots moves, each proposed step must be checked for collisions with other robots and with obstacles. We present CollisionGAT, a graph-attention network that reads the current and proposed states of moving agents together with locally relevant stationary obstacles and returns one co

    multi-agent
  976. arxiv:2609.32781 · cs.LG
    CT-OPD: Counterfactual Trace On-Policy Distillation for Diffusion Vision-Language Models
    Long Qian, Bingke Zhu, Jiaqi Wei, Yu Li +2

    Diffusion vision-language models generate answers by gradually resolving masked tokens, making accurate conditional prediction in partially resolved states central to post-training. Masking completed answers yields coherent contexts and targets, but prescribed masks do not reflect the model's reveal

    post-trainingbenchmark
  977. arxiv:2609.32780 · cs.LG
    OmniMoE-VL: A Sparse Vision-Language Model with Coupled Visual-Depth Routing
    Long Qian, Bingke Zhu, Jiaqi Wei, Yingying Chen +1

    Vision-language models (VLMs) increasingly use sparse mixture-of-experts (MoE) to scale language-side computation, yet visual information is typically routed only after passing through a fixed cross-modal interface. This leaves an important decision unresolved: which intermediate visual representati

    benchmark
  978. arxiv:2609.32779 · cs.RO
    Copper-Policy: Focus on the Representation for Robust Robot Manipulation
    Zexin Feng, Yixu Feng, Lingyu Xiao, Shang Su +5

    World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve c

    embodiedmanipulationliberorobotwin
  979. arxiv:2609.32778 · cs.AI
    Agentic Network Traffic Monitoring
    Manuel Tsoukatos, Hayden Jananthan, Jeremy Kepner

    As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way that is aligned with the users intent. Auditing an agent's network traffic provides a clear record of

    agentai agentagentic
  980. arxiv:2609.32776 · cs.CL
    On the Behavioral Traits of LLM Agents
    Haokai Zhao, Jie Gao, Yunze Xiao, Xintao Wang +4

    Users increasingly describe different AI agents as distinct colleagues to work with. AI personality research aims to quantify such impressions by attributing human-like "traits" to agents. However, existing measures fall short: models' self-reports (S-data) diverge from their actual behavior, while

    agentai agentllm agent
  981. arxiv:2609.32771 · cs.LG
    Continual Learning via Self-Probe Gradients
    Dongkyu Cho, Rumi Chunara, Sungmin Cha

    Adapting pretrained models to new data can cause catastrophic forgetting of previously learned behavior. When only a few past samples remain, they give continual learning methods sparse and narrow evidence about what to preserve. We show that language models can expand this evidence through self-pro

    benchmark
  982. arxiv:2609.32770 · cs.CL
    C-HAT-Bench: Benchmarking Chinese AI-Text Detection Beyond Fully Generated Text
    Qing Yang, Zixiang Luo, Zhenyu Mao, Zezheng Wu +5

    Large Language Models (LLMs) increasingly participate in writing by modifying or extending human drafts, causing machine involvement to vary in both form and extent. Yet most Machine-Generated Text (MGT) detectors are evaluated only on fully human-written versus fully AI-generated text. Because huma

    benchmark
  983. arxiv:2609.32767 · cs.RO
    CLAP: Closed-Loop Alignment with Pressure for Precise Suction Manipulation
    Yixian Zou, Chongyang Xu, Yuling Xin, Ziliang Feng +2

    Stacking and palletising demand precise placement: error left in one layer is inherited by the next, and a flat pad offers no feature to funnel a wrong pose into the right one. Top-down suction suits such dense arrangements, and suction has already been brought into vision-language-action (VLA) poli

    vision-language-actionmanipulationgrasp
  984. arxiv:2609.32766 · cs.LG
    Structuring Relations Among Learning Paradigms via Protocol--Objective--Resource Reductions
    Junwei Su, Changjie Wang, Dongyang Chang

    Modern machine learning spans supervised, transfer, continual, meta-learning, and related regimes that often reuse the same hypothesis classes, architectures, and optimizers but differ in information access, objectives, memory, adaptation, and sample accounting. This makes it difficult to determine

    memory
  985. arxiv:2609.32763 · cs.AI
    Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them
    Yicheng Bao, Zhenkun Gao, Xiahui Guo, Mingqian Yang +6

    Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts,

    manipulationbenchmark
  986. arxiv:2609.32762 · cs.RO
    An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning
    Mino Nakura, Sriram Krishna, Yufei Wang, Shubham Tulsiani +2

    Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study o

    manipulationaction head
  987. arxiv:2609.32761 · cs.CV
    From Feed-Forward to Flow: Unifying Reconstruction and Generation Is Easier Than You Think
    Haoru Wang, Qianfan Shen, Kai Ye, Wenzheng Chen +1

    Reconstruct where the images provide evidence, and generate where they do not: recent success of spatial world models such as Atlas (World Labs Team, 2026) highlights the value of unifying reconstruction and generation in one model. Yet the two have long lived in separate paradigms with distinctive

    world model
  988. arxiv:2609.32759 · cs.LG
    The Extender: A Log-Structured Transformer
    Jakob Eriksson

    We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual $\mathbf{h}$, a superposition channel. The Extender adds a concatenation channel $\mathbf{x}$: each lay

    memorylong-context
  989. arxiv:2609.32757 · cs.LG
    Readout is not Recovery: Dissociating Coordinate Emission from Visual-Corruption Repair in Vision-Language Models
    Drandreb Earl Juanico

    VLM bounding-box localization is both language generation and spatial commitment. Parseable fields such as bbox_2d make localization easy to score, but dimensions that emit coordinate tokens need not repair localization after visual evidence is damaged. We study this readout/recovery separation in Q

    benchmark
  990. arxiv:2609.32756 · cs.LG
    Reuse or Relearn? A Spectral View of Earth Observation Foundation Models
    Mehmet Ozgur Turkoglu, Valerio Marsocci, Dominik J. Mühlematter, Dominik Senti +2

    Foundation models are rarely used as generic, frozen feature extractors; instead, they are fine-tuned for the target downstream application. This practice is particularly prevalent in Earth observation (EO), and it raises a question that downstream accuracy alone cannot answer: does fine-tuning reus

    benchmark
  991. arxiv:2609.32750 · cs.AI
    CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning
    Xin Yan, Zhengbo Jiao, Jiaqi Liu, Zhenglin Wan +10

    Reinforcement learning enables computer-use agents to improve through interaction with real software environments, including websites and desktop applications. However, conventional deployments replicate an initialized runtime for each independent rollout, even when trajectories use the same softwar

    memoryagentevaluator
  992. arxiv:2609.32749 · cs.LG
    Retrospective Distillation Attribution via Normalized Response Similarity
    Minwoo Jang, Jaechang Kim, Minhyeon Oh, Jeongyeon Hwang +1

    Model distillation transfers capabilities through supervised fine-tuning (SFT) on teacher responses, often collected from commercial APIs, raising questions of model provenance. Existing distillation attribution methods have been largely evaluated on students immediately after the SFT step. However,

    post-training
  993. arxiv:2609.32747 · eess.SY
    Path Invariance of a Quadrotor System under Cyber Attacks with Theoretical Guarantees
    Hamza Mahmood, Usman Ali, Adeel Akhtar

    This paper presents a path-following controller for a quadrotor system to guarantee safe maneuvers, in terms of forward path invariance, in the presence of cyber-physical attacks. We assume that an adversarial agent can control any one of the rotors through a false data injection (FDI) type of attac

    agent
  994. arxiv:2609.32746 · cs.LG
    Self-Evolving Multi-Agent Symbolic Discovery for Financial Fundamental Analysis
    Kelvin J. L. Koa, Filip Orestav, Shengqiong Wu, Michael J. Wooldridge +1

    While symbolic regression (SR) has been successfully used in science to discover new equations, its use in financial valuation is hindered by several limitations. Whereas the natural sciences provide objectively correct relationships, financial valuation constitutes a distinct class of symbolic disc

    memorymulti-agentagent frameworkself-evolving
  995. arxiv:2609.32745 · cs.RO
    MORPH: Self-Organising Multi-Robot Task Allocation via Neuroplasticity-Inspired Adaptive Topology
    Xuezhi Niu, Didem Gürdür Broo

    Multi-robot task allocation (MRTA) in dynamic environments faces a fundamental tension: effective coordination requires learned structure, but that structure must adapt when conditions change. Existing methods resolve this by assuming prior task knowledge, a utility function, a cost matrix, or a tra

    multi-agentbenchmark
  996. arxiv:2609.32743 · cs.LG
    Benchmarking EEG Foundation Models at Scale: Lessons from 20,000 Evaluations
    Zhige Chen, Shu Peng, Chengxuan Qin, Rui Liu +3

    Electroencephalography (EEG) foundation models (FMs) promise transferable neural representations, yet their advantages over strong supervised baselines and their prospects for further scaling remain unclear. To address these questions, we introduce EEG-Arena, an open-source benchmark covering 30 EEG

    benchmarkevaluation framework
  997. arxiv:2609.32740 · cs.CV
    AnesTRACE: Benchmarking Intraoperative Anesthesia from Multimodal Perception to Multi-step Decision-Making
    Ziwei Huang, Qi Gao, Zhe Ji, Yuanyuan Yao +4

    Intraoperative anesthesia requires systems to interpret evolving multimodal evidence, recommend timely management, and revise decisions as patient states change, yet existing benchmarks usually isolate perception or single-point reasoning. We introduce AnesTRACE, an evaluation suite comprising AnesT

    benchmarkevaluator
  998. arxiv:2609.32737 · cs.LG
    Gradient-Guided Decoupled Adaptation for Geospatial Vision-Language Models
    Dongdong Wang, Deepak Balakrishnan, Ravi Srinivasan, Shenhao Wang

    Existing geospatial vision-language models (Geo-VLMs) typically optimize diverse geospatial tasks through a unified multi-task adaptation paradigm without explicitly accounting for the heterogeneous optimization characteristics. Our empirical observations reveal heterogeneous gradient characteristic

    benchmark
  999. arxiv:2609.32734 · cs.CV
    REALIS: A Curated Dataset for Studying the Challenges of AI Image Detection
    Aleksandr Gushchin, Khaled Abud, Georgii Bychkov, Ekaterina Shumitskaya +4

    AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for ev

    benchmark
  1000. arxiv:2609.32731 · cs.AI
    SkillVine: Agent Skill Evolution via Branching Exploration
    Kaiwei Liu, Jiqian Dong, Liran Dong, Shuai Mao +6

    Agent skills encapsulate reusable procedural knowledge that enables LLM agents to perform tasks, and they can be improved automatically using trajectories from interactions with the environment. This is the classic problem of skill evolution. Existing approaches predominately follow a linear evoluti

    agentllm agentbenchmark
  1001. arxiv:2609.32720 · cs.LG
    BiasReducer: Adaptive Bias Mitigation for Reward Models
    Shuang Liu, Yongliang Miao, Yanguang Liu, Haoyi Xiong +1

    Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods ei

    benchmark
  1002. arxiv:2609.32712 · cs.AI
    MassAlloc Attention: Let Attention Allocate Its Own Compute
    Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi +5

    FullAttn often assigns negligible normalized mass to much of the causal score space, yet dense kernels execute the complete post-score path after forming each QK tile. We introduce MALA, a fused attention primitive that preserves score access to every legal causal interaction and uses normalized con

    long-contextbenchmark
  1003. arxiv:2609.32706 · cs.AI
    Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models
    Jeongho Yoon, Chanhee Park, Yongchan Chun, Duong Tuan Thanh +4

    Cloud-based large language models (LLMs) require users to disclose plaintext data to service providers, creating privacy risks in sensitive domains. Existing privacy-preserving approaches often trade utility for protection, incur substantial computational or communication overhead, remain vulnerable

    tool calling
  1004. arxiv:2609.32704 · cs.AI
    CoWindow Attention: Full Causal Coverage Is a Collective Property
    Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi +5

    FullAttn repeatedly exposes the complete causal history to every attention head, creating substantial redundant computation and memory traffic even with IO-efficient dense kernels. We introduce CoWA, a structured attention architecture that distributes access to the causal history across KV heads. A

    memorylong-contextbenchmark
  1005. arxiv:2609.32700 · cs.AI
    CAIRN: Dynamic Fact-Intent DAGs for Multi-Agent Exploration
    Zuyao Xu, Yuyang Jia, Junwei Guan, Xiang Li +2

    LLM-powered autonomous systems have demonstrated promising capabilities in mathematical reasoning, engineering, and cybersecurity. Yet how to organize these systems for effective, reliable, and sustained performance remains an open question. In this paper, we present CAIRN, a fact-intent-driven mult

    multi-agent
  1006. arxiv:2609.32698 · cs.RO
    SEES: A Self-Evolving Embodied System via Failure-Guided VLA Policy Adaptation
    Ziwen Li, Hanlue Zhang, Zhenyang Ren, Tianyu Huang +8

    Recent vision-language-action (VLA) policies demonstrate promising generalization across diverse short-horizon tasks. However, they remain unreliable on long-horizon tasks, partly because the large-scale training data is biased toward single-stage manipulation tasks that are cheaper to demonstrate.

    vision-language-actionvlavla policyembodiedmanipulationself-evolving
  1007. arxiv:2609.32696 · cs.LG
    Flat-Consensus Diffusion for Robust Data Reshaping under Noisy Evaluator
    Hongyu Cao, Kunpeng Liu, Fei Xie, Sandip Ray

    Data shape determines how features are structured, how patterns are separated, and how distributions cover the underlying domain. Poor data shape can make models learn noise rather than generalizable structure. This paper studies robust feature-centric data reshaping: generating feature transformati

    evaluator
  1008. arxiv:2609.32694 · cs.AI
    IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents
    Angqing Jiang, Gaoming Zhang, Chaoqun Zhang, Jianchun Song +4

    On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from

    agentbenchmark
  1009. arxiv:2609.32692 · cs.AI
    World Agent: Can Language Models Keep a World Running?
    Weixing Chen, Weipeng Zhang, Nan An, Yang Liu +1

    World models are moving from generating realistic frames to generating playable worlds, yet whether a delivered world can keep running is not tested anywhere. Existing evaluations stop at generation, at delivery, or at single-step transitions, and each stops at a different point along the way. Corre

    world modelagentbenchmark
  1010. arxiv:2609.32691 · cs.LG
    Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection
    Animesh Shaw

    LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that evaluation is rarely examined. We audit

    agentllm agentagentictool usebenchmark
  1011. arxiv:2609.32690 · cs.CV
    MM-OPD: Towards One More Bottleneck Between Perception and Reasoning
    Jintao Tong, Yujing Lou, Zhanming Shen, Jiaqi Gu +5

    Recent multimodal large language models (MLLMs) advance visual reasoning by strengthening both perception and reasoning, implicitly assuming a process that transitions seamlessly from perception to reasoning. However, we observe a counterintuitive phenomenon that challenges this assumption: holding

    benchmark
  1012. arxiv:2609.32689 · cs.LG
    Self-Evolving Time-Series Forecasting Agents with Episodic Memory and Online Policy Learning
    Junyi Wang, Yilin Wang, Wen Wu, Chao Zhang

    LLM-based agents are increasingly used for time-series forecasting because they can organise contextual information, perform multi-step analysis, and guide the sequence of actions required to complete forecasting tasks. Most existing agents focus only on the current forecasting instance. However, in

    memoryepisodic memoryagentself-evolvingbenchmark
  1013. arxiv:2609.32687 · cs.AI
    Dude, Where's My State? Execution Information Requirements for Stateful Agents
    Nikita Mehrotra, Ashish Tiwari, Priyanshu Gupta, Sumit Gulwani

    Long-running agents must preserve information that later steps depend on. We introduce the Execution Information Requirement (EIR), a lower bound on the information that must remain accessible for correct completion under specified task and access conditions. We develop LACUNA, a framework that gene

    semantic graphagent
  1014. arxiv:2609.32685 · cs.LG
    PINNMorph: Evolving Online Adaptation Policies for Physics-Informed Neural Networks
    Xu Yang, Mingyang Yu, Jun Zhang, Keqian Li +1

    Physics-informed neural networks (PINNs) provide a learning-based framework for solving partial differential equations (PDEs), yet their training behavior can change substantially throughout optimization. Residual distributions, gradient interactions, regional learning difficulty, and model-capacity

    benchmark
  1015. arxiv:2609.32683 · physics.optics
    Reconfigurable Linear Optical Transformations in a Single Integrated Multimode Waveguide
    Manoah van der Worp, Jonas Krimmer, Ivo M. Vellekoop, Pepijn W. H. Pinkse +1

    Wavefront shaping enables control over optical fields for applications ranging from imaging to photonic information processing. While conventional wavefront shaping relies on free-space systems comprising bulk optical components, compact and fully integrated approaches are comparatively unexplored.

    quantum photonic
  1016. arxiv:2609.32681 · cs.CV
    RCVLA: 4D Radar-Grounded Semantic Reasoning and Trajectory Arbitration for Autonomous Driving
    Lianqing Zheng, Xiaokai Bai, Yixuan Luo, Runwei Guan +5

    4D radar provides geometric and motion cues that complement visual semantics, but integrating it into vision-language-action (VLA) models requires both radar--language alignment for semantic reasoning and explicit use of radar measurements for trajectory refinement and selection. To support these ca

    vision-language-actionvla
  1017. arxiv:2609.32679 · cs.LG
    The GUI Is Not the State: Diagnosing State Aliasing in GUI World Models
    Dongsheng Liu, Chao Jin, Wenkui Yang, Hejin Wang +6

    GUI World Models (GUI-WMs) are increasingly used to predict future states for agent planning and simulation, yet most existing formulations condition only on the current GUI observation and action. We identify state aliasing, where the vis- ible interface omits transition-relevant environment state,

    world modelagentbenchmark
  1018. arxiv:2609.32677 · cs.LG
    When Better Gets Worse: Improvement Fidelity for Self-Improving Agents in Adaptive Worlds
    Ke Wang, Zijie Zhao, Zhiyi Yuan, Changlun Li

    Self-improving agents increasingly rely on proxy verifiers to choose policy updates, yet deployment can change the world in which those updates are evaluated. An update that looks better to the verifier can therefore become worse after deployment even when the verifier ranks policies well overall. W

    self-improvingself-improvement
  1019. arxiv:2609.32674 · cs.AI
    Expected Reasoning-Step Return Unifies On-Policy Learning from Rewards and Teachers
    Qiangqiang He, Jin Li

    On-policy reasoning models can learn from task rewards or teacher signals, but these sources differ in form and can favor conflicting updates, leaving unclear which should guide a given reasoning action. We introduce \textbf{Expected Reasoning-Step Return (ERSR)}, which treats semantic reasoning ste

    benchmark
  1020. arxiv:2609.32671 · cs.CV
    Concepts Complement Dense Semantics: Learning Compact Sparse Spaces for Text-Image Retrieval
    Yoonseo Kim, Jungwoo Choi, Cheonyoung Park, Youngwook Kim +2

    Cross-modal retrieval has been advanced by vision-language pre-trained models that encode images and texts into a shared dense embedding space. While dense representations effectively capture overall semantic similarity, they often obscure fine-grained visual-textual information needed for precise c

    grasp
  1021. arxiv:2609.32663 · cs.LG
    Distance-KV: Exploiting Relative Distance for Efficient Long-Context Inference
    Xianpeng Shang, Canbin Huang, Jiang Li, Tian Lan +3

    The memory usage and decoding latency of LLM inference grow rapidly with context length. To reduce these costs, key-value (KV) cache compression methods selectively retain cached states based on token importance or differences in attention patterns across heads. However, we discover that retrieval c

    memorylong-contextlong contextbenchmark
  1022. arxiv:2609.32662 · cs.AI
    AsynCodeBench: Benchmarking Collaboration of Asynchronous Multi-Agent Systems in Software Engineering
    Kaituo Zhang, Zhen Xiong, Zhimeng Jiang, Mingyu Zhong +7

    Multi-agent coding has emerged as an increasingly active direction in software engineering, where complex development tasks are decomposed across multiple specialized agents working on different parts of the problem. Despite the shift from individual problem solving to distributed collaboration, mul

    multi-agentagent systembenchmark
  1023. arxiv:2609.32661 · cs.LG
    Equivariant Neural Primal-Dual Assignment for Maximum Common Edge Subgraphs
    Jiaqing Xie, Yanchao Li, Zhuo Yang, Yuxin Wang +2

    Maximum common edge subgraph (MCES) matching finds a partial vertex correspondence between two labeled graphs that preserves as many labeled edges as possible. Molecular similarity search requires matching many graph pairs, making the cost of repeated queries important. The strongest baseline attain

    benchmark
  1024. arxiv:2609.32660 · cs.CV
    InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning
    Hanqian Li, Sirui Huang, Chen Ling, Jungang Li +11

    Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-

    benchmark
  1025. arxiv:2609.32659 · cs.LG
    Quantization-Aware Pre-Training with Constrained Empirical Weight Distribution
    Ningfeng Yang, Tor M. Aamodt

    Quantization-Aware Pre-Training (QAPT) can increase the inference efficiency of DNNs, but a problematic behaviour known as rounding boundary weight oscillation can introduce detrimental noise into the training process and significantly reduce convergence speed. While existing methods can reduce this

    memory
  1026. arxiv:2609.32658 · cs.AI
    Contract Memory Compiler: Resolve, Then Traverse
    Zhi Song, XiMing Xing, Chunhan Li, Weian Mao +8

    External memory lets language-model agents answer questions about histories too long for the answer model's context window. Updates create a harder problem than retrieving a recent fact: changing one relation can redirect a multi-hop question to records about an entity absent from the question. We s

    memoryexternal memory
  1027. arxiv:2609.32657 · cs.AI
    World Models with Predictable Long-Horizon Marginals
    Yuhao Du, Shunian Chen

    Accurate one-step predictions do not ensure that a world model's rollouts retain the data distribution. We make the model's decoded stationary law explicit by learning a decoder of a fixed Gaussian reference and constraining the behaviour-averaged transition to preserve that reference. For controlle

    world modeldreamerv3
  1028. arxiv:2609.32656 · cs.LG
    MixBench-TS: A Multivariate Time Series Forecasting Benchmark Where Channel Mixing Pays Off
    Ibram Abdelmalak, Mischa Putzke, Jungmin Choi, Tom Hanika +2

    Multivariate Time Series Forecasting (MTSF) models that mix information across channels assume that the past of one channel carries information about the future of another. Yet they are evaluated on a small fixed set of standard datasets whose cross-channel structure is rarely examined. We ask two q

    benchmark
  1029. arxiv:2609.32653 · cs.CV
    From Knowing to Abstaining: Bridging the Representation-Action Gap in Vision-Language Models
    Jialuo He, Huangxun Chen

    The ability of vision-language models (VLMs) to abstain from unanswerable questions is as important as their ability to answer answerable ones accurately. Recently, several benchmarks have emerged to evaluate and improve VLM abstention, but they have substantial limitations. First, samples often con

    benchmark
  1030. arxiv:2609.32652 · cs.LG
    Prediction Limits and Koopman Closure of Geometry-Induced Soft State Abstractions
    Mohit Kumar, Somayeh Kargaran

    We study when geometry-induced soft state abstractions admit accurate finite-dimensional linear dynamics. Each state is represented by simplex-valued coordinates obtained from class-specific Kernel Affine Hull Machine (KAHM) reconstruction scores, and a matrix is used to predict the next-state coord

    benchmark
  1031. arxiv:2609.32645 · cs.AI
    From Scene Graphs to Answers: Selective Neuro-Symbolic Reasoning for Autonomous Driving
    Yiyao Wang, Pei Liu, Fangzhou Liu, Jun Ma

    Autonomous-driving question answering requires reasoning over structured scene information, yet existing vision-language approaches largely delegate heterogeneous reasoning operations to a single neural inference process. We argue that this uniform strategy overlooks a fundamental distinction: some

    scene graph
  1032. arxiv:2609.32643 · cs.AI
    Business Compromise Detection with Agentic AI and LLM-driven Knowledge Discovery
    Diego Palma, Kyu Bin Kim, Zhen Han, Allbright Dsouza +1

    Detecting compromised business ad accounts is a challenge in digital advertising, as attackers exploit hijacked accounts to launch fraudulent campaigns. Large Language Model (LLM) agents show promise for integrity enforcement, but hallucinated mistakes on hard cases create business friction. In a st

    agentautonomous agentagenticbenchmark
  1033. arxiv:2609.32641 · cs.LG
    Extremely Fast and Compact Binary Graph Representations via Randomized Operator Sketching
    Srajan Agarwal, Megha P, Bikas C Das, Zakaria Laskar +1

    Graph neural networks typically rely on dense, floating-point node representations, which can impose substantial memory and computational costs. Binary graph hashing offers an alternative by encoding node information as compact bit strings. However, existing approaches either sacrifice global topolo

    memory
  1034. arxiv:2609.32638 · cs.AI
    Can Open-Weight Large Language Models (LLMs) Simulate Human Survey Populations? A Cross-Instrument Calibration Study
    Grandee Lee, Wang Yue

    Large language models (LLMs) are increasingly used to generate synthetic survey respondents and digital twins of real people, but whether their output preserves real human statistical structure, rather than surface plausibility, remains unresolved, and most existing evidence comes from proprietary m

    benchmark
  1035. arxiv:2609.32635 · cs.AI
    Trust the Brand, Lose Control: How Identity Hijacks LLM Agent Orchestration
    Xutao Mao, Rui Qian, Linghan Chen, Yudong Gao +4

    LLM agents now execute tasks end to end with permission to change real systems and increasingly orchestrate subagents that differ in capability and cost. Prior work treats the choice of subagent as an optimization problem. Yet the orchestrator makes this choice from the identities that subagents dis

    agentllm agentagent systembenchmark
  1036. arxiv:2609.32634 · cs.RO
    PF-RL: Progress Field Reinforcement Learning via Goal-Conditioned Value Geometry for Vision-Language-Action Models
    Yunpeng Qing, Yilun Kong, Sixu Lin, Ming Zhou +8

    Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progr

    vision-language-actionvlamanipulationliberorobotwin
  1037. arxiv:2609.32631 · cs.AI
    SWE-MILE: Asynchronous Potential-Induced Milestone Credit Assignment for Long-Horizon Software Engineering Agents
    Chaoqun Cui, Hao Zhou, Meiqi Chen, Fandong Meng +1

    Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an a

    agentevaluator
  1038. arxiv:2609.32630 · cs.CL
    ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis
    Kwangwook Seo, Dongha Lee

    Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable pro

    agentllm agentself-evolving
  1039. arxiv:2609.32626 · cs.RO
    World SLAM Model: Joint World Modeling for SLAM and Navigation
    Minghui Qin, Yijun Yuan, Weicheng Zheng, Kenan Li +6

    We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent

    embodiedworld modelmemorypersistent memory
  1040. arxiv:2609.32622 · cs.CL
    Does CoT-Pass@k Really Check the CoT? A Multilingual Mathematical Audit
    Tarık Tuna Taşaltı, Burcu Hüdaverdi, David Semedo

    Pass@k measures whether a model reaches a correct answer under repeated sampling, but never how: a lucky guess counts the same as sound reasoning. CoT-Pass@k was proposed to close that gap, adding an LLM-as-judge that must assess a solution's reasoning chain before it counts. Its value rests entirel

    benchmarkllm-as-judge
  1041. arxiv:2609.32619 · cs.LG
    Intuition vectors
    Shahar Haim, Daniel C. McNamee

    Large self-supervised vision models learn representations that support scene segmentation and the semantic decomposition of physical objects. We ask whether their representational geometry supports transfer to visual reasoning problems without any task-specific fine-tuning. We hypothesized that rela

    benchmark
  1042. arxiv:2609.32617 · cs.AI
    The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models
    Qingjia Huang, Yakai Li, Jianguo Wu, Qihang Zhou +4

    Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data

    post-trainingbenchmark
  1043. arxiv:2609.32616 · cs.AI
    "You're Right, Let Me Fix It": How LLM Agents Damage Correct Work When Falsely Accused
    Xutao Mao, Rui Qian, Longxiang Wang, Xinjian Yi +5

    LLM agents increasingly keep working after a task succeeds as they resume after compaction or take over handoffs. Their finished work keeps receiving follow-up input that sometimes falsely accuses it for later failures. We call an agent's acceptance of such a false accusation gaslight sycophancy, an

    agentllm agentagenticbenchmark
  1044. arxiv:2609.32613 · cs.LG
    Learning the Graph and the Embedding Together: Classifier-Independent Rewiring for Heterophilic Node Classification
    Harshit Kumar, Sujan Chakraborty, Priyanka Saha, Pritam Kar +1

    Graph neural networks lose much of their advantage on heterophilic graphs, where connected nodes often carry different labels. Graph rewiring is a popular remedy, but rewiring methods are usually evaluated with a single classifier, which makes it hard to tell whether the gains come from the new topo

    benchmark
  1045. arxiv:2609.32612 · cs.CV
    Levy-Driven Correspondence Estimation for Registration
    Qianliang Wu, Jiaqi Yang, Wankou Yang, Le Hui +3

    Finding reliable point correspondences is difficult when point clouds have low overlap or undergo non-rigid deformation. Iterative refinement can correct uncertain matches, but costly network evaluations limit the number of updates. We present LevyMatch, a Lévy-driven method that uses random jumps t

    iterative refinement
  1046. arxiv:2609.32600 · cs.AI
    CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering
    Prince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan +5

    Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studi

    benchmark
  1047. arxiv:2609.32597 · cs.LG
    Not Every Term Adds New Structure: Sobolev Novelty for Symbolic Regression
    Boxiao Wang, Kai Li, Yuheng Jing, Tianyi Liu +3

    Symbolic regression (SR) aims to discover compact and meaningful mathematical equations from data, but searching the vast combinatorial space of symbolic structures remains challenging. Existing methods typically guide this process using expression-level objectives, such as fitting error, which asse

    benchmark
  1048. arxiv:2609.32595 · cs.RO
    RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation
    Incheol Cho, Jintae Park, Jinkyu Kim, Jungbeom Lee +2

    Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However,

    quadruped
  1049. arxiv:2609.32594 · cs.AI
    MA-FPPO: Multi-Agent Flow-Pretrained Policy Optimization
    Guowei Zou, Haonan Chen, Haitao Wang, Beiwen Zhang +2

    Multi-agent flow policies learn cooperative behavior from fixed offline datasets, but often struggle to complete tasks in situations not covered by the offline data. In these situations, agents must both adapt to changes in the environment and coordinate with one another, yet action patterns learned

    multi-agentonline learningevaluation protocol
  1050. arxiv:2609.32592 · cs.CV
    SPACE: Sparse Predictive Attractor via Counterfactual Eviction for Streaming Video Memory
    Hongjin Niu, Weizhan Zhang, Shuo Bao, Jiahao Wang +3

    Fixed-capacity streaming video memory requires repeated eviction decisions whose effects accumulate over time. Yet existing policies are evaluated primarily in terms of retained information or downstream accuracy, leaving how repeated updates alter the futures supported by memory largely unexamined.

    memory
  1051. arxiv:2609.32591 · cs.RO
    Think Fast, Plan Selectively: Adaptive Deliberation for Efficient Data-Driven MPC
    Yi Xian Goh, Sze Jue Yang, Hao Luan

    Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well

    world model
  1052. arxiv:2609.32590 · cs.CV
    Retrieved but Not Delivered: Multimodal Memory Delivery for Long-Term Agents
    Yuhang Jiang, Qingwei Liao, Kaize Yin, Xingling Liu +2

    Work on memory for multimodal agents optimizes what is written, updated and retrieved. Between retrieval and the answer, however, is a stage that multimodal memory evaluations do not isolate: what of the retrieved memory reaches the model, and in what form. We call it delivery, and a controlled deco

    memoryagentbenchmark
  1053. arxiv:2609.32584 · cs.AI
    EMIR$^2$: Evolution-Aware Memory with Intent-Guided Multi-Round Retrieval
    Jinlan Liu, Hongliang Sun, Yong Wang, Bolin Zhang +3

    Long-term memory enables large language model (LLM) agents to leverage historical interactions for future tasks. However, existing memory systems struggle to utilize continuously evolving historical information, as they often rely on static memory representations and single-round retrieval strategie

    memoryagentllm agent
  1054. arxiv:2609.32581 · cs.LG
    HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion
    Junxiang Qiu, Zhengsu Chen, Xinting Hu, Shuo Wang +5

    Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation a

    memory
  1055. arxiv:2609.32579 · cs.LG
    DimPO: Dimensionality Reduction for Attention using Preference Optimization
    Vojtěch Lanz, Yufei Cui, Prasanna Parthasarathi

    A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal tha

    long-context
  1056. arxiv:2609.32577 · cs.LG
    Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
    Jinhao Dong, Liang Zhao, Zihao Yue, Wenhan Ma +6

    Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and

    agentagentic
  1057. arxiv:2609.32576 · cs.RO
    Assisting for Open-Ended Tasks: Goal-Oriented Shared Autonomy as a Particle Filter
    Mengxue Fu, Ethan Xu, Sam Iyer-Singh, Yinlong Dai +2

    A common approach for shared autonomy blends human inputs with autonomous assistance based on the human's likely goal. However, most existing approaches assume that a static set of possible goals is known a priori, which limits the use of such methods in unstructured assistive settings. We instead i

    manipulation
  1058. arxiv:2609.32574 · cs.AI
    CUE-Mem: Benchmarking Long-Term User Memory via Implicit Cues in Multimodal Conversations
    Yulin Hu, Yanyan Zhao, Zimo Long, Xing Fu +5

    Long-term memory is essential for multimodal agents that interact with users across sustained conversations. However, user memories are not always explicitly stated: they may also be implied by recurring background objects in images, ambient sounds in audio, or other peripheral multimodal cues. Exis

    memorybenchmark
  1059. arxiv:2609.32573 · cs.LG
    Trapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy Distillation
    Jinqian Chen, Jihua Zhu, Chang Liu

    On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated collaboration? We study federated OPD and fin

    benchmark
  1060. arxiv:2609.32570 · cs.LG
    CAESAR: Clustering via Autonomous Embedding-Space Agglomerative Reorganization
    Ilan Bacry, Rémi Devaux, Antoine Jardin

    Clustering algorithms that operate on nearest-neighbor graphs, such as FINCH (First Integer Neighbor Clustering Hierarchy), depend heavily on the quality of the embedding space they are given. However, pretrained vision and language model embeddings are not optimized for this purpose. We propose CAE

    benchmark
  1061. arxiv:2609.32569 · cs.CV
    OmniSmartHome: A Multimodal Reasoning Benchmark for Smart-Home Agents
    Jihoo Jung, Suho Yoo, Jeongsoo Choi, Hyebin Cho +4

    Smart-home assistants are expected to handle diverse, realistic requests that arise in daily life. In such interactions, users often rely on the surrounding multimodal context-pointing at objects or referring to what they see or hear, leaving their requests underspecified in language alone. Existing

    memoryagentbenchmark
  1062. arxiv:2609.32566 · cs.LG
    Compositional Objectives: Learning Structure in Structure
    Pranavchandra Vivekananda, Sumukh Bettadapura, Ajan Subramanian

    Intelligence is defined in many ways. One of these definitions defines intelligence as the pursuit of learnable novelty. However, learnable novelty can be meaningless without the ability to compose the learned structures to take action and achieve goals. Learnable novelty builds on epiplexity, which

    benchmark
  1063. arxiv:2609.32564 · cs.AI
    ProTTT: Learning to Learn Semantic User Memory with Test-Time Training
    Sejun Park, Hyoungjo Bhang, Hyein Jeong, Yohan Jo

    Personalization requires language models to capture user-specific knowledge from a growing user history. Existing context-based approaches incur increasing inference costs as user history accumulates and rely on separate retrieval or summarization stages, while parametric-based approaches often requ

    memorybenchmark
  1064. arxiv:2609.32560 · cs.CL
    How to Reduce Whisper Hallucination
    Husein Zolkepli

    Whisper is still what runs in production: one permissively licensed checkpoint, 99 languages, no per-language tuning. But it writes sentences nobody said. On 42 clips of pure room tone, whisper-large-v3 emits words on 61.9% of them and emits something on 100%. The usual response is to distil a stude

    benchmark
  1065. arxiv:2609.32550 · cs.RO
    Are Vision-Language-Action Models Robust to One-Step Observation Perturbations?
    Shojiro Yamabe, Jun Sakuma

    Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation c

    vision-language-actionvla
  1066. arxiv:2609.32547 · cs.LG
    Prioritizing Repeated LLM Evaluation for Hidden Failure Discovery
    Keita Broadwater, Akin Broadwater

    Large language models are commonly evaluated by generating a small number of stochastic responses for each prompt in a benchmark. Because inference budgets are limited, this shallow evaluation may fail to observe low-probability but operationally important failures. A prompt that produces no failure

    benchmark
  1067. arxiv:2609.32544 · cs.AI
    Porimon: An LLM-Based Pokémon Battle Agent Enhanced by Long/Short-Term Knowledge Augmented Generation
    Dongyin Zhuo, Fengjunjie Pan, Nenad Petrovic, Alois Knoll

    In this paper, we use Pokémon Battles as a case study to investigate how to improve the performance of LLM-based agents in tasks that require opponent-aware planning without additional fine-tuning. We propose Long/Short-Term Knowledge Augmented Generation (LSTKAG), a mechanism that enables LLM-based

    agent
  1068. arxiv:2609.32537 · cs.CV
    RIPE-MambaSpike: Resolution-Independent Spiking-State-Space Interfaces for Parameter-Efficient Event-Based Vision
    Md Muhiminul Islam, Shoaib Ahmed Dipu, Sayeed Shafayet Chowdhury

    Spiking-Mamba hybrids reach strong accuracy on event-based vision, but existing designs often require tens of millions of parameters. Much of that cost comes from how the spiking front-end is connected to the state-space backbone rather than from the hybrid architecture itself. In a representative m

    benchmark
  1069. arxiv:2609.32536 · cs.AI
    Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
    Yanjie Zhang, Nanchen Hu, Yushi Sun

    Audio language models can recognize spoken commands and invoke tools, but an agent must first decide whether the acoustic and conversational context warrants action. We introduce VGBench, a 1,018-item diagnostic benchmark for action-level addressedness across side-talk, self-talk, and speaker-switch

    manipulationagentpost-trainingbenchmark
  1070. arxiv:2609.32534 · cs.LG
    DepthBench: Measuring How Residual Connections Enable More Computational Depth
    Keyu Wang, Yangyi Huang, Jiale Kang, David González-Martínez +2

    Depth is a natural way to increase the computational capacity in Transformers, yet the contribution of deeper layers can diminish as depth grows larger. Recent approaches enhance normalization (\text{e.g.}, LayerNorm Scaling) or residual connections (\text{e.g.}, mHC, AttnRes) to enable better infor

    benchmark
  1071. arxiv:2609.32533 · cs.AI
    LLMAdBench: A Human Preference Benchmark for Advertising in LLM Responses
    Rui Ai, Yuqing Liu, Sitao Qiu, Yun Qiao +13

    Inserting advertisements (ads) into consumer-facing LLM output is emerging as a new business model, but there is little shared evidence on how such ad insertion should be evaluated or how it affects user preferences. We introduce LLMAdBench, a human-preference benchmark for studying advertising in L

    benchmark
  1072. arxiv:2609.32528 · cs.AI
    Fail Loudly: An Auditable Runtime for Agentic Data Analysis
    Hanxu Yan, Langxuan Deng, Zhengle Wang, Yibo Wang +1

    Large language models (LLMs) have enabled data-science agents to automate multi-step analyses over heterogeneous files. However, incorrect choices regarding data sources, scope, or statistical definitions often lead to silent errors: computations execute successfully but produce plausible yet incorr

    agentagentic
  1073. arxiv:2609.32527 · cs.LG
    AmbiModBench: Benchmarking Gene Perturbation Prediction Beyond Shared Responses
    Sikai Huang, Zhiwen Yang, Kai Yu, Jiayuan Chen +1

    Predicting cellular responses to genetic perturbations helps prioritize experiments in single-cell genomics, where exhaustive measurement is infeasible. While computational models increasingly predict these responses, three evaluation deficiencies obscure what their scores demonstrate. First, absolu

    benchmark
  1074. arxiv:2609.32525 · cs.LG
    Does Transolver really need a Transformer?
    Shizheng Wen, Siddhartha Mishra

    The widely used Transolver family of neural operators is based on physics-attention, which softly assigns the points of an unstructured mesh to a small number of slices, applies self-attention among the resulting tokens, and broadcasts the result back to the points. We provide a comprehensive empiri

    memorybenchmark
  1075. arxiv:2609.32522 · cs.AI
    Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
    Wenxu Jia, Xize Cheng, Zihan Zhang, Dongjie Fu +4

    Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational conte

    memoryagentagentic
  1076. arxiv:2609.32521 · cs.AI
    MemAgent: Learning to Manage Heterogeneous Memory Providers for LLM Agents
    Yongxian Wei, Yilin Zhao, Runxi Cheng, Xinrui Chen +4

    Current agents remain largely stateless across tasks, limiting their ability to continually improve from prior interactions and making memory essential for long-horizon agentic behavior. Existing memory methods seek to reuse past experience, but most rely on a single memory representation (e.g., tra

    memoryagent memoryagentllm agentagenticbenchmark
  1077. arxiv:2609.32520 · cs.AI
    When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
    Yanjie Zhang, Bowen Cao, Zixin Chen, Yushi Sun

    LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that con

    llm agentbenchmark
  1078. arxiv:2609.32518 · cs.CV
    UnStep: Training-Free Acceleration of Causal Video Diffusion with Fewer Steps Than Distillation
    Youssef Mansour, Enis Simsar, Fadime Sener, Markos Georgopoulos +3

    Distilling bidirectional multi-step video diffusion transformers into few-step causal models has become a common approach for streaming video generation. While these few-step students are significantly faster than the teachers they are distilled from, they remain slow for real-time generation. In th

    memory
  1079. arxiv:2609.32517 · cs.LG
    LocalProp: Neuro-Localized Memory-Efficient Backpropagation
    Diana-Nicoleta Grigore, Iuliana Georgescu, Radu Tudor Ionescu

    The current deep learning training paradigm employs end-to-end backpropagation, regardless of the training stage, i.e. pre-training or fine-tuning. However, backpropagating through the entire model is neither biologically plausible nor memory efficient, since learning inside the brain is highly loca

    memorybenchmark
  1080. arxiv:2609.32516 · cs.LG
    REFINE: A Resilient Evolution Framework for Intelligent Enterprise Alert Triage in Security Operations Centers
    Huimin Chen, Quan Long, Yanhao Wang

    Security Operations Centers (SOCs) process large volumes of alerts daily. Alert triage prioritizes high-risk threats while reducing manual review of benign alerts. LLM agents can reason over logs and threat intelligence, but struggle to keep aligned with organization-specific, rapidly evolving SOC o

    llm agentagent framework
  1081. arxiv:2609.32514 · cs.AI
    From Anomalies to Failures: Constructing Causal Error Graphs for Agentic Trace Diagnosis
    Shu-Xun Yang, Yidong Wang, Zhuoer Feng, Bosi Wen +6

    LLM-driven agents are increasingly deployed in complex applications, where long agentic traces make failures difficult to diagnose. Existing trace diagnosis methods often conflate anomalies, errors, and failures, making diagnostic targets ambiguous; they also lack structured modeling of how causally

    agenticbenchmark
  1082. arxiv:2609.32512 · cs.LG
    What Do Latent Predictive Vehicle Representations Retain? Measuring State, Geometry, and Local Response
    Enzo Nicolás Spotorno, Josafat Leal Filho, Antônio Augusto Fröhlich

    Models of vehicle dynamics learned from logged states and commands complement physics-based models, and latent world models, which predict in a learned representation, are used to plan and train controllers in other domains. Vehicle controllers are usually specified in physical terms: costs, limits,

    world modelaction-conditioned
  1083. arxiv:2609.32511 · cs.AI
    Learning from Others, Acting for You: Cross-User Memory Sharing for LLM Agents
    Jinming Hu, Haodong Zhao, Qi Jia, Die Chen +3

    Large language model (LLM) agents serving different users often solve related tasks, yet separate user histories can leave reusable experience inaccessible to other agents. Pooling memories expands access but risks transferring preferences that conflict with the receiving user's requirements. We int

    memorymemory architecturellm agent
  1084. arxiv:2609.32499 · cs.AI
    Learning an Anchored Prompt Space for Continual Adaptation of Large Language Models
    Rongguang Ye, Zhan Zhuang, Yichen Wu, Ming Tang +1

    Continually adapting large language models requires acquiring new knowledge while preserving previously learned capabilities. Jointly adapting model parameters and task-specific soft prompts offers a promising solution, but faces two key limitations: historical prompts may become less effective as t

    benchmark
  1085. arxiv:2609.32498 · cs.AI
    DAAF: From Failure Localization to Editable System Assets in LLM Agents
    Xiaoyang Yuan, Qi Liu, Yubin Ruan, Xinyi Mou +8

    Deployed LLM agents increasingly rely on persistent, versioned system assets such as routing rules, knowledge segments, prompt instructions, and reusable skills. Failure-localization methods can identify where an error manifests in an agent or execution trace, but repair requires a different decisio

    agentllm agent
  1086. arxiv:2609.32496 · cs.CL
    Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning
    Bohao Chu, Hendrik Damm, Qianli Wang, Hui Wang +3

    Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound yet globally insufficient. Local soundne

    benchmark
  1087. arxiv:2609.32495 · cs.AI
    Hearsay: Can an Auditor Trust the Record a Deployed Agent Harness Writes?
    Jiahong Dai, Zhuochen Yang, Pengyang Shao, Kelvin Ng +4

    An agent harness, the code that turns a model into an agent, writes its own record of each run, and that record is all a later reader gets when a run is disputed, investigated or audited. We call a record evidentiary when a reader who was not there can check it without trusting the writer. Across si

    agentbenchmark
  1088. arxiv:2609.32493 · cs.LG
    SoFT: Soft Targets for Generalizable LLM Fine-Tuning
    Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi +6

    Distillation enables student language models to acquire new capabilities from expert teachers. However, integrating knowledge from multi-teacher, multi-domain demonstrations into a single student remains challenging. We study supervised fine-tuning (SFT) in this setting, where students must acquire

    agentic
  1089. arxiv:2609.32492 · cs.AI
    Beyond Prompt or Skill? Attribution-Guided Optimization of Modular LLM Programs
    Haoran Shou, Haoyue Liu, Yu Huo, Kun Zeng +1

    Large language models can solve increasingly diverse reasoning tasks, yet their performance remains highly sensitive to task prompts, intermediate instructions, and the way reusable problem-solving knowledge is incorporated. Existing optimization methods usually focus on only one part of this design

    benchmark
  1090. arxiv:2609.32490 · cs.AI
    RepoMAS: Solving Progressively Specified Tasks with Issue-Driven Multi-Agent Systems
    Yuchen Song, Andong Chen, Wenxin Zhu, Muyun Yang +1

    LLM-based multi-agent systems (MASs) have shown strong potential for solving complex tasks, but most assume that task requirements are sufficiently specified before execution. In practice, user requests are often incomplete, and additional requirements may only become clear during reasoning, tool us

    multi-agentagent frameworkagent systemtool usebenchmark
  1091. arxiv:2609.32488 · cs.LG
    When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections
    Maojun Sun, Yancheng Yuan, Jian Huang, Ruijian Han

    Dense retrieval powers retrieval-augmented generation, semantic search, and question answering, yet the theoretical basis for choosing between shared and dual query-document projections remains unclear. We introduce a bias-variance theory for low-rank bilinear scoring. Shared projections induce posi

    retrieval-augmented
  1092. arxiv:2609.32485 · cs.LG
    What Should Federated LoRA Share? FedSAIL via Input-aware Subspace Alignment
    Junye Du, Shuaida He, Long Feng

    Federated low-rank adaptation (LoRA) requires identifying an update structure that is shared across heterogeneous clients. Prior work reports strong similarity among trained LoRA projection matrices across clients; however, such agreement may be largely induced by common initialization and collapses

    benchmark
  1093. arxiv:2609.32481 · cs.LG
    JEPA Learns What the Mask Leaves Unrecoverable
    Peng Xie, Amr Alanwar

    Joint-embedding predictive architectures are unusually sensitive to how the input is masked: block masks work, scattered masks do not, and the explanations are empirical. We give a measurement account. A mask is a linear measurement, and in a compactly supported wavelet basis every atom whose suppor

    v-jepa
  1094. arxiv:2609.32474 · cs.CL
    PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization
    Ziyi Zhang, Shuang Cui, Haotian Zhang, Xiaoyu Wang

    While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle'' phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, m

    long-contextbenchmark
  1095. arxiv:2609.32473 · cs.AI
    VPEvolve: A Self-Evolving Virtual Process Engineer for Computational Lithography
    Tianyi Li, Wenxuan Dong, Donger Luo, Nan Wang +4

    Optical proximity correction (OPC) recipes grow as engineers add local rules to repair newly discovered lithography hotspots. Each correction can interact with existing rules, while lessons from commercial-tool trials remain scattered across code and logs. \system combines a Virtual Process Engineer

    self-evolvingbenchmark
  1096. arxiv:2609.32472 · cs.CL
    AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
    Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen +5

    Document rerankers determine what evidence reaches the downstream model in RAG and deep research, yet mainstream rerankers select by relevance matching, and individually relevant documents rarely constitute the complete, complementary, non-redundant set a complex information need demands. Prior work

    ragbenchmark
  1097. arxiv:2609.32471 · cs.RO
    GlowTact: Simple and Compact Vision-Based Tactile Sensing with High Sensitivity and Spatial Resolution
    Yuxiang Ma, Megha Tippur, Pengfei Ye, Sandra Q. Liu +3

    Vision-based tactile sensors (VBTS) provide rich contact information for robotic manipulation, but existing designs can be hard to simplify and adapt to the size and constraints of humanoid fingertips. We introduce \textbf{GlowTact}, a pressure-responsive vision-based tactile sensing mechanism that

    manipulationhumanoidtactile
  1098. arxiv:2609.32470 · cs.LG
    On the Pitfalls of Verbalized Confidence Priors for Calibrating Large Reasoning Models
    Shuoyuan Wang, Beier Luo, Hao Zeng, Chengyao Yu +4

    Large reasoning models (LRMs) often suffer from overconfidence when expressing their uncertainty. Confidence-aware reinforcement learning (RL) offers a promising way to optimize calibration. However, it relies on on-policy rollouts and is thus constrained by the model's pre-RL confidence distributio

    benchmark
  1099. arxiv:2609.32469 · cs.LG
    PULSE: Identifying Demonstration-Utility Features with Sparse Autoencoders
    Chenduo Hao, Chuanbao Gao, Pinjun Zeng, Jingze Zhu +3

    In-context learning is highly sensitive to demonstration choice, yet most methods select demonstrations using external query-demonstration similarity. Such criteria can miss model-specific signals: Similar demonstrations may activate different internal features and downstream behaviors. We introduce

    benchmark
  1100. arxiv:2609.32467 · cs.LG
    Bison: Cross-Dataset Learning for Unseen-Compound Perturbation Prediction
    Yunfan Liu, Kasra Ghorbani, Yufei Huang, Zicheng Liu +5

    Predicting transcriptional responses to unseen compounds is limited by fragmented chemical coverage and heterogeneous experimental platforms and gene panels. To assess molecular generalization across these settings, we build on Chem-PerturBridge to benchmark eight datasets with 16,771 compounds, wit

    benchmark
  1101. arxiv:2609.32465 · cs.LG
    PolyStepOR: Learning to Decide Without Optimal Decisions
    Viet The Nguyen, Gunther Gust, An Thai Le

    Decision-focused learning (DFL) trains predictors for downstream decision quality, but often relies on optimal reference decisions that are expensive to obtain. We present PolyStepOR, which trains directly from realized decision costs without pre-computed optima and extends to in-constraint predicti

    benchmark
  1102. arxiv:2609.32464 · cs.LG
    Adapting Nonstationary Multi-output Gaussian Processes to Bayesian Optimization
    Zikai Xie

    Multi-objective Bayesian optimization (MOBO) commonly relies on independent Gaussian processes (GPs) with stationary kernels, limiting its ability to represent nonstationary structure and share information between objectives. However, expressive nonstationary GPs do not necessarily make reliable BO

    benchmark
  1103. arxiv:2609.32462 · cs.CV
    Can Motion-Language Models Ground Structure? STRIDE for Evaluating the Evaluators
    Lixing Tan, Qing Xia, Yuting Guo, Shuai Li +1

    Motion-language models are typically scored by motion-language evaluators, but how well these evaluators ground language structure remains unclear. Here, we introduce the Structure grounding via Temporal-order, Reflection, and Identity Diagnostic Evaluation (STRIDE) benchmark to systematically evalu

    benchmarkevaluatorevaluation protocol
  1104. arxiv:2609.32460 · cs.CV
    REMEDY: How Far Is Video Generation from Medical Education World Models?
    Lixing Tan, Yanghao Zhou, Qing Xia, Yuting Guo +2

    Recent video generation models produce realistic videos and show potential as a foundation for world models. These advances create opportunities for generating medical teaching demonstrations, which requires both convincing visual quality and precise procedural actions. However, whether current gene

    world modelbenchmark
  1105. arxiv:2609.32458 · cs.CL
    Streamlined Reflective Evolution for Task-Adaptive Self-Refinement Pipelines
    Xiaofan Zhou, Lu Cheng

    Reflective prompt optimization improves large language model (LLM) systems without updating model weights, but fixed architectures constrain how self-refinement is organized. We introduce Workflow-Designing Agents (WDA), a framework for streamlined reflective evolution of task-adaptive self-refineme

    agenticself-refinementbenchmark
  1106. arxiv:2609.32457 · cs.LG
    Write Back the $Δ$: Revisiting the Same Tokens with Fresh Representations
    Wencheng Ye, Anning Hu, Xiangdong Zhang, Tianyi Wang +4

    Transformers process information strictly forward through depth, preventing deeper computation from revisiting and refining earlier representations. To augment the standard forward pass, existing approaches either re-execute depth, incurring additional computation, or modify the residual stream usin

    benchmark
  1107. arxiv:2609.32456 · cs.CV
    Seeing Parts, Reasoning about Worlds: Visual Inference under Partial Observation
    Wei Wang, Wenqiao Zhang, Yutong Lin, Jun Xiao +1

    World modeling under partial observation requires reasoning about the complete worlds that remain compatible with limited visual evidence. Occluded objects and unseen regions can leave several world states possible; additional views can exclude alternatives and strengthen the conclusions supported b

    world model
  1108. arxiv:2609.32453 · cs.RO
    DRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation Policies
    Xinyu Zhao, Yixiang Shan, Tao Yang, Runyu Lei +4

    Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation win

    manipulationmemorymemory modulepost-training
  1109. arxiv:2609.32449 · cs.CL
    Self-Reports Do Not Identify Self-Models: An Identifiability Test for Counterfactual Reports
    Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz

    Language-model self-reports are evidence about behavior in a prompt environment, not by themselves evidence of a self-model. We investigate counterfactual reports about affect-like states under activation interventions and ask whether the report remains bound to the named intervention when the demon

    benchmark
  1110. arxiv:2609.32448 · cs.AI
    ForkLeft: Entropy-First Rollouts for Prefix-Aligned Autoregressive-to-Diffusion Distillation
    Junming Liu, Jicheng Wang, Yifeng He, Hao Chen +1

    Autoregressive Next-Token Prediction (NTP) has enabled strong reasoning capabilities in language models, while Diffusion Language Models (DLMs) offer flexible token orders and parallel generation. We ask whether DLMs can acquire NTP-style reasoning through distillation without giving up their native

    benchmark
  1111. arxiv:2609.32447 · cs.LG
    Length-Independent State Tracking Under a Parallel Scan
    Julien Brandoit, Arthur Fyon, Thomas Braipson, Tom Clara +4

    Learning robust and scalable finite-state tracking is fundamental to sequence processing. While linear recurrent neural networks (RNNs), linear attention, and state space models enable scalable parallel training through affine recurrences, their theoretical expressivity guarantees assume idealized a

    benchmark
  1112. arxiv:2609.32444 · cs.LG
    Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It
    Tianrun Yu, Kaixiang Zhao, Shangzhe Li, Yuxiao Yang +3

    We study training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To acco

    benchmark
  1113. arxiv:2609.32434 · cs.AI
    From Latents to Wires: Surgical Post-Editing on Large Language Models
    Jiankai Jin, Xiangzheng Zhang, Zhao Liu, Wenzhuo Xu +3

    Given a large language model (LLM), can whoever holds the weights name a semantic target (e.g., the model's identity), locate the model components that produce it, and edit them so that the target no longer appears while other capability is preserved? We call such an edit on a trained model a post-e

    post-training
  1114. arxiv:2609.32430 · cs.AI
    Multi-Agent System Search via Active Substructure-aware Policy Optimization
    Beicheng Xu, Bowen Fan, Weitong Qian, Lingching Tung +1

    LLMs enable multi-agent systems (MAS) to tackle complex tasks, but manually designing agent roles, prompts, and communication structures requires substantial expertise and effort. This motivates learning policies that construct query-specific MAS from execution reward. Existing approaches typically

    agentmulti-agentagent systembenchmark
  1115. arxiv:2609.32429 · cs.LG
    PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers
    Yanlong Chen, Yining Chen, Song Zhang, Amirhossein Habibian +1

    Smaller activation outliers do not necessarily imply better low-bit quantization: their alignment with the quantizer matters. We introduce PrismQuant, a quantizer-aware rotation framework that aligns the leading activation eigenspace with the constant group subspace of asymmetric grouped INT4. The a

    memory
  1116. arxiv:2609.32428 · cs.AI
    Authorization Closure Graph: Minimal Repair for LLM Agents with Evolving User Instructions
    Qingzhuo Wang, CaiYi Wang, Jinglu Meng, Ruiyang Qin +3

    Tool-using large language model (LLM) agents increasingly perform state-changing actions that require user authorization. Yet existing approaches do not provide a principled mechanism for selectively updating prior authorization when only part of an instruction changes. To this end, we propose an Au

    llm agent
  1117. arxiv:2609.32426 · cs.LG
    HoTS: Homophily-Aware Temperature Scaling for Graph Neural Network Calibration
    Inwoo Tae, Yoontae Hwang, Yongjae Lee

    For graph node classification, calibrated class probabilities are needed when confidence scores, usually the maximum predicted class probability, are used to rank predictions, defer uncertain nodes to human review, or control risk. Existing post-hoc calibrators either apply one global temperature or

    benchmark
  1118. arxiv:2609.32424 · cs.AI
    CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance
    Qi Chen, Fushuo Huo, Hangli Shen, Jingcai Guo +2

    Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulne

    long-contextagentllm agentmulti-agentagent systembenchmark
  1119. arxiv:2609.32423 · cs.AI
    PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins
    Yaorui Shi, Yuchun Miao, Yuxin Chen, Jiayuan Zhang +4

    The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomiz

    agent
  1120. arxiv:2609.32416 · cs.RO
    RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation
    Jiawei Zhang, Xiangrong Zhang, Rui Song, Huanbin Zhou +2

    Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the superv

    embodiedcode-as-policy
  1121. arxiv:2609.32407 · cs.AI
    Opening LLM Judges: Recovering Preference Signals Beyond the Final Verdict
    Sourabrata Mukherjee, Sunayana Sitaram

    LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal re

    benchmarkevaluator
  1122. arxiv:2609.32405 · cs.CV
    Toward On-Chip Training of Spiking Neural Networks for Dense Event-Based Vision
    Maxime Vaillant, Axel Carlier, Lai Xing Ng, Christophe Hurter +1

    Event cameras provide low-latency, asynchronous visual sensing for resource-constrained robotics. Spiking neural networks (SNNs) process event streams naturally, but training deep SNNs with backpropagation through time (BPTT) requires substantial memory and remains difficult on neuromorphic hardware

    memorybenchmarkevent camera
  1123. arxiv:2609.32401 · cs.AI
    Shared Worlds, Private Minds: Structured Memory for Long-Form Writing as World Creation
    Qiuyu Tian, Xiaowen Gu, Hang Su, Jianghan Chao +8

    LLM agents that write long-form fiction need an explicit memory of the evolving storyworld to keep new events consistent with established facts. Such memory must keep heterogeneous narrative information distinct, integrate story developments across granularities, and recover dependencies that a writ

    memoryllm agentbenchmark
  1124. arxiv:2609.32400 · cs.AI
    SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback
    Pengyu Zhu, Jingyi Yang, Yi Liu, Li Sun +1

    Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill m

    agent
  1125. arxiv:2609.32396 · cs.CL
    FA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy Conditions
    Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han +10

    Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices o

    benchmark
  1126. arxiv:2609.32395 · cs.LG
    Memory as a cache: Exact context reuse and deletion by construction
    Shengyao Wang, Jiang Liu

    The KV cache of a transformer entangles every token's representation with its entire prefix: a passage encoded once cannot be reused under a different prefix or removed without recomputing everything after it, so exact cache reuse is limited to shared prefixes. We present SMem, an architecture whose

    memory
  1127. arxiv:2609.32394 · cs.AI
    Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization
    Minghao Li, Rui Tan, Ruihang Wang

    Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate

    manipulationagentllm agentagentic
  1128. arxiv:2609.32391 · cs.LG
    SCLATE: a Substrate for Continual-Learning Agent Training and Evaluation
    Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan +2

    Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such as session stop and start, crons, and memory consolidation. Yet existing benchmarks and training fra

    memoryagentbenchmark
  1129. arxiv:2609.32390 · cs.AI
    Reward Hacking and Agent Containment Failure: A Monte Carlo Study Based on the 2026 Hugging Face Incident
    Murat Ozer, Bulent Erenay, Ibrahim Berber

    The July 2026 intrusion into Hugging Face production infrastructure showed how reward hacking can become an external cybersecurity incident when a capable agent encounters weak containment boundaries. This study develops a probabilistic risk model linking five stages: reward hacking, containment esc

    agentai agent
  1130. arxiv:2609.32389 · cs.CV
    RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion
    Sai Sri Teja Kuppa, Parth Shinde, Priyadharsan Balaji S, Jinka Harshavardhan +1

    Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tok

    memory
  1131. arxiv:2609.32385 · cs.AI
    DashAct: A Progressive Diagnostic Benchmark for GUI Agents in Interactive Dashboard Analysis
    Chuhan Zhang, Qi Xie, Ziyue Wang, Jianing Yin +3

    Interactive dashboards require users to reveal and connect evidence across stateful interactions. Although graphical user interface (GUI) agents could automate this process, existing dashboard benchmarks primarily report final answers or task success. They provide limited insight into whether failur

    agentbenchmark
  1132. arxiv:2609.32379 · cs.LG
    Measurement Boundaries in LLM Financial Agent Evaluation: Fixed-Tape Execution and Multi-Defect Auditing
    Weicheng Xue

    What controls are needed to interpret execution performance and audit scores in LLM agent evaluations? We study two limits on these interpretations in a financial agent harness. In Study~A, comparing independent runs under idealized and stressed execution on three synthetic settings that share one 2

    agentllm agent
  1133. arxiv:2609.32378 · cs.AI
    AuthorityLens: Rethinking LLM-Based Agent Systems Through the Lens of Authority
    Shaojin Chen, Huihao Jing, Wun Yu Chan, Wenbin Hu +6

    LLM-based agents are increasingly deployed with authority over consequential resources and decisions in real systems. These agents often operate alongside human and LLM-based participants who hold different forms of authority. Yet workflow roles, permission settings, and review mechanisms do not nec

    agentagent system
  1134. arxiv:2609.32367 · cs.LG
    STRIDE: State-Transition Representation via Increment Dynamics and Evolution
    Yuchen Xiong, Siming Huang, Jianfeng Sun

    We introduce STRIDE (State-Transition Representation via Increment Dynamics and Evolution), which defines states through derivative fingerprints and learns local functions for state transitions (qpairs), recasting continuous forecasting as transition prediction. Trailing convolution windows estimate

    benchmark
  1135. arxiv:2609.32365 · cs.LG
    Graph Memory: Spectral Associative Memory via Dirichlet Energy
    Zhaoyang Shi

    Dense associative memories have traditionally focused on storing and retrieving vector-valued patterns. Many modern machine learning problems, however, are naturally graph-structured, requiring memory mechanisms for relational patterns, graph diffusion geometries, community structures, and graph-bas

    memory
  1136. arxiv:2609.32363 · cs.LG
    DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting
    Weiwei Ye, Dongyuan Li, Hangchen Liu, Haotong Jiang +2

    Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model (DDPM)-based approaches have shown promise by equipping the dif- fusion process with pretrained mean and variance estimators to accommodate

    benchmark
  1137. arxiv:2609.32362 · cs.CV
    StegGNN: Learning Graphical Representation for Image Steganography
    Abhinav Kumar, Shorya Singhal, Agam Pandey, Tushar Kumar +1

    Image steganography refers to embedding secret messages within cover images while maintaining imperceptibility. Recent advances in deep learning - primarily driven by Convolutional Neural Networks (CNNs) and architectures such as inverse neural networks, autoencoders, and generative adversarial netw

    benchmark
  1138. arxiv:2609.32361 · cs.LG
    Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation
    Derui Wang, Zewei Shi, Rayne Holland, Ruoxi Sun +3

    Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradation, in which adaptation preserves monitor

    multi-agentbenchmark
  1139. arxiv:2609.32354 · cs.RO
    Proactive Motion Planning for Human-Robot Cooperation
    Elena Basei, Edoardo Lamon, Matteo Saveriano, Daniele Fontanelli +1

    This abstract addresses the incorporation of human motion prediction into proactive and dynamic human-aware motion planning, with the goal of enabling safe collaboration between humans and robots. A deep learning, graph-based model is used to forecast human motion and is integrated into a planning f

    manipulator
  1140. arxiv:2609.32353 · cs.LG
    Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
    Junxian Li, Ruixuan Yang, Tianao Zhang, Tiange Xu +2

    Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step fur

    benchmark
  1141. arxiv:2609.32352 · cs.CV
    EyeVQA: Benchmarking Ophthalmic Vision-Language Models from Recognition to Spatial Grounding
    Gujie Shao, Zixun Xie, Xuechun Xing, Ruixiang Wang +6

    Vision-language models (VLMs) have shown increasing potential for medical image understanding, yet their capabilities in ophthalmic imaging remain insufficiently characterized. Existing ophthalmic datasets are typically designed for individual diseases or specialized tasks, making it difficult to sy

    benchmark
  1142. arxiv:2609.32344 · cs.AI
    ALLOT: Budgeted Hybrid-Memory Routing for Knowledge Updates in LLMs
    Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du

    For large language models (LLMs), parametric adaptation is costly when retrieval already suffices. We introduce ALLOT, a hybrid-memory routing framework that separates learned write priority from a hard parametric budget. A memory-aware router combines frozen text representations, retrieval confiden

    external memorybenchmark
  1143. arxiv:2609.32343 · cs.CV
    OpenMASC: An Open-Source Pipeline for Cross-Trajectory Metal-Aware Sampling and Correction in Accelerated MRI
    Zhengyi Lu, Ming Lu, Chongyu Qu, Junchao Zhu +10

    Metal implants corrupt MRI measurements throughout $k$-space, yet existing accelerated MRI methods assume clean data and most metal artifact reduction approaches assume fully sampled acquisitions. No public dataset provides paired $k$-space and images with and without metal for the same anatomy, and

    agent
  1144. arxiv:2609.32341 · cs.LG
    A Comparative Analysis of Attention versus State-Space Models for In-Context Learning
    Enes Arda, Semih Cayci, Atilla Eryilmaz

    Transformers and state-space models (SSMs) are two prominent sequential learning architectures, yet their comparison remains largely empirical and existing theoretical analyses are typically task-specific or architecturally restricted. In this paper, we develop belief geometry, a unified analytical

    memory
  1145. arxiv:2609.32339 · cs.AI
    Enabling Timely Guidance before Skill Retrieval: Retaining Helpful Warm Tips in Agent Context
    Feng Liang, Yupeng Li, Runhao Zeng, Francis C. M. Lau +1

    Reusable skills help LLM-based agents solve complex tasks, but the agent must receive guidance before it commits to an ineffective approach. Existing skill mechanisms often expose only metadata and load full content on demand, leaving useful guidance unavailable until the agent decides to retrieve i

    memoryagentbenchmark
  1146. arxiv:2609.32336 · cs.RO
    WSM-Aware HRI: An IoT-Enhanced Framework for Early Detection and Norm-Guided Repair of Failures with LLM Guidance
    Hanlin Zhang, Yuquan Wang, Tianwei Zhang, Zhenglong Sun

    Human-robot interaction (HRI) failures remain a major barrier to deploying robots in real-world environments. Prior work often treats failures as isolated technical faults or focuses on post-hoc recovery behaviors. In practice, many breakdowns arise because humans and robots operate under inconsiste

    world model
  1147. arxiv:2609.32333 · cs.CV
    Progressive-View On-Policy Distillation for Regional-to-Global Transfer in Multimodal LLMs
    Shanfeng Huang, Zhou Fang, Song Xiao, Hai Du

    Regional-to-global distillation uses crop-conditioned guidance to improve full-image understanding. The challenge is to effectively transfer the teacher's crop-based advantage to the student's full-image inference. We propose progressive-view on-policy distillation (PVD), which shifts the student's

    benchmark
  1148. arxiv:2609.32327 · cs.AI
    HyperReCo: Retrieving and Connecting Evidence with Hypergraph Neural Networks for LLM Multi-hop Reasoning
    Zicheng Zhao, Linhao Luo, Junnan Dong, Haoran Luo +3

    Large language models (LLMs) have shown strong capabilities, with retrieval-augmented generation (RAG) supporting complex multi-hop reasoning by retrieving evidence distributed across documents. Graph-based approaches exploit connections among evidence, and hypergraph-based retrieval further preserv

    retrieval-augmentedbenchmark
  1149. arxiv:2609.32325 · cs.LG
    Active Feature Acquisition With Incomplete Training Data
    Reza Rezvan, Valter Schütz, Han Wu, Linus Aronsson +1

    In many prediction tasks, acquiring all features can be a prohibitively expensive or outright impossible task. Further, in many cases a static subset of features may not be enough to solve the problem sufficiently across various instances. Active Feature Acquisition (AFA) addresses these problems by

    benchmark
  1150. arxiv:2609.32322 · cs.LG
    Not All Errors Matter: Decision-Relevant Prediction Error Predicts Planning Quality
    Linhao Wang, Yiyan Fan, Dongjin Huang

    World models are typically trained and evaluated by prediction error, assuming that more accurate predictions lead to better decisions. We show that this assumption can fail because models with similar total error can differ substantially in planning performance when their errors occur on different

    world modelevaluation protocol
  1151. arxiv:2609.32318 · cs.LG
    What Can a Leaderboard Certify? Compositional Controllability for Fair Evaluation and Training of Biomedical Literature-Review Agents
    Zhaowei Han, Xiang Zhang, Lingxiao Guan, Danqi Hu +3

    Leaderboards rank long-horizon agents by their final outputs. Yet a higher score alone does not establish whether two systems are comparable or which stage accounts for the difference. Unequal evidence, inputs, or budgets can affect scores, and statistical corrections do not remove this mismatch. We

    ai agentbenchmarkleaderboard
  1152. arxiv:2609.32316 · cs.CV
    One Perception, All Maneuvers: Directional Traffic Signal Understanding for Maneuver-Level Signal Intent Prediction
    Ang Zou, Runzhe Zheng, Zhigang li, Zhen Yang +4

    Traffic lights are a key regulatory signal for autonomous driving at urban intersections, yet existing traffic signal perception is still predominantly formulated as instance-level detection or color recognition. Such formulations identify where traffic lights are and what colors they display, but l

    benchmark
  1153. arxiv:2609.32313 · cs.RO
    MemTransfer: Benchmarking Memory Beyond Matched Experience in Embodied Decision-Making
    Haiming Tang, Xianjie Dai, Gujie Shao, Zuyi Guo +4

    Memory lets an embodied agent reuse past experience, yet retaining useful information does not ensure that the agent can apply it when conditions change. We present MemTransfer, a benchmark comparing six memory representations, a working-memory baseline and five representations of past experience, u

    embodiedmemoryepisodic memoryagentembodied agentbenchmark
  1154. arxiv:2609.32312 · cs.AI
    Delayed Supervision for Test-Time Language Models
    Jinha Kim, Taksh Kothari

    Test-time language models adapt a compact memory while processing the input sequence. This perspective encompasses nonlinear fast-weight learning in LaCT, associative delta-rule updates in DeltaNet, and generalized delta-rule state updates in RWKV-7. Training these models to predict the next token d

    memorypost-training
  1155. arxiv:2609.32303 · cs.LG
    Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging
    Jingyuan Huang, Zuming Huang, Yucheng Shi, Zhongzhi Li +3

    Domain experts trained from a shared checkpoint can be merged into one model through on-policy distillation (OPD), where they act as teachers supervising a student on its own trajectories. One upstream choice is rarely examined: whether to build each expert with supervised fine-tuning (SFT) or reinf

    agentic
  1156. arxiv:2609.32297 · cs.AI
    Agentsensus: Consensus-Compressed Shared Memory for Multi-Agent Story Worlds
    Yu Pan

    A agentic story world is a dynamic system simulating who learned what, when, and from whom -- yet the standard design gives each character a private memory stream. A shared event is therefore stored once per witness, large duplication will be incurred in terms of storage. We present Agentsensus, a s

    memorymulti-agentagentic
  1157. arxiv:2609.32296 · cs.LG
    FUND: Density Flow for Sampling Unnormalised Distributions
    Vikas Kanaujia, Vipul Arora

    Efficient sampling from Boltzmann distributions is central to modelling complex physical systems. Markov Chain Monte Carlo (MCMC) methods suffer from critical slowing down, high autocorrelation, and poor mode-mixing, limiting their scalability. Recent advances, like Boltzmann Generators, offer a pro

    benchmark
  1158. arxiv:2609.32295 · cs.AI
    GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents
    Wei Zhu, Yiming Wang, Rui Wang, Lixing Yu +2

    LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalib

    agentllm agent
  1159. arxiv:2609.32292 · cs.RO
    Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation
    Xuekang Yang, Lu Chen, Shuang Luo, Jialing Zhu +3

    Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the

    embodiedagent
  1160. arxiv:2609.32289 · cs.LG
    Neural ODEs Meet Concurrent Learning: Stable Online Learning with Lyapunov Guarantees
    Omkar Sudhir Patil

    Neural ODEs learn dynamics from trajectory losses, but their adjoint gradients lack the regressor-times-parameter-error structure on which Lyapunov analyses of online adaptation rest, so training on streaming data comes without stability guarantees. We show that this structure is in fact present: th

    online learning
  1161. arxiv:2609.32288 · cs.LG
    A Solvable Theory of Pre-training Data Poisoning: Regime-Dependent Scaling Exponents
    Indranil Halder, Rastri Dey, Cengiz Pehlevan

    Pre-training data poisoning of large language models is usually studied using targeted backdoors and their survival through safety post-training, which leaves open a more basic question: how does a model's clean data performance degrade as the poison rate $\varepsilon$ grows? Motivated by our contro

    post-training
  1162. arxiv:2609.32280 · cs.CV
    HeroFrame-Bench: Reference-Anchored Evaluation via Rubric--Ranking Co-Evolution for Movie Hero Frame Selection
    Weitai Kang, Hanieh Deilamsalehy, Yumo Xu, Dewang Sultania +2

    Hero frames are in-film stills used as source imagery for theatrical posters, streaming cover art, film database listings, and other promotional placements. As the first visual entry point, they shape audiences' initial impressions of the movie and their subsequent willingness to watch it. Selecting

    benchmark
  1163. arxiv:2609.32276 · cs.LG
    HyperLabel: Multi-Label Classification via Hypergraph-Based Label Correlation Modeling
    Peiyu Zhang, Heng Ping, Nikos Kanakaris, Yucheng Zhao +4

    Multi-label classification (MLC) requires predicting multiple relevant labels for each instance, where a central challenge is modeling complex label dependencies arising from co-occurrence patterns. Existing approaches are limited in capturing high-order label correlations, relying on implicit learn

    benchmark
  1164. arxiv:2609.32275 · cs.LG
    Topology-Adaptive Hyperbolic Graph Attention Networks Guided by the Hyperbolic Sombor Index
    Haifang Cao, Boan Tao, Xiyuan Gao, Timing Li +2

    Hyperbolic geometry has emerged as a principled space for representing hierarchical graphs. However, existing hyperbolic graph neural networks typically rely on shared curvature configurations and feature-driven attention, failing to explicitly exploit local hierarchical topological patterns. To bri

    benchmark
  1165. arxiv:2609.32274 · cs.AI
    When Does a Skill Add Value? Task-Conditional Gain Prediction for Selective Skill Use
    Anjie Xu, Zhiyu Zhang, Ruiqing Ding, Fengli Xu +1

    Agent skills are expected to improve task performance. Yet we find that they often provide no benefit, and can even hurt performance while incurring additional token costs. Can we predict whether a skill will help before the agent acts? We introduce SkillDelta, a framework for predicting task-condit

    agentbenchmark
  1166. arxiv:2609.32269 · cs.LG
    Before Answering: Evidence Sufficiency under Size-Matched Memory Construction
    Joyanta Jyoti Mondal, Md. Shifatul Ahsan Apurba, Mridul Banik, Md Masud Al Mahmud +1

    Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memor

    memorybenchmark

02 US SEMI · SEC 8-K FILINGS

2 items

scanned: NVDA / AVGO / MRVL / COHR / LITE / AMD / TSM / SMCI / ANET / CRDO / POWL / VECO

  1. $AMD · 8-K · filed 2026-09-28
    Advanced Micro Devices Inc
    Items: 3.02
    8-K
  2. $MRVL · 8-K · filed 2026-09-25
    Marvell Technology Inc
    Items: 8.01,9.01
    8-K

03 HUMANOID · COMPANY NEWS

54 items

scanned: figure-ai / 1x / boston-dynamics / unitree / apptronik / sanctuary-ai / neura-robotics / agility-robotics / physical-intelligence / agibot

04 CN PHOTONICS · 公告流

0 items
CN 源 尚未实装 (TIER-1 下一步)